Redundant power flow constraint screening method based on spatio-temporal feature learning

The method leverages GCN and TCN to enhance the identification of redundant flow constraints in SCUC by integrating spatial and temporal features, reducing computational complexity and improving model accuracy.

CN120320291APending Publication Date: 2025-07-15SOUTHEAST UNIV +2
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510362937.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-26
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

In the existing safety constraint unit combination problem (SCUC), the current constraint screening method lacks in-depth research, especially in the direction of constraint screening by mining the spatial and temporal characteristics of the data, resulting in increased computational complexity and inefficiency.

Method used

Using a method based on spatiotemporal feature learning, a model combining graph convolutional neural network (GCN) and temporal convolutional network (TCN) is used to integrate the spatial and temporal features of current distribution to identify redundant current constraints. Specific steps include data preparation, preprocessing, spatial feature extraction, temporal feature extraction, model training and optimization, and use transformed loss functions to improve model performance.

Benefits of technology

It realizes accurate identification of redundant trend constraints, reduces the complexity of SCUC calculations, and improves the solution efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120320291A_ABST
    Figure CN120320291A_ABST
Patent Text Reader

Abstract

The invention discloses a redundant power flow constraint screening method based on spatial-temporal feature learning, and the method comprises the steps: preparing a power grid topological structure, node load, power flow distribution and constraint state data, carrying out the preprocessing of the data through a sliding window mechanism, carrying out the deep learning of spatial-temporal features through a GCNN (Graph Convolutional Neural Network) and a TCN (Time Convolutional Neural Network), and carrying out the prediction of the spatial-temporal features. Therefore, redundant power flow constraints are effectively identified, and model parameters are further debugged through prediction and evaluation of a test set, so that prediction precision and model stability are improved. According to the method, the accuracy of redundant power flow constraint screening can be improved in a complex power grid environment, and service is provided for optimizing the power grid dispatching efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of power system dispatching, and particularly relates to a redundant power flow constraint screening method based on spatio-temporal feature learning. Background Art

[0002] There are a variety of security constraints involved in the Security-Constrained Unit Commitment (SCUC) problem, mainly including power flow constraints, generator output constraints, frequency constraints, stability constraints, etc. Many constraints in the SCUC problem are redundant, that is, they do not affect the final unit commitment scheme. Power flow constraints account for a large proportion among the redundant constraints. In some cases, they may not affect the optimal solution, but greatly increase the computational complexity. Therefore, an effective power flow constraint screening method is particularly important. By screening the constraints highly relevant to the current dispatching scheme, the number of constraints in the solution process can be effectively reduced, and the solution efficiency can be improved.

[0003] The existing constraint screening methods are mainly divided into two categories: screening methods based on physical models and screening methods based on data-driven. However, there is a lack of in-depth research on constraint screening methods from the perspective of considering constraint characteristics. Especially in the direction of realizing constraint screening by mining spatio-temporal characteristics of data, the existing research has not fully solved the key challenges. Summary of the Invention

[0004] Object of the Invention: The present invention provides a redundant power flow constraint screening method based on spatio-temporal feature learning, which realizes the accurate identification of redundant power flow constraints by fusing the spatial and temporal features of power flow distribution.

[0005] Technical Solution: A redundant power flow constraint screening method based on spatio-temporal feature learning according to the present invention includes the following steps:

[0006] Step 1, Prepare data; Prepare historical data of the power grid topology structure, node load demand, and line power flow distribution; If the power flow of a certain line reaches its maximum transmission capacity, it is considered that the constraint of this line is bound; Then, count whether the power flow constraints of each line at each moment are bound. If bound, mark it as 1, and if not bound, mark it as 0;

[0007] Step 2, Data preprocessing; Use a sliding window mechanism to divide the input features (power grid topology structure, load demand, and power flow distribution) and label features (power flow constraint binding status) of the input data into multiple time-step segments, and perform standardization and normalization processing on the data within each time step to eliminate the influence of different dimensions;

[0008] Step 3, Spatial feature extraction: Use a graph convolutional neural network (GCN) to perform spatial feature extraction on the data at a single time step to capture the spatial dependencies between power grid nodes. The input data is the adjacency matrix and the load feature matrix, where the adjacency matrix represents the power grid topology, and each element represents the connection relationship between nodes;

[0009] Step 4, Temporal feature extraction: Input the node features extracted by the GCN into a temporal convolutional network (TCN) model, and use causal convolution, dilated convolution, and residual connections to perform time series feature modeling to capture the dynamic evolution law of node loads in the time dimension;

[0010] Step 5, Train the TCN-GCN model: Use the feature window as the input to predict the power flow binding status of the corresponding label window. Adopt a modified loss function and appropriate strategies to improve the model training efficiency and generalization performance;

[0011] Step 6, Input the test set data into the trained model to obtain the predicted values (0 or 1) of whether the power flow constraint is bound or not. Then, compare the predicted values with the true values and calculate metrics such as classification accuracy, recall rate, and F1 score;

[0012] Step 7, According to the test results, further optimize the key parameters of the model, including the number of GCN layers, the size of the TCN convolutional kernel, the learning rate, and the batch size, to improve the model performance and accuracy.

[0013] Furthermore, in Step 1, the historical data includes, but is not limited to, unit loads, unit characteristics, and the power flow calculation results after unit combination; the judgment method for binding is as follows: if the power flow of a certain line reaches the maximum transmission capacity of this line, then this line constraint is a binding constraint; count whether the power flow constraints of each line at each moment are bound. If bound, record it as 1, and if not bound, record it as 0. Denote the variable p as the power flow of a certain line at a certain moment, and the variable P max as the maximum transmission capacity that this line can withstand. If p ≥ P max or p ≤ -P max then the state of the power flow constraint of this line at this moment is bound. If -P max < p < P max then it is not bound.

[0014] Furthermore, in Step 2, the data preprocessing specifically includes the following steps:

[0015] Step 21, The dimension of the load feature matrix is where 360 represents a time span of 360 days, 24 represents that each day contains 24 time steps, and 118 represents 118 load nodes in the power grid. The dimension of the corresponding label matrix is Among them, 186 represents the power flow status of 186 transmission lines in the power grid;

[0016] Step 22: The sliding window technique is used to construct the input feature sequence and the prediction label sequence. Each window contains the features of consecutive time steps. The length of the input feature window is set to 72 hours, the prediction length is set to 24 hours, and the sliding step (stride) is set to 24 hours. The input feature sequence and the prediction label sequence are as shown in the following formula:

[0017] χ = [X i , X i+1 ,..., X i+T-1

[0018] Y = [Y i+T , Y i+T+1 , …, Y i+T+pre_len-1

[0019] In the formula: T represents the total time span of the feature window. In the 118-node system, 72 hours can be selected; χ represents the load feature matrix of each node within this time span; X i represents the load feature matrix at the i-th time step; Y represents the corresponding prediction label sequence, including the power flow binding status data for the next pre_len time steps. Here, pre_len is 24; Y i represents the prediction label at the i-th time step.

[0020] Step 23: Standardization and normalization processing; the node feature matrix X (load demand data) is standardized to scale the feature values to the range of [0, 1] to reduce the difference in the range of node load feature values; then, the load features of each node are further normalized to a standard distribution with a mean of 0 and a variance of 1.

[0021]

[0022] In the formula: max(X) is the maximum value of the load demand data among all nodes; x i is the feature value of a certain node at time step i; N is the number of time steps; μ is the average value of the feature values of this node.

[0023] Furthermore, in step 3, the spatial feature extraction specifically includes the following steps:

[0024] Step 31: The input of the GCN network is the adjacency matrix A and the node feature matrix X. Among them, the adjacency matrix A represents the power grid structure and indicates the connection relationship between nodes; the node feature matrix X is the initial feature representation of each node for each time step, obtained from step 2;

[0025] Step 32: The spatial feature propagation rule of graph convolution is shown in the following formula:​​

[0026]

[0027] Where: H (l) is the node feature matrix of the l-th layer, which is the feature representation of all nodes in the power grid topology graph at the l-th layer. Among them, the node feature matrix H (0) of the initial layer = X; W (l) is the learnable parameter matrix of the l-th layer, which is used for feature transformation; is the sum of the adjacency matrix and the identity matrix, that is is the degree matrix of is a diagonal matrix, where the elements on the diagonal represent the degree of each node, that is, the number of neighbor nodes directly connected to the node, that is σ is the activation function, specifically the tanh hyperbolic tangent function, which is integrated into the non-linear model;

[0028] Step 33: The features of each time step after extraction are organized into a four-dimensional tensor X ∈ R B×T×N×G , where B is the batch size, T is the number of time steps, N is the number of nodes, and C is the feature dimension.

[0029] Furthermore, in step 4, the time feature extraction specifically includes the following steps:

[0030] Step 41: Input the four-dimensional tensor obtained in step 3 into the TCN model with time as the main axis, and use causal convolution, dilated convolution, and residual connection to model the time series features, capturing the dynamic dependence and cross-time scale causal relationship of the load data;

[0031] Step 42: The mathematical formula of the causal convolution in the TCN model is shown as follows:

[0032]

[0033] Where: y[t] represents the output at the t-th moment, x[t - i] represents the input at the t - i-th moment, w[i] represents the weight of the convolution kernel, k represents the size of the convolution kernel, and the output y[t] is only related to the inputs x[t], x[t - 1], …, x[t - (k - 1)] at the current moment and previous moments, and is not affected by the future inputs x[t + 1], x[t + 2]…;

[0034] Step 43: The mathematical formula of the dilated convolution in the TCN model is shown as follows:

[0035]

[0036] Where: F(s) represents the output at the s-th time step, x s-d·iDenote the value of the input sequence at s-d·i, f(i) represents the i-th weight in the convolutional kernel, k represents the size of the convolutional kernel, and d represents the dilation rate, that is, the jump interval between adjacent convolutional kernel elements;

[0037] Step 44, the residual connection in the TCN model is mathematically shown as follows:

[0038] O = Activation(F(x) + x)

[0039] In the formula: O is the output of the residual block, F(x) is the output after a series of non-linear transformations of the input x through causal convolution, dilated convolution, activation function, and Dropout, x is the original input of the residual block, Activation is the non-linear activation function ReLU. This design allows information to be passed across different layers, avoiding the problem of gradient disappearance caused by the deepening of the network layers.

[0040] Further, in step 5, the process of predicting the corresponding label window power flow binding state by the node feature matrix H is as follows:

[0041] The node feature matrix H after time feature extraction by the TCN network is passed to the edge through a mapping mechanism to form a line feature representation, and its mathematical expression is as follows:

[0042] e ij = g(h i , h j ; W)

[0043] In the formula: e ij is the feature representation of line (i, j); h i , h j are the feature vectors of the two end nodes of the line, W is the learnable parameter matrix; g is the feature aggregation function;

[0044] In actual implementation, a fully connected layer is used to map the node features; the weight matrix W r is used to perform a linear transformation on the node feature matrix H, and finally the feature matrix E of the line is obtained. This process is as follows. After the mapping by the fully connected layer, the shape of the line feature matrix E is M×F', where M is the number of lines and F' is the dimension of the edge feature;

[0045] E = HW r

[0046] Further, in step 5, the modified loss function specifically includes the following steps:

[0047] Step 51: The loss function provides an optimization direction for the model by calculating the prediction error, calculates the gradient during backpropagation, and gradually updates the model's parameters, including the weight coefficients and the weight matrix of the classifier; the optimization objective of the model is to minimize the predicted over-line probability p i from the difference with the true label y i . Map the network output to the probability space (0, 1) through the Sigmoid function, and then use the binary cross-entropy loss function (BCE) for classification training. Its formula is shown as follows:

[0048]

[0049] In the formula: M is the number of lines; y i is the true label of the i-th line, taking values of 0 (invalid constraint) or 1 (valid constraint); p i is the predicted value of the over-limit probability of the i-th line by the model;

[0050] Step 52: Optimize the design of the loss function; first, calculate the loss directly on the model output without being processed by the Sigmoid function to solve the calculation redundancy and potential gradient disappearance risk existing in the Sigmoid + BCE combination in actual situations. The mathematical expression is shown as follows:

[0051]

[0052] In the formula: N is the number of samples; z i is the original output of the i-th sample;

[0053] Step 53: On the basis of the loss function in Step 52, introduce the joint optimization design of the focal loss function and the regularization term. The focal loss function is used to solve the problem that the model pays insufficient attention to positive samples, and the regularization term can limit the unconstrained growth of model parameters and suppress the overfitting phenomenon. The final loss function expression is shown as follows:

[0054]

[0055] L Total = L Focal + L Reg

[0056] L Focal is the focal loss function: L BCE,i is the binary cross-entropy loss of the i-th sample; p t is the confidence of the current predicted probability; α is the balance factor used to adjust the weights of positive and negative samples; γ is the focusing parameter used to amplify the loss contribution of difficult-to-classify samples and improve the recall rate of positive samples; L Reg is the regularization term: λ is the regularization strength control factor, and w is the trainable parameter of the model; LTotal is the final loss function, which is jointly composed of the focal loss function and the regularization term.

[0057] Furthermore, in step 6, accuracy and recall are used to evaluate the precision of redundant and effective power flow constraint screening, as shown in the following formula:

[0058]

[0059] In the formula: N correct is the number of samples correctly predicted by the model, and N total is the total number of all samples; TP is the number of samples with a true label of positive class and a predicted label of positive class, and FN is the number of samples with a true label of positive class and a predicted label of negative class.

[0060] Beneficial effects: Compared with the prior art, the present invention has the following remarkable advantages: First, historical data of the power grid topological structure, node load demand, and line power flow distribution are prepared; Second, it is statistically determined whether the power flow of each line at each moment reaches its maximum transmission capacity. If it reaches, the constraint of that line is considered to be bound and marked as 1; if it does not reach, the power flow constraint of that line is considered to be unbound and marked as 0; Then, a TCN-GCN framework is built; Information such as the power grid topological structure, load demand, and power flow distribution is used as the input data of the GCN to extract the spatial features of the line power flow distribution; The spatial features extracted by the GCN are used as the input data of the TCN to extract the temporal features of the line power flow distribution; Thus, by fusing the spatial and temporal features of the power flow distribution, accurate identification of redundant power flow constraints is achieved. Brief Description of the Drawings

[0061] Figure 1 is a schematic diagram of the method flow of the present invention.

[0062] Figure 2 is a schematic diagram of constructing the input feature sequence and the prediction label sequence by the sliding window mechanism of the present invention.

[0063] Figure 3 is a schematic diagram of spatial feature extraction by the GCN network of the present invention.

[0064] Figure 4 is a schematic diagram of causal convolution of the TCN network of the present invention.

[0065] Figure 5 is a schematic diagram of dilated convolution of the TCN network of the present invention. Detailed Embodiments

[0066] As Figure 1 shown, a method for screening redundant power flow constraints based on spatio-temporal feature learning includes the following steps:

[0067] S1. Prepare data. Prepare historical data of the power grid topological structure, node load demands, and line power flow distributions; if the power flow of a certain line reaches its maximum transmission capacity, it is considered that the constraint of this line is bound; then, count whether the power flow constraints of each line at each moment are bound. If bound, mark it as 1; if not bound, mark it as 0.

[0068] S2. Data preprocessing. Use a sliding window mechanism to divide the input features (power grid topological structure, load demands, and power flow distributions) and label features (power flow constraint binding status) of the input data into multiple time-step segments, and perform standardization and normalization processing on the data within each time step to eliminate the influence of different dimensions.

[0069] S3. Spatial feature extraction. Use a graph convolutional neural network (GCN) to perform spatial feature extraction on the data of a single time step to capture the spatial dependence relationships between power grid nodes. The input data is the adjacency matrix and the load feature matrix, where the adjacency matrix represents the power grid topological structure, and each element represents the connection relationship between nodes.

[0070] S4. Temporal feature extraction. Input the node features extracted by GCN into a temporal convolutional network (TCN) model, and use causal convolution, dilated convolution, and residual connections to perform time series feature modeling to capture the dynamic evolution law of node loads in the time dimension.

[0071] S5. Train the TCN-GCN model. Use the feature window as the input to predict the power flow binding status of the corresponding label window. Adopt a modified loss function and a suitable strategy to improve the model training efficiency and generalization performance.

[0072] S6. Input the test set data into the trained model to obtain the predicted values (0 or 1) of whether the power flow constraint is bound. Then, compare the predicted values with the true values and calculate metrics such as classification accuracy, recall rate, and F1 score.

[0073] S7. According to the test results, further optimize the key parameters of the model, including the number of GCN layers, the size of the TCN convolutional kernel, the learning rate, and the batch size, etc., to improve the model performance and accuracy.

[0074] In one embodiment, in S2:

[0075] The dimension of the load feature matrix is where 360 represents a time span of 360 days, 24 represents 24 time steps per day, and 118 represents 118 load nodes in the power grid. The corresponding dimension of the label matrix is where 186 represents the power flow status of 186 transmission lines in the power grid.

[0076] AsFigure 2 As shown, the sliding window technique is used to construct the input feature sequence and the prediction label sequence. Each window contains consecutive time-step features. The window length is set to 48 or 72 hours, the prediction length is set to 24 hours, and the sliding stride is set to 24 hours. For example, the first time window starts from time step [t0,t1,…,t 71 , and the prediction range is [t 72 ,t 73 ,…,t 95 ; the second time window is [t 24 ,t 25 ,…,t 95 , and the prediction range is [t 96 ,t 97 ,…t 119 . This sliding is repeated until the entire time series is covered. The input feature sequence is shown in Equation (1), and the prediction label sequence is shown in Equation (2).

[0077] χ = [X i ,X i+1 ,...,X i+T-1 (1)

[0078] Y = [Y i+T ,Y i+T+1 ,…,y i+T+pre_len-1 (2)

[0079] Where: T represents the total time span of the feature window. In the 118-node system, 72 hours can be selected; χ represents the load feature matrix of each node within this time span; X i represents the load feature matrix at the i-th time step; Y represents the corresponding prediction label sequence, including the power flow binding state data for the next pre_len time steps. Here, pre_len is 24; Y i represents the prediction label at the i-th time step.

[0080] Standardization and normalization processing. As shown in Equation (3), the node feature matrix X (load demand data) is standardized to scale the feature values to the range [0,1] to reduce the difference in the range of node load feature values. Then, as shown in Equation (4), the load features of each node are further normalized to a standard distribution with a mean of 0 and a variance of 1.

[0081]

[0082] Where: max(X) is the maximum value of the load demand data among all nodes; x i is the feature value of a certain node at time step i; N is the number of time steps; μ is the average value of the feature values of this node.

[0083] In one of the embodiments, the detailed process of spatial feature extraction in S3 is as follows:

[0084] As Figure 3 shown, the input of the GCN network is the adjacency matrix A and the node feature matrix X. Among them, the adjacency matrix A is the power grid structure, representing the connection relationship between nodes; the node feature matrix X is the initial feature representation of each node at each time step.

[0085] The spatial feature propagation rule of graph convolution is shown in Equation (5).

[0086]

[0087] In the formula: H (l) is the node feature matrix of the l-th layer, which is the feature representation of all nodes in the power grid topology graph at the l-th layer. Among them, the node feature matrix H of the initial layer (0) = X; W (l) is the learnable parameter matrix of the l-th layer, used for feature transformation;

[0088] is the sum of the adjacency matrix and the identity matrix, that is is the degree matrix of is a diagonal matrix, where the elements on the diagonal represent the degree of each node, that is, the number of neighbor nodes directly connected to this node, that is σ is the activation function, specifically the tanh hyperbolic tangent function, which is integrated into the non-linear model.

[0089] The features at each time step after extraction are organized into a four-dimensional tensor X ∈ R B×T×N×G , where B is the batch size, T is the number of time steps, N is the number of nodes, and C is the feature dimension.

[0090] In one of the embodiments, the detailed process of temporal feature extraction in S4 is as follows:

[0091] The principle of causal convolution in the TCN model is as Figure 4 shown, and the mathematical expression is shown in Equation (6).

[0092]

[0093] In the formula: y[t] represents the output at the t-th moment, x[t-i] represents the input at the t-i-th moment, w[i] represents the weight of the convolution kernel, and k represents the size of the convolution kernel. The output y[t] is only related to the input x[t], x[t-1], …, x[t-(k-1)] at the current moment and previous moments, and is not affected by the input x[t+1], x[t+2] … at future moments.

[0094] The principle of dilated convolution in the TCN model is as Figure 5 shown. Mathematically, dilated convolution is shown in Equation (7):

[0095]

[0096] In the formula: F(s) represents the output at the s-th time step, and x s-d·i represents the value of the input sequence at s - d·i, f(i) represents the i-th weight in the convolutional kernel, k represents the size of the convolutional kernel, and d represents the dilation rate, that is, the jump interval between adjacent convolutional kernel elements.

[0097] The residual connection in the TCN model is shown in Equation (8):

[0098] O = Activation(F(x) + x) (8)

[0099] In the formula: O is the output of the residual block, F(x) is the output after a series of non-linear transformations of the input x through causal convolution, dilated convolution, activation function, and Dropout, x is the original input of the residual block, and Activation is the non-linear activation function ReLU. This design allows information to be passed between different layers, avoiding the problem of gradient disappearance caused by the deepening of the network layers.

[0100] In one of the embodiments, the process of predicting the corresponding label window power flow binding state from the node feature matrix H is as follows:

[0101] The node feature matrix H after time feature extraction by the TCN network is passed to the edge through a mapping mechanism to form a line feature representation. Its mathematical expression is shown in Equation (9).

[0102] e ij = g(h i , h j ; W) (9)

[0103] In the formula: e ij is the feature representation of line (i, j); h i , h j are the feature vectors of the two end nodes of the line, W is the learnable parameter matrix; g is the feature aggregation function.

[0104] In actual implementation, a fully connected layer is used to map the node features. Using the weight matrix W r to perform a linear transformation on the node feature matrix H, and finally obtaining the line feature matrix E. This process is shown in Equation (10). After the mapping by the fully connected layer, the shape of the line feature matrix E is M × F', where M is the number of lines and F' is the dimension of the edge feature.

[0105] E = HW r (10)

[0106] In one of the embodiments, the loss function in S5 is as follows:

[0107] Based on the binary cross-entropy loss function, a joint optimization design of the focal loss function and the regularization term is introduced. The binary cross-entropy loss function is shown in Equation (11), the focal loss function is shown in Equations (12) and (13), the regularization term is shown in Equation (14), and the final loss function is shown in Equation (15).

[0108]

[0109] L Total = L Focal + L Reg (15)

[0110] Where: L BCE,i is the binary cross-entropy loss of the i-th line; N is the number of samples; y i is the true label of the i-th line, taking values of 0 (invalid constraint) or 1 (valid constraint); z i is the original output of the i-th line; L Focal is the focal loss function; α is the balance factor used to adjust the weights of positive and negative samples; p t is the confidence of the current prediction probability; γ is the focusing parameter used to amplify the loss contribution of difficult-to-classify samples and improve the recall rate of positive samples; L Reg is the regularization term; λ is the regularization intensity control factor, w is the trainable parameter of the model; L Total is the expression of the final loss function, which is jointly composed of the focal loss function and the regularization term.

[0111] In one of the embodiments, the evaluation metrics in S6 are as follows:

[0112] Accuracy and Recall are used to evaluate the precision of redundant and effective power flow constraint screening. The mathematical formula of Accuracy is shown in Equation (16), and the mathematical formula of Recall is shown in Equation (17).

[0113]

[0114] Where: N correct is the number of samples correctly predicted by the model, N total is the total number of all samples; TP is the number of samples with true label as positive class and predicted as positive class, FN is the number of samples with true label as positive class and predicted as negative class.

Claims

1. A redundant power flow constraint screening method based on spatio-temporal feature learning, characterized in that The steps are as follows: Step 1: Prepare data. Prepare historical data of the power grid topology structure, node load demands, and line power flow distributions. If the power flow of a certain line reaches its maximum transmission capacity, it is considered that the constraint of this line is bound. Then, count whether the power flow constraints of each line at each moment are bound. If bound, mark it as 1; if not bound, mark it as 0. Step 2: Data preprocessing. Use a sliding window mechanism to divide the input features and label features of the input data into multiple time-step segments, and perform standardization and normalization processing on the data within each time step. Step 3: Spatial feature extraction. Use the graph convolutional neural network GCN to perform spatial feature extraction on the data of a single time step to capture the spatial dependence relationship between power grid nodes. The input data is the adjacency matrix and the load feature matrix, where the adjacency matrix represents the power grid topology structure, and each element represents the connection relationship between nodes. Step 4: Temporal feature extraction. Input the node features extracted by GCN into the temporal convolutional network TCN model, and use causal convolution, dilated convolution, and residual connections to model the time series features to capture the dynamic evolution law of node loads in the time dimension. Step 5: Train the TCN-GCN model. Use the feature window as the input to predict the power flow binding status of the corresponding label window. Adopt a modified loss function and appropriate strategies to improve the model training efficiency and generalization performance. Step 6: Input the test set data into the trained model to obtain the predicted values of whether the power flow constraints are bound. Then, compare the predicted values with the true values, and calculate the classification accuracy, recall rate, and F1-score metrics. Step 7: According to the test results, further optimize the key parameters of the model, including the number of GCN layers, the size of the TCN convolutional kernel, the learning rate, and the batch size, to improve the model performance and accuracy.

2. The redundant power flow constraint screening method based on spatio-temporal feature learning according to claim 1, wherein In step 1, the historical data includes but is not limited to the unit load, unit characteristics, and the power flow calculation results after unit combination; the binding judgment method is as follows: If the power flow of a certain line reaches the maximum transmission capacity of this line, then this line constraint is a binding constraint; count whether the power flow constraints of each line at each moment are bound. If they are bound, record it as 1. If they are not bound, record it as 0. Denote the variable p as the power flow of a certain line at a certain moment, and the variable P max as the maximum transmission capacity that this line can withstand. If p ≥ P max or p ≤ -P max then the state of the power flow constraint of this line at this moment is bound. If -P max < p < P max then it is not bound.

3. The redundant power flow constraint screening method based on spatio-temporal feature learning according to claim 1, wherein In Step 2, the data preprocessing specifically includes the following steps: Step 21, the dimension of the load characteristic matrix is where 360 represents a time span of 360 days, 24 represents 24 time steps per day, and 118 represents 118 load nodes in the power grid. The dimension of the corresponding label matrix is where 186 represents the power flow status of 186 transmission lines in the power grid; Step 22: The sliding window technique is used to construct the input feature sequence and the prediction label sequence. Each window contains consecutive time-step features. The length of the input feature window is set to 72 hours, the prediction length is set to 24 hours, and the sliding step is set to 24 hours. The input feature sequence and the prediction label sequence are as shown in the following formula: Where: T represents the total time span of the feature window, which can be selected as 72 hours in the 118-node system; represents the load feature matrix of each node within this time span; X i represents the load feature matrix at the i-th time step; represents the corresponding prediction label sequence, including the power flow binding state data for the next pre_len time steps, where pre_len is 24; Y i represents the prediction label at the i-th time step; Step 23: Standardization and normalization processing. Standardize the node feature matrix X, that is, the load demand data, and scale the feature values to the range of [0, 1]. Then, further perform standard distribution normalization on the load features of each node to make its mean 0 and variance 1. Where: max(X) is the maximum value of the load demand data among all nodes; x i is the eigenvalue of a certain node at time step i; N is the total number of time steps; μ is the average value of the eigenvalues of this node.

4. The redundant power flow constraint screening method based on spatio-temporal feature learning according to claim 1, wherein In Step 3, the spatial feature extraction specifically includes the following steps: Step 31: The input of the GCN network is the adjacency matrix A and the node feature matrix X. The adjacency matrix A is the power grid structure, representing the connection relationship between nodes. The node feature matrix X is the initial feature representation of each node for each time step, obtained from Step 2. Step 32: The spatial feature propagation rule of graph convolution is as shown in the following formula: Where: H (l) is the node feature matrix of the l-th layer, which is the feature representation of all nodes in the power grid topology graph at the l-th layer. Among them, the node feature matrix H (0) of the initial layer = X; W (l) is the learnable parameter matrix of the l-th layer, which is used for feature transformation; is the sum of the adjacency matrix and the identity matrix, that is is the degree matrix of is a diagonal matrix, where the elements on the diagonal represent the degree of each node, that is, the number of neighbor nodes directly connected to this node, that is σ is the activation function, specifically the tanh hyperbolic tangent function, which is integrated into the non-linear model; Step 33, the features at each time step after extraction are organized into a four-dimensional tensor X ∈ R B×T×N×G , where B is the batch size, T is the number of time steps, N is the number of nodes, and C is the feature dimension.

5. The redundant power flow constraint screening method based on spatio-temporal feature learning according to claim 1, wherein In Step 4, the temporal feature extraction specifically includes the following steps: Step 41: Input the four-dimensional tensor obtained in Step 3 into the TCN model with time as the main axis, and use causal convolution, dilated convolution, and residual connection to model the time series features, capturing the dynamic dependencies of the load data and the causal relationships across time scales; Step 42: Mathematically, the causal convolution in the TCN model is shown as follows: In the formula: y[t] represents the output at the t-th moment, x[t-i] represents the input at the t-i-th moment, w[i] represents the weight of the convolution kernel, k represents the size of the convolution kernel, and the output y[t] is only related to the inputs x[t], x[t-1], …, x[t-(k-1)] at the current and previous moments, and is not affected by the inputs x[t+1], x[t+2] … at future moments; Step 43: Mathematically, the dilated convolution in the TCN model is shown as follows: where: F(s) represents the output at the s-th time step, x s-d·i represents the value of the input sequence at s - d·i, f(i) represents the i-th weight in the convolutional kernel, k represents the size of the convolutional kernel, and d represents the dilation rate, that is, the jump interval between adjacent convolutional kernel elements; Step 44: Mathematically, the residual connection in the TCN model is shown as follows: O = Activation(F(x) + x) In the formula: O is the output of the residual block, F(x) represents the output after a series of non-linear transformations of the input x through causal convolution, dilated convolution, activation function, and Dropout, x is the original input of the residual block, and Activation is the non-linear activation function ReLU, i.e., the rectified linear unit.

6. The redundant power flow constraint screening method based on spatio-temporal feature learning according to claim 1, wherein In Step 5, the process of predicting the corresponding label window power flow binding state from the node feature matrix H is as follows: The node feature matrix H after time feature extraction by the TCN network is passed to the edge through a mapping mechanism to form a line feature representation, and its mathematical expression is shown as follows: e ij = g(h i , h j ; W) where: e ij is the feature representation of line (i, j); h i ,h j are the feature vectors of the nodes at both ends of the circuit, and W is a learnable parameter matrix; g is the feature aggregation function; In actual implementation, a fully connected layer is used to map node features; the weight matrix W r is used to perform a linear transformation on the node feature matrix H, and finally the feature matrix E of the circuit is obtained. This process is shown in the following formula. After the mapping of the fully connected layer, the shape of the circuit feature matrix E is M×F', where M is the number of circuits and F' is the dimension of edge features; E = HW r .

7. The redundant power flow constraint screening method based on spatio-temporal feature learning according to claim 1, wherein In Step 5, the reformed loss function specifically includes the following steps: Step 51. The loss function provides an optimization direction for the model by calculating the prediction error, calculates the gradient in backpropagation, and gradually updates the parameters of the model, including the weight coefficients and the weight matrix of the classifier; the optimization objective of the model is to minimize the predicted over-line probability p i and the true label y i The difference is that the network output is mapped to the probability space (0, 1) through an explicit Sigmoid function, and then the binary cross-entropy loss function BCE is used for classification training. Its formula is shown as follows: Where: M is the number of lines; y i is the true label of the i-th line, taking values of 0 (invalid constraint) or 1 (valid constraint); p i is the predicted value of the model for the probability of the i-th line exceeding the limit; Step 52: The loss function is optimized and designed; first, the loss is calculated directly on the output of the model without being processed by the Sigmoid function to solve the calculation redundancy and potential gradient vanishing risk existing in the Sigmoid + BCE combination in actual situations, and its mathematical expression is shown as follows: where: N is the number of samples; z i where is the original output of the i-th sample; Step 53: Based on the loss function in Step 52, a joint optimization design of the focal loss function and the regularization term is introduced. The focal loss function is used to solve the problem that the model pays insufficient attention to positive samples, and the regularization term can limit the unconstrained growth of model parameters and suppress the overfitting phenomenon. The final loss function expression is shown as follows: L Total = L Focal + L Reg L Focal is the focal loss function: L BCE,i is the binary cross-entropy loss of the i-th sample; p t is the confidence of the current predicted probability; α is the balancing factor used to adjust the weights of positive and negative samples; γ is the focusing parameter, which is used to amplify the loss contribution of difficult-to-classify samples and improve the recall rate for positive samples; L Reg is the regularization term: λ is the regularization strength control factor, and w is the trainable parameter of the model; L Total is the final loss function, which is jointly composed of the focal loss function and the regularization term.

8. The redundant power flow constraint screening method based on spatio-temporal feature learning according to claim 1, wherein In Step 6, the accuracy Accuracy and recall Recall are used to evaluate the accuracy of redundant and effective power flow constraint screening, as shown in the following formula: Where: N correct is the number of samples correctly predicted by the model, and N total is the total number of all samples; TP is the number of samples with a true label of positive class and a predicted label of positive class, and FN is the number of samples with a true label of positive class and a predicted label of negative class.

Citation Information

Cited By

  • Power distribution network reconstruction optimization method for improving distributed new energy bearing capacity

    CN121840596A