Power system dynamic blocking high-credibility probability early warning method and device

By constructing a congestion risk quantization perception network using a hybrid attention mechanism and a tensor LSTM network, the problem of poor interpretability of multiple input variables in power systems is solved, and highly reliable prediction and decision support for power system congestion states are achieved.

CN121524924APending Publication Date: 2026-02-13WUHAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511668061.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-14
Publication Date
2026-02-13

AI Technical Summary

Technical Problem

Existing technologies in power systems have poor interpretability for multiple input variables, making it difficult to effectively predict and explain dynamic congestion events, resulting in passive and unreliable control methods.

Method used

A congestion risk quantification perception network is constructed by employing a hybrid attention mechanism, tensor LSTM network units, and static covariate knowledge information fusion. The model is trained using multidimensional time series data to predict the congestion status of the power system.

Benefits of technology

It improves the interpretability and credibility of the model, can accurately predict the congestion state of the power system, provides reliable decision support, and enhances the auxiliary reference value of dispatching operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121524924A_ABST
    Figure CN121524924A_ABST
Patent Text Reader

Abstract

The invention discloses a power system dynamic blocking high-credibility probability early warning method and device, and the method comprises the steps: obtaining multi-variable multi-dimensional time series data, and constructing a supervision data set; based on a mixed attention mechanism, a tensor LSTM network unit and a static covariable knowledge information fusion network, constructing a blocking risk quantitative sensing network; training a blocking risk quantitative sensing network based on the supervision data set to obtain a trained blocking risk quantitative sensing network; and inputting the multi-dimensional time sequence data to be measured into the trained blocking risk quantitative sensing network to predict the blocking state of the power system. The technical problem of poor explanation of multiple input variables in the prior art can be solved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of power system congestion early warning, and particularly relates to a power system dynamic congestion high-confidence probability early warning method and device. BACKGROUND

[0002] With the increasing demand for deep decarbonization worldwide and the continuous advancement of policies, the development of renewable energy and the electrification of traditional industries have made significant progress. However, the uneven development of source and load locations and the relatively lagging construction of transmission networks may lead to the modern power system often operating in a "tight balance" state close to the safety boundary, inducing regional section dynamic congestion events, and posing a significant risk challenge to the safe and stable operation of the system. Among them, section dynamic congestion refers to the transmission power of the inter-regional tie line approaching or exceeding the section transmission limit, while the transmission limit is the comprehensive dynamic upper limit calculated according to the static security and transient stability rules. Due to the insufficient current system perception level and the low degree of decision-making intelligence, when facing real-time dynamic congestion events, the existing regulation mode will fall into a passive situation in terms of resources and time, and there is a great regulation pressure. With the rapid development of data measurement and storage technology and information science theory, digital technology brings new ideas to solve the above problems. However, unlike other application fields of digital technology, the power industry is the backbone of national economic development, and the first priority is to ensure safe operation, which puts high requirements on the reliability of risk early warning models and the reliability of regulation and control decisions.

[0003] The reliability (confidence) of a data-driven model is a decisive factor for dispatching and operating personnel to adopt its output results, which in turn will affect its actual assistance to the regulation and control optimization process. Therefore, improving the reliability of data-driven models is the key to realizing their perception value. Currently, research work in this direction mainly focuses on improving the model's explainability, and researchers try to increase users' trust by deconstructing and understanding the prediction reasoning logic inside the model. However, the applicability of commonly used post-hoc explainability methods (SHAP and LIME) when facing time series data is still questionable. In their typical forms, the time sequence correlation of input features is not considered. For example, for LIME, its proxy model is built on each data point separately; for SHAP, data features are considered separately in adjacent time steps. Since there is usually significant correlation between adjacent time steps in time series data, the above post-hoc analysis methods may result in poor explanation effect, making it difficult to adapt to the explanation needs of multivariate time series learning models and help understand time dynamic trends.

[0004] Using a decision tree model is another commonly used explainability enhancement strategy, which makes decisions in a conditional reasoning manner and can be intuitively displayed through visualization tools, so it can be better understood by users to support the analysis and calculation process. However, the decision tree belongs to a shallow learning model, and the structure is relatively simple, so its performance is difficult to surpass that of a deep learning model when facing complex prediction tasks. At the same time, in practical applications, as the dimension of the feature variable grows rapidly, the complexity of the safe and stable evaluation rules will also increase significantly, and it is difficult to obtain an explainable evaluation rule using a shallow learning-based method. In recent years, attention mechanisms have attracted attention due to their ability to improve prediction performance and explainability, and some network architectures based on attention mechanisms have been gradually proposed, which can better adapt to time series data, such as the Transformer architecture. The core of the attention mechanism is to autonomously and dynamically enhance the feature weight related to the task and suppress the contribution proportion of redundant features, and the corresponding attention weight can be extracted to reveal the internal law of variable contribution and time dynamics, thereby helping to analyze the importance of different variables and different event development stages. However, in the current research, attention is mainly applied to the multiple time step hidden states calculated by the recurrent neural network layer and the feature vector information extracted by the convolutional neural network layer, and these data contents usually contain information from multiple input variables, which may cause deviation in the explanation of the importance of a single feature variable. SUMMARY

[0005] The application provides a power system dynamic congestion high-confidence probability early warning method and device, electronic equipment and medium, which solves the technical problem of poor explanation of multiple input variables in the prior art.

[0006] According to an aspect of the application, a power system dynamic congestion high-confidence probability early warning method is provided, comprising: obtaining multi-dimensional time series data of multiple variables and constructing a supervised data set; constructing a congestion risk quantitative perception network based on a hybrid attention mechanism, a tensorized LSTM network unit, and a static covariate knowledge information fusion network; training the congestion risk quantitative perception network based on the supervised data set to obtain a trained congestion risk quantitative perception network; inputting the to-be-tested multi-dimensional time series data into the trained congestion risk quantitative perception network to predict the congestion state of the power system.

[0007] According to another aspect of the application, a power system dynamic congestion high-confidence probability early warning device is provided, comprising: a data set construction unit configured to obtain multi-dimensional time series data of multiple variables and construct a supervised data set; A model construction unit is configured to construct a congestion risk quantitative perception network based on a hybrid attention mechanism, a tensorized LSTM network unit, and a static covariant knowledge information fusion network. A training unit is configured to train the congestion risk quantitative perception network based on the supervision data set to obtain a trained congestion risk quantitative perception network. A prediction unit is configured to input the to-be-tested multi-dimensional time series data into the trained congestion risk quantitative perception network to predict the congestion state of the power system.

[0008] The technical scheme of the embodiment of the present application comprises the following steps: obtaining multi-variable multi-dimensional time series data and constructing a supervision data set; constructing a congestion risk quantitative perception network based on a hybrid attention mechanism, a tensorized LSTM network unit, and a static covariant knowledge information fusion network; training the congestion risk quantitative perception network based on the supervision data set to obtain a trained congestion risk quantitative perception network; and inputting to-be-tested multi-dimensional time series data into the trained congestion risk quantitative perception network to predict the congestion state of the power system. The present application solves the technical problem of poor interpretability of the prior art for multiple input variables.

[0009] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present application, nor is it intended to limit the scope of the present application. Other features of the present application will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS

[0010] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0011] Figure 1 is a flow chart of a power system dynamic congestion high-confidence probability early warning method provided by an embodiment of the present application; Figure 2 is a schematic diagram of an update derivation process of a hidden state matrix based on tensor dot product in an embodiment of the present application; Figure 3 is a general architecture of a congestion risk perception network and an ETS confidence calibration module in an embodiment of the present application; Figure 4 is a time covariant attention weight histogram distribution diagram of a congestion risk perception network extracted in an embodiment of the present application; Figure 5 is a variable attention weight histogram distribution diagram of a congestion risk perception network extracted in an embodiment of the present application; Figure 6 Variable average time attention weight columnar heat map extracted by the congestion risk perception network in an embodiment of the present application; Figure 7 Schematic diagram of a dynamic congestion risk perception case in an embodiment of the present application; Figure 8 Schematic diagram of a dynamic congestion risk perception case in another embodiment of the present application; Figure 9 Probability estimation reliability graph of the congestion risk perception network before and after confidence calibration in an embodiment of the present application; Figure 10 Partial variable time sequence trajectory graph during a dynamic congestion event in an embodiment of the present application; Figure 11 Class prediction and probability estimation result graph of the congestion risk perception network without confidence calibration in an embodiment of the present application; Figure 12 Class prediction and probability estimation result graph of the congestion risk perception network after confidence calibration in an embodiment of the present application; Figure 13 It is a structural diagram of a power system dynamic congestion high-confidence probability early warning device according to an embodiment of the present application. DETAILED DESCRIPTION

[0012] In order to make the personnel in the art better understand the present application scheme, the technical solutions in the embodiments of the present application will be described clearly and completely below in combination with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by the person skilled in the art without creative labor should belong to the scope of protection of the present application.

[0013] It should be noted that the terms "first", "second", and the like in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily indicate a specific order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device including a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0014] Embodiment one Figure 1A flowchart of a power system dynamic congestion high-probability early warning method is provided for the first embodiment of the present application. As shown in Figure 1 the method comprises: S101, acquiring multi-dimensional time series data of multiple variables and constructing a supervised data set.

[0015] Among them, the multi-dimensional time series data is time series data including multiple variables. Specifically, the multiple variables can include state variables and control variables of the power system, etc.

[0016] In a specific embodiment, the multi-dimensional time series data can be extracted based on time window sliding, and after identifying the system dynamic congestion state of the multi-dimensional time series data, the cross section is marked to construct the supervised data set.

[0017] Among them, for the multi-dimensional time series data, the multi-dimensional time series data can be extracted by time window sliding to constitute a sample data, that is, the multi-dimensional time series data can be divided into sample data by time window, such as 10 seconds of multi-dimensional time series data as a sample data. It should be noted that the size of the time window in this embodiment can be set as needed. In addition, the multi-dimensional time series data can be discrete data obtained every second as sequence data, that is, 10 seconds of multi-dimensional time series data includes 10 discrete data, and each variable in the multiple variables in each sample data contains 10 time series point data. It should be noted that the time interval between adjacent sequence points in the time series data can be set as needed.

[0018] In addition, for each sample, a time point can be taken as a cross section to mark the current power system dynamic congestion state, which can include normal stage, early stage and dynamic congestion stage. Based on each marked sample data, a supervised data set is constructed.

[0019] In this embodiment, the cross section transmission power can be compared with the limit transmission capacity to determine whether the current is in a dynamic congestion state. Among them, the limit transmission capacity measures the ability of the power transmission network to reliably transmit power from one area to another area without compromising system safety. In this embodiment, the transmission dynamic congestion is considered as the cross section transmission power exceeding 95% limit transmission capacity, thereby indicating that the system safety margin is insufficient or facing safety risk.

[0020] Considering computational convergence and the flexibility of embedding complex security constraints, the ultimate transmission capacity can be calculated based on the Repeated Power Flow (RPF) method. During the calculation, AC power flow security, N-1 static security, and transient stability rules are comprehensively considered. Specifically, the calculation of the ultimate transmission capacity considering the above complex security constraints can be modeled as the following optimization problem: (1) (2) (3) (4) (5) (6) (7) In the formula, is the power generation and load growth index, used to adjust the power growth in the power generation and receiving areas; u is the system control variable, including active power injection at generator nodes, node voltage amplitude, and load power, etc. For system dependency variables; These are the state variables and algebraic variables of the system during the transient process, respectively. These are the sets of equality and inequality constraints before the fault, respectively. This represents the set of static equality constraints of the system after the N-1 fault is interrupted. For N-1 static security verification rules, Represents the set of N-1 line disconnection faults; Equation (4) represents the initial values ​​of the system state variables calculated by solving the power flow equations before the fault. Equations (5) and (6) represent the differential equation and algebraic equation in the transient process, respectively. For transient stability verification rules, For the purpose of anticipating accidents, For the transient stability study period, referring to existing research, this invention mainly considers transient power angle stability constraints and adopts the following judgment rules: (8) In the formula, and These represent the power angle of generator g and the power angle at the center of inertia after the fault is cleared, respectively. To constrain the upper limit, it is often set to 180 degrees. This represents the generator's inertia constant. The above-mentioned ultimate transmission capacity calculation model is based on a certain generation and load growth pattern, and is adjusted... To iteratively change the system's operating conditions: (9) where, is the sending-end area generator set, is the receiving-end area load set, , , , are the initial states of generator active power, voltage, load active power and load reactive power, respectively, , and represent the generator active power, voltage and load growth pattern, respectively. Models (1)-(9) can be solved by the RPF method, which continuously increases the unit output of the sending-end area and the load demand of the receiving-end area in the solving process, thereby continuously increasing the transmission of the cross-section power. At the same time, in each iteration, various safety constraints are comprehensively checked to obtain the cross-section transmission limit: (10) where, is the optimal growth index obtained by iterative calculation, is the set of inter-regional tie lines, is the limit transmission capacity of tie line l under the operation scenario , is a function for calculating the transmission power of tie line l. Therefore, by comparing the tie line power flow and the transmission limit under this operation scenario, it can be determined whether there is dynamic congestion.

[0021] S102, based on a hybrid attention mechanism, a tensorized LSTM network unit, and a static covariant knowledge information fusion network, a congestion risk quantification perception network is constructed.

[0022] In this embodiment, the hybrid attention mechanism can include time attention and variable attention, which is used to explain the role of the input physical variables on the feature and time level on the prediction results; the tensorized LSTM network unit is used to capture the dynamic historical information of each variable; the static covariant knowledge information fusion network is used to inject static covariant data information, so that the heterogeneous information fusion network has good data behavior explainability and dynamic congestion risk perception performance.

[0023] In an embodiment, the process in which the congestion risk quantification perception network processes the input multi-dimensional time series sample and static covariate data comprises: inputting the multi-dimensional time series sample into a tensorized LSTM network unit to obtain a corresponding hidden state matrix sequence; applying time attention to the historical hidden state sequence, i.e. weighting the hidden state sequence by time attention weights to obtain historical context vector information of each input feature variable; concatenating the historical context vector information with the current hidden state sequence to obtain concatenated dynamic historical information; inputting the static covariate to obtain a static enriched context vector; concatenating the dynamic historical information with the static enriched context vector to obtain a context tensor; calculating network scores and variable attention weights based on the context tensor; calculating final network scores based on the network scores and the variable attention weights; and outputting a congestion state prediction result based on the final network scores.

[0024] wherein inputting the multi-dimensional time series sample into a tensorized LSTM network unit to obtain a corresponding hidden state matrix sequence comprises: In this embodiment, the congestion risk quantification perception network based on hybrid attention is established on the basis of high-dimensional feature combination optimization, i.e. using the preferred feature variable combination obtained by high-dimensional feature combination optimization as model input. At the same time, based on the preferred feature variable combination , the original multi-variable time series data can be reconstructed to obtain the multi-dimensional time series data (multi-time series) supervision data set used by the congestion risk quantification perception network. Specifically, assuming that the feature variable combination contains variables, i.e. , . By retaining only the time series data of the variables, the original multi-dimensional time series supervision data set S can be reconstructed as : (11) wherein , , L is the length of each sequence segment in the data matrix, is the multi-variable data vector of the mth reconstructed multi-dimensional time series sample at the lth time step, , ; M is the total number of samples, and y is the congestion state label. In the remaining part of this section, for the sake of simplicity, the sample index m will be omitted.

[0025] The main idea of the tensorized LSTM network proposed in this embodiment is to develop a new network cell update mechanism, so that each element (e.g. each row) in the hidden state matrix can correspond to the information of a certain input variable. Specifically, the tensorized hidden state matrix at the l-th time step can be represented as: , ,

[0026] The overall size of the hidden layer is defined as , is the hidden size corresponding to each variable; element is the hidden state vector corresponding to the n-th input variable; for a given input vector at time step l and the hidden state matrix of the previous time step, the update rule of the cell includes: (12) (13) .(14) (15) The update structure of the tensorized LSTM cell includes the input gate , the forget gate , the output gate , and the memory cell state ; in which, , , , represents the update element of the hidden state with respect to the input variable n; represents the input-to-hidden state transition tensor, , , ; represents the transition tensor between hidden states at different time steps, , , ; represents the bias, ; similarly, represents the input-to-hidden state transition tensor corresponding to the gate , , , , , , ; represents the transition tensor between hidden states at different time steps, , , ; This indicates the corresponding gate. , , The bias, ; and They captured data from the previously hidden state. and new input data Knowledge and information; operation Indicates along the variable dimension The tensor dot product of two tensors is calculated as follows: (16) (17) tanh(.) represents the tanh activation function. Represents the activation function, operation This indicates element-wise multiplication.

[0027] As an example Figure 2 The hidden state matrix is ​​shown in detail. The derivation process, where different colors correspond to different input variable information. By performing the tensor dot product operation shown in equations (16) and (17), the input gate... Forgotten Gate Output gate and memory cell state Both are matrices, and have the same properties as... The same matrix shape, their elements are the same as the matrix Similarly, it can specifically correspond to information about a particular input variable. This is due to the gate matrix. and Only for and Scaling (see Equation (14)) can be performed, and the data organization corresponding to the variables can be passed to Similarly, due to the gate matrix Only for Scaling is performed, therefore in the hidden state matrix The data organization corresponding to the variables can be preserved, that is... It encodes only knowledge information derived from the input variable n. This characteristic allows the hidden state matrix to capture the dynamic pattern information of each variable separately, improving the traceability and transparency of data flow within the model and facilitating the interpretation of the data behavior characteristics of each input feature variable in conjunction with the attention mechanism.

[0028] The time attention is applied to the historical hidden state sequence, i.e. the hidden state sequence is weighted by the time attention weight, to obtain the historical context vector information of each input feature variable, including: wherein, for the multi-dimensional time series sample After inputting the tensorized LSTM network unit, the hidden state matrix sequence is obtained, wherein, for the hidden state sequence of the nth input feature variable (non-static covariate), i.e. the time attention is applied to the hidden state sequence of the input feature variable, i.e. , respectively; then, the variable attention is derived to combine these aggregated historical information corresponding to the input variable one by one, and the static rich context vector , so as to summarize the knowledge information learned by the model from the time and variable two levels.

[0029] The time attention weight is:

[0030] wherein, denotes the time attention weight corresponding to the nth input feature variable at the lth time step, ; denotes the feedforward neural network activated by the tanh(.) function corresponding to the variable n, and the subscript T represents the time attention; for the same input feature variable, the parameters are shared at different time steps. The time attention weight reflects the contribution factor of the data information of the input variable n at the time step l to the current prediction, and the time attention has the following numerical characteristics: , , .

[0031] The hidden state sequence is weighted by the time attention weight , to obtain the historical context vector information corresponding to the feature variable n, denoted as , ; The historical context vector information is spliced with the current hidden state sequence to obtain the spliced dynamic historical information; the static covariate is inputted to obtain the static rich context vector; the dynamic historical information and the static rich context vector are spliced to obtain the context tensor, including:

[0032] wherein, represents the context tensor; represents the vertical concatenation operation, represents the history context vector information of the concatenation feature variable n and the current hidden state information, represents the static rich context vector. As shown in Figure 3 , since the time attention is applied separately for the hidden state sequence of each input variable, and only data concatenation operation is performed subsequently, the data organization characteristics corresponding to the variables (similar to the hidden state matrix ) can still be retained in the context tensor , which is conducive to information backtracking.

[0033] Based on the network part logit and the variable attention weight in the context tensor, including: The variable attention weight is:

[0034] In the formula, represents the variable attention weight of the i-th variable, has the following numerical characteristics: , , ; represents the i-th row vector of the matrix , , ; when , ; when , ; is a feedforward neural network activated by the tanh(.) function shared by all variables, and the subscript V represents variable attention; in order to maintain the correspondence between the model information flow and the physical variable until the output layer, the variable attention will be applied to the network part logit (logit) calculated by the fully connected (FC) network corresponding to the physical variable.

[0035] For the i-th physical variable, the calculated network part logit is , is a fully connected network without activation function corresponding to variable i, responsible for mapping the data representation in the context matrix to the label space, and the subscript FC represents full connection; represents the total number of classes in the target classification problem; Based on the network part logit and the variable attention weight, the final network part logit is calculated, including:

[0036] wherein, denotes the final network score log vector calculated by weighting, ; Based on the final network score log output, the congestion state prediction result includes: The synthesized final network score log will be activated by a softmax layer, thereby forming a dynamic congestion event perception model:

[0037]

[0038] wherein, denotes the softmax activation function, denotes the synthesized network score log vector, containing elements, denotes the set of class labels, in the dynamic congestion event perception task, h is the class index, , denotes the probability estimate about class h output by the softmax layer, denotes the network score log corresponding to class h, ; Based on the probability estimate vector , the class prediction result (dynamic congestion event phase label) corresponding to the input sample (and static covariate information ) can be obtained, that is , , and the estimated confidence score corresponding to the class prediction result (indicating the probability of being correct, i.e., the probability of class occurrence).

[0039] Based on the above units, the congestion risk quantification perception network that integrates hybrid attention mechanism and tensorized LSTM units is formed, and the overall network architecture is shown in Figure 3 . For the sake of simplicity, this network will be referred to as MA-LSTMT network in the following text. Through the data organization structure corresponding to each data variable throughout the dynamic congestion risk quantification perception MA-LSTMT network, the contribution of each data variable to the model prediction can be automatically identified and accurately quantified through the hybrid attention mechanism, so that the variable attention weight and the time attention weight The input physical variables are explained in terms of their effects on the prediction results in the feature and time dimensions, where a higher attention weight indicates a greater impact. Meanwhile, the attention weights are dynamically allocated according to different input data, so as to adaptively focus on important parts of different input multi-dimensional time series samples.

[0040] S103, training the blockage risk quantification perception network based on the supervision data set to obtain a trained blockage risk quantification perception network.

[0041] The objective function used to train the blockage risk quantification perception network is to minimize the cross-entropy loss:

[0042] In the formula, m is a sample index, is the sample input corresponding to the final output network score log, is the real one-hot encoded label corresponding to the mth sample, , and only if the real class of the mth sample is h, ; otherwise .

[0043] In an embodiment, a confidence post-processing unit is further included, which calibrates the confidence output of the blockage risk quantification perception network without affecting the class prediction result by minimizing the mapping relationship between the heterogeneous information fusion network score log and the dynamic blockage event posterior probability.

[0044] In this embodiment, the probability output of the blockage risk perception network is modified by the confidence calibration post-processing technology. Considering that the high accuracy of the model is a basic requirement for dynamic blockage event perception, the confidence calibration process should not reduce the accuracy of the trained model. At the same time, since the dynamic blockage supervision data samples are usually limited, the confidence calibration process should have a certain data efficiency. In addition, considering the complexity of the probability calibration task itself, the calibration method should also have good data representation ability. In view of the above three aspects of the demand for the confidence calibration method: 1) accuracy maintenance; 2) data efficiency; and 3) data representation ability, this embodiment proposes a confidence calibration post-processing method based on ensemble temperature scaling (ETS), which is used to reconstruct the mapping relationship between the network score log and the dynamic blockage event posterior probability.

[0045] Specifically, for the input multi-dimensional time series data sample and the static covariate information , the synthesized network score log calculated by the model is , , Represents the number of network splits corresponding to category h; based on The corresponding softmax score vector can be obtained. This refers to the original probability vector estimated by the network, from which the predicted labels for the dynamic blocking event stages can be obtained. and the corresponding confidence score. :

[0046] Based on the true labels of the samples and network logarithmic vector The goal of confidence calibration is to produce corrected confidence scores. At the same time, maintain the category prediction results Unchanged. For example... Figure 3 As shown on the right, the ETS confidence calibrator is a device that includes optimizable parameters. and The three-component integrated post-processing module. For samples... and and network logarithm vector The ETS calibration mapping has the following form:

[0047] In the formula, This represents the post-calibration confidence level corresponding to category h after processing by the ETS confidence calibrator, u and The introduced calibration parameters to be optimized This indicates the total number of event phase label types, for ={C0: Normal phase, C1: Early phase, C2: Dynamic blocking phase} , This refers to the softmax function.

[0048] In the above formula Corresponding to the original temperature scaling confidence calibration method, this method introduces a parameter u, which is equivalent to maximizing the entropy of the output probability distribution under the network logarithm constraint, thereby alleviating the overconfidence problem of the original deep learning network; assuming given The network logarithm vector corresponding to each input sample and the corresponding real category labels The temperature scaling model is the unique solution to the following maximum entropy problem. :

[0049]

[0050]

[0051]

[0052] The above formula ensures is a probability distribution and restricts the range of the distribution. Intuitively, this constraint states that the average true class of the network score log equals the average weighted network score log.

[0053] In addition, and can be regarded as temperature scaling mapping, but the parameter u takes 1 and respectively. The second term aims to improve the stability of the confidence calibration when the probability estimates of the original model are already relatively reliable, avoiding reverse adjustment. The third term outputs uniform probabilities for each class, similar to the label smoothing technique. The ETS calibration parameters and can be determined by minimizing the overall negative log-likelihood loss on the calibration (validation) dataset, where the negative log-likelihood (NLL) is a standard metric for evaluating the quality of a probability model. The optimization model for solving the negative log-likelihood loss of the ETS calibration parameters and is shown below:

[0054]

[0055] In the formula, , is the number of multi-dimensional time series data samples used for confidence calibration, i.e., the number of validation dataset samples, is the model parameter of the trained MA-LSTMT congestion risk quantification perception network; these parameters are fixed during the confidence calibration process. The above nonlinear constraint optimization problem can be solved by gradient-based optimization algorithms.

[0056] Since the ETS confidence calibration module uses a convex combination of strictly monotonic functions, the relative size order of the elements in the probability estimate vector output by the original network can be preserved, so the class prediction results output by the calibrated model are consistent with the original network, and the accuracy of the original model is not affected:

[0057] In the formula, is the class prediction result output by the calibrated model, which is consistent with , This outputs the corresponding calibration confidence level. Compared to the original temperature scaling method, the ETS confidence calibration module adds only three additional parameters, namely... Therefore, theoretically, it should inherit the data efficiency of the temperature scaling method. Simultaneously, due to the introduction of additional calibration parameters, the data characterization capability of the ETS confidence calibration module is expected to be improved. The relationship between the ETS confidence calibration module and the MA-LSTMT network is as follows: Figure 3 As shown.

[0058] To evaluate the reliability of model probability estimates, in addition to reliability plots, the Expected Calibrated Error (ECE) and Maximum Calibrated Error (MCE) metrics are also commonly used. Specifically, based on the test dataset (containing... (samples), obtained from model estimation The confidence scores will be divided into R intervals, each interval having a size of R. The r-th interval can be represented as... , .use This indicates that the confidence estimate score falls within the interval The set of samples within, The classification accuracy of the samples within can be expressed as:

[0059] In the formula, Indicates the predicted label, Indicates the true label, Indicates falling within the interval The number of samples within. Simultaneously, it can be calculated. The average estimated confidence score within the sample set:

[0060] In the formula, This represents the calibrated confidence score.

[0061] based on and This allows for the calculation of the ECE index, which quantifies the deviation between the confidence estimate and the probability of correctness (weighted accuracy minus confidence difference). For R confidence intervals, the ECE index is calculated as follows:

[0062] In the formula, To test the number of samples, the difference between the accuracy probability (acc) and the confidence (conf) in an arbitrary interval reflects the calibration bias. For the interval , if , it means underconfidence; if , it means overconfidence. In addition, since the dynamic congestion risk perception belongs to the power system security auxiliary application, the robustness of the model has high requirements, and the worst case of confidence estimation needs to be evaluated. Therefore, the MCE index can be used to evaluate the maximum deviation between acc and conf:

[0063] In addition, as a standard measure to evaluate the quality of the probability model, the NLL index on the test data set will also be used to evaluate the confidence calibration performance. The confidence calibration optimization process will be carried out on the reserved validation data set, and the confidence calibration performance evaluation process will be carried out using the test data set. Specifically, the reserved multi-dimensional time series validation data set and static covariate information will be input into the trained MA-LSTM congestion risk quantification perception network, so as to calculate and store the corresponding network score log, that is, . Based on and the real dynamic congestion event stage label information , the parameters and can be determined by confidence calibration optimization, so as to complete ETS calibration.

[0064] Based on the congestion risk quantification perception network described above based on the fusion hybrid attention mechanism and the tensorized LSTM unit, and the confidence calibration module based on the ETS post-processing method proposed in this section, the credibility of the dynamic congestion risk perception model output decision can be comprehensively enhanced, and more reliable auxiliary reference information can be provided for dispatching personnel.

[0065] In a specific embodiment, static covariate data refers to variables that do not change much in the observation time window time scale, such as day information, month information, whether it is a holiday, and location information, etc. For the mth window sliding of the observation time window (OTW), the observation time window starts at time and ends at time + . Therefore, the static covariate information located at time + + can be added to the learning network to enrich the model knowledge. Let denote the static covariate data information corresponding to the mth observation time window, , , The total number of introduced static covariates. Due to the easy availability and advance knowability of time information variables (such as season), the present application mainly incorporates them as static covariates into the dynamic congestion quantification perception network. In order to make the dynamic congestion quantification perception network better understand the time covariates, the present application uses the trigonometric transformation to encode the periodic time information into a numerical vector, the calculation method is as follows:

[0066] In the formula, The period of the nth time covariate is denoted, where the sample index m is also omitted to simplify the description. By trigonometric transformation, the original time covariate data vector is transformed into :

[0067] In the formula, In order to make more generalized modeling expression, here d0 is used to represent the data dimension of a single static covariate after input encoding, and is used to represent the converted data, . Based on the encoded data, the context vector for static knowledge information fusion can be formed based on the tensorized feedforward neural network:

[0068] In the formula, The static rich context vector is denoted, , , The context vector of the nth static covariate is denoted, . The dimension of the single static covariate context vector is set to 2d, so as to be combined with the context vectors of other variables (referring to the preferred feature variables contained in by variable attention. The weight parameter tensor is denoted, , , , The bias parameter is denoted, . In order to establish a data organization structure corresponding to the input variables one by one, the tensor dot product operation is still used to calculate the context vector , so as to maintain the data information traceability of the congestion risk quantification perception network:

[0069] In the formula, is the encoded information corresponding to the nth static covariate, The data flow of the static covariant knowledge information fusion part is as shown in Figure 3

[0070] S104, input the to-be-tested multi-dimensional time series data into the trained congestion risk quantization perception network to predict the congestion state of the power system. In this embodiment, by inputting the to-be-tested multi-dimensional time series data into the trained congestion risk quantization perception network, the congestion state of the power system can be predicted.

[0071] The present application is directed to the high reliability probability early warning problem of power system dynamic congestion. Firstly, based on high-dimensional feature combination optimization, a heterogeneous information fusion network (MA-LSTM network) is constructed based on a mixed attention mechanism, a tensorized LSTM network unit and a static covariant knowledge information fusion method, which can be used for efficient online application of dynamic congestion risk quantization perception. Further, a confidence enhancement calibration method for output decision of the heterogeneous information fusion network is proposed, and a confidence integrated calibration post-processing method is used to improve the comprehensive confidence of the model output result, thereby enhancing its auxiliary reference significance to the subsequent decision-making process and practical application value. The present application has the following advantages: 1. Based on the tensorized LSTM network unit, the hidden state matrix corresponding to the input variable in the proposed heterogeneous information fusion network can help capture the dynamic data patterns from different variables, thereby supporting the explicit tracking of the encoded information flow from each input variable. 2. The mixed attention mechanism realizes the visualization analysis of the internal reasoning process of the model, and each variable and its attention weight at each time step can be extracted to analyze the model prediction reasoning logic, which can help users check whether the attention pattern of the heterogeneous information fusion network conforms to the domain cognitive experience, thereby measuring the trust degree of the model. 3. The static covariant knowledge information fusion method is integrated into the congestion risk quantization perception network, which can help conveniently inject covariant data information into the heterogeneous information fusion network, so that the heterogeneous information fusion network has good data behavior explainability and dynamic congestion risk perception performance. 4. The model confidence post-processing module based on the integrated scaling method is proposed, which can effectively improve the reliability of the model probability output, so as to reflect the real correctness likelihood of the prediction as much as possible, and obtain better probability calibration performance compared with other methods.

[0072] In an embodiment, a synthetic dynamic congestion data set can be simulated on an IEEE 39-node system, and based on this, the dynamic congestion perception performance of the heterogeneous information fusion network in the present application, the data behavior based on the mixed attention, the confidence calibration test of the model output, and the dynamic congestion risk quantization perception online simulation are comprehensively verified.

[0073] ​The proposed dynamic congestion risk perception model output decision credibility enhancement calibration method is verified by an example. The experiment is based on the IEEE 39-node test system multi-dimensional time series data supervision data set and the corresponding high-dimensional feature combination optimization results, verifying the effectiveness and advantages of the proposed method, and clearly showing the whole process of the proposed method.

[0074] First, the class prediction performance of the MA-LSTMT congestion risk quantification perception network is tested, compared with commonly used multi-dimensional time series deep learning models, and the effect of integrating the proposed static covariate knowledge information is verified. This section experiment is based on the IEEE 39-node test system multi-dimensional time series data supervision data set and the corresponding high-dimensional feature combination optimization results, all methods use the same input feature variables and data sets for testing. Specifically, the feature variables shown in Table 1 are used to reconstruct the multi-dimensional time series supervision data set, only the time series data of the variables in the table are retained. The reconstructed multi-dimensional time series supervision data set has 3548 samples, of which the number of samples belonging to the "normal stage", "early stage" and "dynamic congestion stage" is 1640, 710 and 1198 respectively, the number of variables contained in each multi-dimensional time series data sample is 16, and the time step is 12 (t = 180 min). The above samples can be divided into three parts: training set (contains 1896 samples), validation set (contains 722 samples) and test set (contains 930 samples). In order to integrate static covariate knowledge based on time information, each sample is additionally attached with an absolute time information, if the observation time window corresponding to the mth multi-dimensional time series sample starts at time , and ends at time + , the absolute time information attached to the sample is taken from + + According to this absolute time information, the following time static covariates are used in this example: ① the number of hours in a day , ; ② the number of days in a week , ; ③ the number of weeks in a month , ; ④ month , ; ⑤ season , . Therefore, the input data variables used by the MA-LSTMT dynamic congestion risk perception model in this example are shown in Table 1.

[0075] Table 1 ​

[0076] In this example, the commonly used multi-dimensional time series deep learning models for comparison include: ① standard long short-term memory network LSTM: containing two LSTM hidden layers and two fully connected layers, respectively with {100, 50, 50, 3} neurons; ② Gated Recurrent Unit (GRU) network: containing two GRU hidden layers and two fully connected layers, respectively with {100, 50, 50, 3} neurons. In order to test the effect of the proposed static covariate knowledge information integration, this example also includes the MA-LSTMT network without time information integration, which is denoted as MA-LSTMT(D) model for subsequent convenience. In the MA-LSTMT and MA-LSTMT(D) networks, the hidden size corresponding to each feature variable is set to 32, the learning rate is set to 0.001, and the step decay strategy is used, reducing the learning rate by 10% every 60 epochs. Other settings are the same for the above models (maximum period: 300, batch size: 20, Dropout Rate: 20%, loss function: cross-entropy loss, optimizer: Adam). LSTM, GRU and MA-LSTMT(D) models use the same input features, i.e. the non-static covariate features in Table 1. For the MA-LSTMT network, the periodic time information is encoded into a numerical vector using the triangular transformation, so that the dynamic congestion quantization perception network can better understand the time covariates, therefore, Each method will be run 10 times to report the average results, and the experimental comparison results are shown in Table 2, including the average accuracy (AvgAcc), F-beta score (AvgF-beta), recall, precision of the "early stage" class samples, and the average training time of the model on the test data set.

[0077] Table 2

[0078] ​​​As shown in Table 2, in addition to the training time, the dynamic blockage event perception model based on the MA-LSTMT network achieved better test performance, with an average accuracy of dynamic blockage event perception of 96.645% and an F-beta score of 0.966. For the "early stage" category samples, the present method increased the recall rate and precision rate from 88.960% and 92.820% of the MA-LSTMT(D) network to 91.040% and 93.398%, respectively, which can more accurately help identify the early stage of dynamic blockage events, thereby more effectively supporting subsequent dynamic blockage prevention control. In terms of parameter size, after adding 5 time covariates, the MA-LSTMT network has 75280 trainable parameters, which is only 1935 (2.638%) more than the 73345 of the MA-LSTMT(D) network; in terms of training time, the average model training time of the MA-LSTMT network is only 36.964 seconds more than that of the MA-LSTMT(D).

[0079] Data behavior explanation based on hybrid attention: This section will extract the hybrid attention weights in the MA-LSTMT blockage risk perception network, so as to visualize and analyze the specific data behavior of the model at the variable level and the time level.

[0080] (1) Global pattern analysis based on hybrid attention Based on the MA-LSTMT blockage risk perception network trained, input the multi-dimensional time series sample and the corresponding time covariate information , the hybrid attention weight values corresponding to the sample can be calculated and extracted, including variable attention weights and time attention weights . In this example, , (see Table 1). This experiment then inputs the training data, validation data and test data into the MA-LSTMT blockage risk perception network, and can collect the hybrid attention weight sets for each data set, i.e. , , and , , . In order to analyze the global data characteristics of the variables, Figure 4 the variable attention weight distribution histograms of the time covariate (time information) and the covariate (week information) on the three data sets are drawn, respectively. In Figure 5 , the variable attention weight distribution histograms of the variable (variables) of (Table 1, No. 3) and (variables) of (Table 1, No. 14) are shown in the left histogram. Figure 5 The vertical dashed line represents the average value of the attention weight.

[0081] As Figure 4 The variable attention weight of the time static covariate (time information) is mainly distributed in the interval [0.02, 0.06] and the interval [0.13, 0.21]. The number of samples with variable attention weight greater than 0.13 in the training data, validation data, and test data is 782, 270, and 313, respectively, accounting for about 38.472% of the total samples, indicating that the time information covariate is helpful for predicting dynamic congestion events in about 1 / 3 of the cases. In comparison, as The variable attention weight of the time static covariate Figure 4 (time information) is mainly distributed in the interval [0.02, 0.06] and the interval [0.13, 0.21]. The number of samples with variable attention weight greater than 0.13 in the training data, validation data, and test data is 782, 270, and 313, respectively, accounting for about 38.472% of the total samples, indicating that the time information covariate is helpful for predicting dynamic congestion events in about 1 / 3 of the cases. In comparison, as The variable attention weight of the time static covariate Figure 4 (time information) is mainly distributed in the interval [0.02, 0.06] and the interval [0.13, 0.21]. The number of samples with variable attention weight greater than 0.13 in the training data, validation data, and test data is 782, 270, and 313, respectively, accounting for about 38.472% of the total samples, indicating that the time information covariate is helpful for predicting dynamic congestion events in about 1 / 3 of the cases. In comparison, as The variable attention weight of the time static covariate Figure 4 (time information) is mainly distributed in the interval [0.02, 0.06] and the interval [0.13, 0.21]. The number of samples with variable attention weight greater than 0.13 in the training data, validation data, and test data is 782, 270, and 313, respectively, accounting for about 38.472% of the total samples, indicating that the time information covariate is helpful for predicting dynamic congestion events in about 1 / 3 of the cases. In comparison, as The variable attention weight of the time static covariate

[0082] As Figure 5 The variable attention weight of the time static covariate (line 18-3 transmission active power) is basically concentrated in the interval [0.11, 0.21], and the number of samples with variable attention weight greater than 0.11 accounts for about 87.176% of the total, indicating that the variable For the important role of predicting dynamic congestion events. This experimental phenomenon is basically consistent with the domain knowledge of the power system. As one of the tie lines between system area 1 and area 2, line 18-3 bears the responsibility of transmitting power to area 2, Figure 5 The results in the right column of the middle histogram show that the MA-LSTM T congestion risk perception network automatically identifies the importance of this variable through the attention mechanism. In contrast, the variable attention weight value distribution of the variable (node 14 photovoltaic active power output variable) is more dispersed, and the variable contribution robustness is not as good as that of the variable As shown in Figure 5 , the MA-LSTM T congestion risk perception network effectively identifies the characteristic differences between the variables and , and accurately gives the variable a higher and more robust attention. The above experimental results show that the MA-LSTM T congestion risk perception network can effectively automatically identify the contribution of variables to prediction, and its variable attention distribution pattern is consistent with general experience and domain knowledge.

[0083] The following further analyzes the variable time attention global data pattern, and the average time attention weight columnar heat map of the variable Figure 6 (Table 1 No. 8) and the variable (Table 1 No. 11) on the three data sets is drawn respectively, i.e. , , , , , , . In the figure, the lag time step represents the number of lag time points relative to the current time (i.e. 0), the time step interval is 15 minutes, and time 0 represents the time when the model performs prediction inference based on the current collected sample data.

[0084] As shown in Figure 6 , the MA-LSTM T congestion risk perception network gives each time step a time attention weight, reflecting the contribution factor of the corresponding input variable data information at that time step to the current prediction inference process. For the node 12 injection active power variable, since the experimental example system only contains a load object at node 12, the characteristics actually follow The active power of load is usually time-dependent and has a time trend due to the combined effects of human life patterns, business work patterns, and changes in the external environment, such as rising temperatures. In the process of developing a scheduling strategy, the load power is generally considered a boundary condition and is avoided as much as possible, so its power value generally follows its own characteristics. Therefore, focusing on the numerical characteristics of load power at multiple historical time steps plays an important role in accurately perceiving its future time trend. As shown in Figure 6 , the time attention of the model on the variable gradually increases as the time step approaches (relative to the current time 0), and the number of time steps with an average time attention weight of more than 0.09 is 4 (about 60 minutes), meaning that the MA-LSTM model will pay attention to data information at multiple historical time steps of the variable , and focus on the load trend in the last 60 minutes, helping the model to predict its future development trend and assist in perceiving the future dynamic congestion risk. This time attention distribution pattern is consistent with the above analysis of the characteristics of the load variable. In contrast, as shown in the right graph in Figure 6 , the time attention of the model on the variable is basically at a low level at the beginning, and suddenly increases greatly at the last time step, meaning that the model mainly focuses on data information at the last time step. The variable represents the active output of the generator at node 30, which is usually treated as a decision variable in the process of developing a scheduling strategy, so the generation data generally implies the influence of the scheduling strategy, and its time correlation is weaker than that of the load variable. Therefore, we should focus on its data information at the latest time to help the model understand the recent operating state of the system without paying too much attention to earlier generation information. Figure 6 The time attention weight distribution pattern in the right graph of is consistent with the above analysis and meets the needs of using generation time series data for dynamic congestion event perception tasks.

[0085] (2) Specific case analysis based on mixed attention In this section, based on specific dynamic congestion event cases, we use mixed attention weight analysis to understand the internal prediction and inference process of the MA-LSTMT congestion risk perception network when it faces specific data inputs. Figure 7 and Figure 8 respectively give the time series curves of some characteristic variables, time attention weight heat maps, and variable attention weight bar charts in two specific dynamic congestion event cases for specific data behavior analysis.

[0086] Figure 7The presented dynamic congestion event case 1 occurred at 20:00:00 on December 18th (16th time step). As shown in the figure, during this time period, the photovoltaic output was at a low level, while the loads in multiple nodes in region 2 were rapidly rising, such as variable 2: node 8 load, variable 4: node 4 load. In order to track the rising load power, the unit output in region 2 continued to increase, such as variable 5: generator output at node 32. However, due to the rapid rising speed of the load, the power demand of the external region increased rapidly, eventually causing the power grid dynamic congestion. Therefore, the rapid load rise is the main factor leading to this dynamic congestion event, especially the node 8 and node 4 loads represented by variable 2 and variable 4. According to the forward marking strategy (t = 30 min), the multi-dimensional time series samples collected by the observation time window in the left figure of belong to the “early stage” category, which are input into the trained MA-LSTMT network, and after the hidden state matrix calculation, they can enter the time attention calculation stage (see Figure 7 ). As shown in the second figure in Figure 3 , the second row and the fourth row, the MA-LSTMT dynamic congestion perception network effectively applies time attention to the steep rising trend of the two loads, and the time sequence growth pattern of the attention weight basically coincides with the rising trend of the load power. At the same time, as shown in the third figure in Figure 7 , the MA-LSTMT network also applies high variable attention to variable 2 and variable 4, indicating that the model effectively captures the key factors in this dynamic congestion event through the mixed attention mechanism. In addition, this dynamic congestion event occurred at 20:00 in the evening, which is a time period when dynamic congestion frequently occurs, as shown in the third figure in Figure 7 , the static covariate Figure 7 (time information) is also given high attention, indicating that the model effectively uses time information to assist in predicting dynamic congestion risk.

[0087] Figure 8 The presented dynamic congestion event case 2 occurred at 17:30:00 on August 9th (16th time step). As shown in Figure 8 ​As shown in the first plot, the photovoltaic power output at system node 9 (variable 7) exhibits a continuous rapid decreasing trend during this time period, and the photovoltaic power output at node 14 (variable 14) also exhibits a decreasing trend. Both the node 9 photovoltaic and the node 14 photovoltaic are located in the same area (i.e., area 2) of the power grid, and in order to make up for the power shortage, the generator active power output at node 32 (variable 5) in the area and the generator active power output at node 30 (variable 11) outside the area are both continuously increasing, and the continuous increase in transmission power eventually leads to power grid dynamic congestion. Therefore, the photovoltaic sudden drop is the main factor leading to this dynamic congestion event, especially the node 9 photovoltaic output represented by variable 7. Similarly, the corresponding dynamic congestion "early stage" sample (as shown in the first plot of the observation time window of Figure 8 , the trained MA-LSTM T congestion risk perception network can be input to calculate the time attention and variable attention weight corresponding to the sample. As shown in the second plot of Figure 8 , the MA-LSTM T congestion risk perception network effectively applies time attention to the decreasing trend of the node 9 photovoltaic output. At the same time, as shown in the third plot of Figure 8 , the MA-LSTM T network effectively captures the key factors in this dynamic congestion event and applies the highest variable attention to the node 9 photovoltaic output (variable 7). The mixed attention mechanism provides an effective way to visualize the decision-making process of the model, and each variable and the attention weight of each variable at each time step can be extracted to analyze the prediction reasoning logic of the model, so as to help the user check whether the attention mode of the model is correct and consistent with the domain cognition and general experience, and thus measure the trust degree of the model.

[0088] Model output confidence calibration test comparative analysis: This section will verify the effectiveness of the output confidence calibration module of the congestion risk perception network (see Figure 4 right) and carry out comparative analysis with other confidence calibration methods. All confidence calibration tests will be based on the trained dynamic congestion risk perception model, and the network parameters will not be changed. All confidence calibration processes will be based on the same verification data set (containing 722 samples), and the calibrated model will use the same test data set (containing 930 samples) to evaluate the dynamic congestion event probability prediction performance.

[0089] (1) Model confidence calibration test analysis This experiment uses the SLSQP algorithm of the SciPy computing library to perform confidence calibration parameter optimization. The experimental results of 10 repeated runs show that (calculation time: 0.0674 0.0208 seconds, number of optimization iterations: 8.700 2.268 times), based on the optimization solver used in this work and the experimental platform hardware condition, the iteration number of ETS confidence calibration parameter optimization process is basically within 11 times, and the average calculation time is in the order of milliseconds. Figure 9 The probability estimation reliability figures of MA-LSTMT blockage risk perception network on test dataset before and after ETS confidence calibration are shown, as well as the corresponding calibration error expectation ECE and maximum calibration error MCE evaluation indexes.

[0090] In this ETS confidence calibration experiment case, the optimized confidence calibration parameters are as follows: u = 2.111, (0.906, 0.094, 1.079×10-19). As Figure 9 shown, based on the proposed ETS confidence calibration post-processing method, the probability estimation performance of MA-LSTMT blockage risk perception network has been effectively corrected, and its ECE index is reduced from 3.52% before calibration to 1.70% (error reduction of about 52%). At the same time, its MCE and NLL indexes have also been significantly improved. The NLL of the model before calibration is 0.225, and the MCE is 21.70%; the NLL of the model after calibration is 0.167, and the MCE is only 8.41%, and the improvement ratios of the two indexes are about 26% and 61% respectively. In particular, for the MA-LSTMT blockage risk perception network without confidence calibration, as shown in the left figure of Figure 9 , in the interval where its observation accuracy acc and average confidence conf deviate the most, i.e. [0.7, 0.8], the sample classification accuracy acc is only 54.17%, while the average confidence conf reaches 75.87%, showing obvious overconfidence. In the decision-making scenario based on confidence score, considering that 75.87% is not a low confidence level, such overconfidence of the model may mislead the decision-making. For example, if the system dispatcher accepts the prediction inference result of the model with a confidence threshold of 70% to guide the decision-making process (i.e. if the confidence of a prediction result exceeds 70%, the prediction result of the model is believed and the decision is made based on it), only 54.17% of the decisions will be correct. Therefore, such overconfidence of the model is very unfavorable to the decision-making process and may have serious consequences. In addition, in the intervals [0.6, 0.7] and [0.8, 0.9], the MA-LSTMT model without confidence calibration also shows obvious overconfidence, and the deviations are 14.87% and 7.40% respectively.

[0091] In contrast, as shown in the right figure of Figure 9 , the calibrated model no longer has obvious overconfidence phenomenon Within the interval [0.6, 0.7], the sample classification accuracy (acc) was 61.77%, the average confidence level (conf) was 65.60%, and the bias was only 3.83%. Apart from this, there were no other significant differences. The maximum calibration error (MCE) of the calibrated model occurs in the interval [0.5, 0.6], where the sample classification accuracy (acc) is 63.64% and the average confidence score (conf) is 55.23%, indicating that the model is underconfident. In the interval [0.8, 0.9], the sample classification accuracy (acc) is 91.30% and the average confidence score (conf) is 86.04%. If a 70% confidence threshold is still used, then 91.03% of decisions in this interval will be correct. Therefore, in high-confidence intervals, overconfident models are generally more harmful to decision-making than underconfident models. Furthermore, in other intervals, the sample classification accuracy and average confidence score of the calibrated MA-LSTMT congestion risk perception network are basically matched, indicating that the model's confidence estimate can basically reflect the true correctness likelihood. The above results show that the ETS confidence calibration method proposed in this invention can effectively improve the reliability of confidence estimation of the congestion risk perception network, thereby improving the credibility of the model output, building user trust, and thus better assisting the decision-making process.

[0092] (2) Comparative analysis of confidence calibration methods This section compares the proposed confidence calibration method based on Integrated Temperature Scaling (ETS) with other state-of-the-art confidence post-processing methods to test the advantages of ETS calibration. The methods compared include: ① Original Temperature Scaling (OTS); ② Vector Scaling (VS); ③ Multi-class Isotonic Regression (MIR). To further test the proposed ETS confidence calibration method, the ETS confidence calibration module is applied to different deep learning networks and compared with other confidence calibration methods to comprehensively test the effectiveness and advantages of the proposed method. In addition to the MA-LSTMT network (including static covariates), the deep learning models used include: ① LSTM; ② GRU; ③ Convolutional Neural Network (CNN). The compared confidence calibration methods include OTS, VS, and MIR. Each method was executed 10 times to report the average calibration performance. The experimental results are shown in Table 3.

[0093] Table 3

[0094] As shown in Table 3, compared with the case without calibration, the probability estimation reliability of the models calibrated using different methods is improved to some extent, which reflects the effectiveness of the confidence calibration post-processing module. And when using different deep learning models, the proposed ETS confidence calibration method still achieves relatively better calibration results overall. The proposed ETS confidence calibration method ranks first in 10 out of 12 index comparisons, and ranks second in the other 2. When no calibration is performed, the probability estimation performance of each deep learning model is poor. Especially the MA-LSTMT network based on hybrid attention and tensor LSTM unit, although it performs best in accuracy (category prediction performance) (see Table 2), its average ECE, MCE and NLL indicators reach 3.59%, 27.26% and 27.13% respectively, showing the worst probability estimation reliability. This is mainly because the MA-LSTMT network has a more complex structure to help improve model interpretability and prediction accuracy, and a complex network structure is usually more prone to overfitting to the NLL loss on the training set, thereby reducing the probability estimation test performance of the model, which may affect the subsequent decision-making process. Based on the proposed ETS confidence post-processing method, the ECE, MCE and NLL indicators of the calibrated MA-LSTMT network are reduced to 1.31%, 8.42% and 15.48% respectively, with a reduction of 63.51%, 68.97% and 42.94% respectively compared with before calibration, and the ECE and MCE indicators rank first among all methods, showing relatively better probability estimation reliability. At this time, the MA-LSTMT congestion risk perception network not only has the best category prediction performance, but also has more reliable confidence output, which helps to make more accurate decisions.

[0095] Dynamic congestion risk quantification perception online simulation case analysis: This section will compare the online probability prediction performance of the MA-LSTMT congestion risk quantification perception network before and after confidence calibration based on a specific power grid dynamic congestion event case. The transmission dynamic congestion event case occurred on August 5, 15:45:00, and Figure 10 plots the time series trajectories of some characteristic variables during the dynamic congestion event. The dynamic congestion risk quantification perception online rolling simulation test results are shown in Figure 11 , including the category prediction results and probability estimation results output by the MA-LSTMT congestion risk perception network, where the dashed line represents the corresponding dynamic congestion event stage prediction probability.

[0096] As Figure 10As shown, the dynamic congestion event starts at time 15:45 (25th time step), and according to the forward-looking labeling strategy, the true labels of the multi-dimensional time series samples collected at the 23rd and 24th time steps are “early stage”. Based on the time series trajectories of the multi-dimensional feature variables, the MA-LSTM T congestion risk perception network makes an accurate judgment on the sample collected at the 23rd time step, accurately predicting it as an early stage sample, as shown in Figure 11 However, at the 24th time step, the trend of some variables (such as the load power at node 8 of variable 2, the active power output of the generator at node 32 of variable 5) shows a downward trend, and the downward trend of the photovoltaic output at node 9 (variable 7) slows down, which confuses the model and incorrectly predicts the sample as “normal stage”, forming a missed judgment, as shown in Figure 11 If the probability estimation results of the pre-calibration model are used to assist in the judgment of this situation, at the 23rd time step, the user will get the category prediction result: dynamic congestion early stage C1, and the corresponding confidence estimation result: 91.18%. Since 91.18% belongs to a high confidence level, the user is likely to choose to believe the prediction result and start to develop a dynamic congestion active control strategy starting from the 24th time step, ready to be issued at the 24th time step. However, by the 24th time step, the new output category prediction result of the model is normal stage, and the confidence reaches 89.19%. If the user accepts the prediction result at the same confidence level, he is likely to choose not to issue the control strategy, thinking that the system is really in a normal state, so he continues to observe the trend of its state, resulting in the missed prevention and control time window. At the same time, the dynamic congestion early stage and the dynamic congestion probability output by the model are only 9.31% and 1.50% respectively, which is difficult to arouse the vigilance of the dispatching and operation personnel. As shown in Figure 12 As shown, the pre-calibration model usually outputs over-confident probability estimation results, resulting in a polarized probability curve, i.e. close to 1 or close to 0, as shown in Figure 12 which will further affect the online decision-making process, easily causing misjudgment and making it difficult to play an actual auxiliary role.

[0097] As shown in Figure 10 After ETS confidence calibration, the MA-LSTMT congestion risk perception network outputs a more smooth probability estimation curve, and its category quantization probability prediction result can basically reflect the development trend of the dynamic congestion event (see Figure 11), the dynamic congestion risk is accurately quantified on the timing probability estimation curve. If the probability estimation results based on the calibrated model are used to assist in judging this dynamic congestion event case, at the 23rd time step, the user will get the category prediction result: early dynamic congestion, and the corresponding confidence estimation result: 71.81%. Considering that 71.81% is not a low confidence level, the dispatch operation personnel will remain vigilant and begin to develop an active control strategy for dynamic congestion. At the 24th time step, the new output category prediction result of the model is the normal stage, with a confidence of 43.93%, but at the same time, the early dynamic congestion and dynamic congestion probabilities are 40.24% and 15.83%, respectively. Although the category prediction result is still normal (the calibration process does not change the category prediction result), its probability is only 43.93%, and the abnormal probability is 56.07%. At the same time, based on the timing probability estimation curve, it can be seen that the early dynamic congestion probability has been increasing since the 17th time step, and the dynamic congestion probability has been increasing since the 22nd time step. Therefore, based on the above information, the dispatch operation personnel have sufficient reason to question the category prediction result given by the model, enhance the confidence of implementing active control, and choose to issue the developed control strategy, thereby achieving accurate defense against dynamic congestion risk. As can be seen, based on the ETS confidence calibration method proposed, the MA-LSTM T congestion risk perception network can output more reliable probability estimation information, and based on the calibrated probability estimation curve, the relevant dispatch personnel can obtain more comprehensive auxiliary reference information, thereby helping them to identify possible errors of the model and correct their own decisions. Even if the category prediction result output by the model is biased, based on the comprehensive multi-category probability curve trend, the dispatch operation personnel still have the opportunity to make accurate decisions, thereby ensuring system safety. In contrast, the pre-calibration model has a significant overconfidence problem, and the probability curve output by the model is basically polarized, as shown in Figure 10 , so it is difficult to obtain more effective reference information from the curve trend. In summary, the ETS confidence calibration method can effectively improve the reliability of the model's probability output, make it reflect the true correctness likelihood of the prediction as much as possible, and provide more comprehensive auxiliary reference information for the dispatch operation personnel, which helps them to make more accurate decisions.

[0098] Embodiment Two Figure 10 A schematic diagram of a high-confidence probability early warning device for dynamic congestion of a power system according to Embodiment Two of the present application is shown in ​ The device in includes: The model construction unit 22 is configured to construct a congestion risk quantization perception network based on a hybrid attention mechanism, a tensorized LSTM network unit, and a static covariant knowledge information fusion network. The training unit 23 is configured to train the congestion risk quantization perception network based on the supervision data set to obtain a trained congestion risk quantization perception network. The prediction unit 24 is configured to input the to-be-tested multi-dimensional time series data into the trained congestion risk quantization perception network to predict the congestion state of the power system.

[0099] The power system dynamic congestion high-confidence probability early warning device provided in the embodiments of the present application can execute the power system dynamic congestion high-confidence probability early warning method provided in any of the embodiments of the present application, and has the corresponding function modules and beneficial effects of the execution method.

[0100] It should be understood that the various forms of flow shown above can be used to reorder, add or delete steps. For example, each step described in the present application can be executed in parallel, sequentially or in a different order, as long as the desired results of the technical solutions of the present application can be achieved, which is not limited herein.

[0101] The above specific embodiments do not constitute a limitation on the protection scope of the present application. Those skilled in the art should understand that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent replacements and improvements made within the spirit and principles of the present application shall be included in the protection scope of the present application.

Claims

1. A power system dynamic congestion high-confidence probability early warning method, characterized in that, include: Obtain multivariate, multidimensional time series data and construct a supervised dataset; A blocking risk quantization perception network is constructed based on a hybrid attention mechanism, tensor LSTM network units, and a static covariate knowledge information fusion network. Based on the supervised dataset, a blocking risk quantification and perception network is trained to obtain the trained blocking risk quantification and perception network. The multidimensional time series data to be tested is input into a trained congestion risk quantification and perception network to predict the congestion status of the power system.

2. The method of claim 1, wherein the high confidence probability of power system dynamic congestion early warning is characterized by, The process of acquiring multivariate, multidimensional time series data and constructing a supervised dataset includes: The multidimensional time series data is extracted by sliding through a time window, and cross-sectional marking is performed after identifying the system dynamic blocking state of the multidimensional time series data to construct a supervised dataset.

3. The method of claim 1, wherein the high confidence probability of power system dynamic congestion early warning is characterized by, The process by which the blocking risk quantification perception network processes the input multidimensional time series samples and static covariate data includes: Multidimensional time series samples are input into tensorized LSTM network units to obtain the corresponding hidden state matrix sequence; Temporal attention is applied to the historical hidden state sequence, that is, the hidden state sequence is weighted by temporal attention weights to obtain the historical context vector information of each input feature variable; The historical context vector information is concatenated with the current hidden state sequence to obtain the concatenated dynamic historical information; Input static covariates to obtain static enrichment context vectors; The dynamic historical information is concatenated with the static enriched context vector to obtain the context tensor; Calculate the network logarithm and variable attention weights based on the context tensor; The final network logarithm is calculated based on the network logarithm and the attention weights of the variables. The final network logarithmic output shows the predicted blocking state.

4. The method of claim 3, wherein the high confidence probability of power system dynamic congestion early warning is characterized by, The step of inputting multidimensional time series samples into tensor-quantized LSTM network units to obtain the corresponding hidden state matrix sequence includes: The tensorized hidden state matrix at time step l can be represented as: , , The overall size of the hidden layer is defined as , is the hidden size corresponding to each variable; element is the hidden state vector corresponding to the n-th input variable; The update structure of a tensorized LSTM unit includes an input gate , a forget gate , an output gate , and a memory cell state ; in which, , , , represents the update element of the hidden state with respect to the input variable n; represents the input-to-hidden state transition tensor, , , ; represents the transition tensor between hidden states at different time steps, , , ; represents the bias, ; similarly, represents the input-to-hidden state transition tensor corresponding to the gate , , , , , , ; represents the transition tensor between hidden states at different time steps, , , ; represents the bias corresponding to the gate , , , ; and capture the knowledge information from the previous hidden state and the new input data , respectively; the operation represents the tensor point product of the two tensors along the variable dimension , and the calculation process is as follows: ; tanh(.) denotes the tanh activation function, denotes the activation function, operation denotes element-wise multiplication.

5. The method of claim 4, wherein, The application of temporal attention to the historical hidden state sequence involves weighting the hidden state sequence using temporal attention weights to obtain the historical context vector information for each input feature variable. include: wherein, for the multi-dimensional time series sample After inputting the tensorized LSTM network unit, , a hidden state matrix sequence is obtained, wherein, for the hidden state sequence of the nth input feature variable, i.e. , a time attention is applied to the hidden state sequence of the input feature variables respectively, i.e. , ; The time attention weights are: wherein denotes the time attention weight at the l-th time step corresponding to the n-th input feature variable, ; denotes a feed-forward neural network activated by a tanh(.) function corresponding to variable n, and T denotes time attention. Using temporal attention weights on the hidden state sequence weighted, a historical context vector information corresponding to the feature variable n is obtained, denoted as , ; The process involves concatenating the historical context vector information with the current hidden state sequence to obtain concatenated dynamic historical information; inputting static covariates to obtain a static enriched context vector; and concatenating the dynamic historical information with the static enriched context vector to obtain a context tensor, including: wherein represents a contextual tensor; represents a vertical concatenation operation, represents a history contextual vector information of a concatenation feature variable n and current hidden state information, represents a static rich contextual vector.

6. The method for high-reliability early warning of dynamic congestion in power systems according to claim 5, characterized in that, The calculation of network logarithms and variable attention weights based on context tensors includes: The attention weights of the variables are: In the formula, This represents the variable attention weight for the i-th variable. It has the following numerical characteristics: , , ; Representation matrix The i-th row vector, , ;when hour, ;when hour, ; It is a feedforward neural network activated by the tanh(.) function and shared by all variables; the subscript V represents variable attention. For the i-th physical variable, the calculated network logarithm is: , It is a fully connected network without an activation function corresponding to variable i, responsible for the context matrix. The data representation in the text is mapped to the label space, and the subscript FC indicates a fully connected layer; | | indicates the total number of categories in the target classification problem; The final network logarithm is calculated based on the network logarithm and the attention weights of the variables, including: in, This represents the final network logarithm vector obtained through weighted calculation. ; Based on the final network logarithmic output of the blocking state prediction results, including: The final network logarithm of the synthesis Activation will be performed through a softmax layer, thus forming a dynamic blocking event-aware model: In the formula, This represents the softmax activation function. Represents the logarithmic vector of the synthesized network, containing | | elements, This represents the set of category labels. In dynamic blocking event awareness tasks, h is the category index. , This represents the probability estimate of class h from the output of the softmax layer. This represents the number of network partitions corresponding to category h. ; Based on probability estimation vector Then you can obtain the input sample The corresponding category prediction results, i.e. , And the estimated confidence score corresponding to the prediction results of the corresponding categories. .

7. The method for high-reliability early warning of dynamic congestion in power systems according to claim 6, characterized in that, The objective function used to train the blocking risk quantification awareness network is to minimize the cross-entropy loss: In the formula, m is the sample index. Input for sample The corresponding final output network logarithm, For the true one-hot encoded label corresponding to the m-th sample, The true class of the m-th sample is h if and only if the true class of the m-th sample is h. ;otherwise .

8. The method for high-reliability early warning of dynamic congestion in power systems according to claim 1, characterized in that, The blocking risk quantification and perception network further includes a confidence post-processing unit, comprising: The confidence post-processing unit reconstructs the mapping relationship between the logarithm of the heterogeneous information fusion network and the posterior probability of the dynamic blocking event by minimizing the negative log-likelihood, thereby calibrating the confidence output of the blocking risk quantification perception network without affecting the category prediction results.

9. The method for high-reliability early warning of dynamic congestion in power systems according to claim 8, characterized in that, The confidence post-processing unit specifically includes: For input multidimensional time series data samples and static covariate information The number of logarithmic segments of the synthetic network calculated by the model is , , Represents the number of network splits corresponding to category h; based on The corresponding softmax score vector can be obtained. This refers to the original probability vector estimated by the network, from which the predicted labels for the dynamic blocking event stages can be obtained. and the corresponding confidence score. : Based on the true labels of the samples and network logarithmic vector The goal of confidence calibration is to produce corrected confidence scores. At the same time, maintain the category prediction results The confidence post-processing unit remains unchanged; it is a unit containing optimizable parameters. and The three-component integrated post-processing module; for samples and and network logarithm vector The confidence post-processing unit calibration mapping has the following form: In the formula, This represents the calibrated post-confidence corresponding to category h after processing by the confidence post-processing unit, u and The introduced calibration parameters to be optimized This indicates the total number of tag types for the event phase. It is the softmax function; Corresponding to the original temperature-based confidence calibration method, this method introduces a parameter u, which is equivalent to maximizing the entropy of the output probability distribution under the logarithm constraint of the network, thereby alleviating the overconfidence problem of the original deep learning network; assuming given The network logarithm vector corresponding to each input sample and the corresponding real category labels The temperature scaling model is the only solution to the lower maximum entropy problem. : Solving for confidence post-processing unit calibration parameters and The negative log-likelihood loss optimization model is shown below: In the formula, , The number of multidimensional time series data samples used for confidence calibration, i.e., the number of validation dataset samples. To block the model parameters of the risk quantization perception network for the already trained model; In the formula, The class prediction results output by the calibrated model, compared with Maintain consistency Output the corresponding calibration confidence level.

10. A power system dynamic congestion high-reliability probability early warning device, characterized in that, include: Dataset building units are used to acquire multivariate, multidimensional time series data and build supervised datasets. The model building unit is used to construct a blocking risk quantification perception network based on a hybrid attention mechanism, a tensor LSTM network unit, and a static covariate knowledge information fusion network. The training unit is used to train the blocking risk quantification perception network based on the supervised dataset to obtain the trained blocking risk quantification perception network. The prediction unit is used to input the multi-dimensional time series data to be tested into the trained congestion risk quantification perception network in order to predict the congestion status of the power system.