Virtual power plant fault early warning method based on hierarchical interactive causality graph Transform
By constructing a three-level hierarchical structure and interactive causal graph of the virtual power plant, combined with multi-resolution causal distillation and conditional variational graph diffusion models, the causal correlation modeling problem of the multi-level complex system of the virtual power plant is solved, and high-accuracy fault warning is achieved.
Patent Information
- Application Number
- CN202510721816.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-30
- Publication Date
- 2025-09-09
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing technologies make it difficult to effectively establish a multi-level causal relationship modeling mechanism and are unable to accurately identify the complex interactive relationships between multi-level complex systems in virtual power plants, resulting in low accuracy in virtual power plant fault warnings.
A three-level hierarchical structure consisting of the micro-device layer, the meso-system layer, and the macro-market layer is constructed, and a three-level hierarchical interactive causal graph is established. Cross-level information transmission and fault path identification are achieved through methods such as multi-resolution causal distillation mechanism, multi-level comparative causal attention mechanism, and conditional variational graph diffusion model.
It significantly improves the accuracy of early warning for virtual power plants in complex fault scenarios, realizes full-link causal tracing from underlying equipment anomalies to market-level chain reactions, and improves the accuracy and reliability of early warning.
Smart Images

Figure CN120611270A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of virtual power plant fault warning technology, and in particular to a virtual power plant fault warning method based on a hierarchical interactive causal graph Transformer. Background Art
[0002] In the global energy transition, distributed energy, renewable energy, and energy storage systems are becoming increasingly integrated, and virtual power plants (VPPs), an innovative energy management model, are attracting significant attention. Leveraging advanced information and communications technologies, VPPs integrate diverse equipment, effectively enhancing system flexibility and renewable energy absorption capacity. However, VPPs are complex internal structures, diverse equipment, and volatile operating environments, leading to significant operational uncertainties and increased difficulty in fault early warning.
[0003] Traditional power system fault warning methods mainly fall into two categories: those based on physical models and statistical methods. Physical model-based methods require precise modeling of the system's physical characteristics. While they have a solid theoretical foundation and clear physical meaning, they are difficult to establish accurate full-system models for complex systems such as virtual power plants. Furthermore, they are computationally complex and have poor real-time performance. Statistical methods utilize historical operating data and employ a variety of statistical learning methods to identify abnormal patterns. While simple to implement, they struggle to capture nonlinear relationships and complex interactions, and their prediction capabilities are limited in areas with sparse data. With the development of machine learning, early warning methods based on deep learning have emerged. Research institutions in various countries have proposed different methods that have demonstrated advantages in improving early warning accuracy, identifying complex correlated faults, and handling long-sequence dependencies.
[0004] However, existing technologies still have shortcomings in virtual power plant applications. Virtual power plants are complex, multi-layered systems. Existing methods struggle to effectively model the complex interactions between these layers, and their ability to identify cross-layer correlated faults is limited. Warning accuracy and lead time are insufficient, and sensitivity to weak fault signals is low. Deep learning models are computationally complex, making real-time responses difficult. They lack interpretability, making it difficult to identify cascading fault paths. Furthermore, large models make efficient deployment on edge devices difficult, limiting their practical applications.
[0005] Chinese patent publication number CN117951577A discloses a method for energy status perception of a virtual power plant. It merely implements data feature extraction and time series learning through a hybrid architecture of stacked convolutional neural networks and Transformers. It is unable to effectively establish a multi-level causal association modeling mechanism and can only capture surface data associations but cannot parse the complex interactive relationships between multi-level complex systems in a virtual power plant. Consequently, it is difficult to accurately identify potential cascading failure paths, resulting in low early warning accuracy when faced with complex failure situations. Summary of the Invention
[0006] To this end, the present invention provides a virtual power plant fault warning method based on a hierarchical interactive causal graph Transformer, which is used to overcome the problem that the existing technology cannot effectively establish a multi-level causal association modeling mechanism, can only capture surface data associations but cannot analyze the complex interactive relationships between multi-level complex systems in the virtual power plant, and thus it is difficult to accurately identify potential cascading failure paths, resulting in low accuracy of warning when facing complex fault situations.
[0007] To achieve the above objectives, the present invention provides a virtual power plant fault warning method based on a hierarchical interactive causal graph Transformer, comprising the following steps: S1. Based on the virtual power plant, a three-level hierarchical structure consisting of micro-equipment layer, meso-system layer and macro-market layer is constructed, and a three-level interactive causal diagram is established based on the three-level hierarchical structure; S2. Construct a multi-resolution causal distillation mechanism based on the three-level hierarchical interactive causal graph, and use the multi-resolution causal distillation mechanism to model the global causal relationship of the Transformer encoder-decoder architecture to obtain the Transformer model; S3. Introducing the multi-level contrastive causal attention mechanism into the Transformer model for automatic quantitative analysis to obtain complex causal relationships between layers; S4. Eliminate false correlations in the complex causal relationships between layers based on the high-order Markov random field framework to obtain the actual complex causal relationships between layers; S5. Construct an inter-layer information transmission mechanism model based on the conditional variational graph diffusion model and the actual complex causal relationship between layers; S6. Optimizing the inter-layer information transfer mechanism model based on the inter-layer information transfer function and the layered sparsification mechanism to obtain an optimized inter-layer information transfer mechanism model; S7. Based on the optimized inter-layer information transmission mechanism model, the potential cascading failure paths of the virtual power plant are identified to obtain comprehensive fault warning information.
[0008] In this solution, a three-level hierarchical structure is defined from the bottom up through the virtual power plant operation scenario, which can comprehensively and accurately cover the full-dimensional elements of the virtual power plant from equipment to the market, avoiding analysis blind spots, and establishing a three-level hierarchical interaction causal graph based on the three-level hierarchical structure to clarify the dynamic correlation between the elements of each layer. By implementing multi-resolution causal distillation on the causal interaction graph, the causal relationships of different granularities are compressed into the Transformer model, forming a basic framework with the ability to capture temporal correlations. After further introducing the multi-level comparative causal attention mechanism, the attention weights are automatically calibrated to the key causal paths, and the Transformer model can automatically quantify the causal effects between layers. In order to eliminate the mixed false correlations in the data, a high-order Markov random field framework is used to perform conditional probability constraints on the complex causal relationships between layers, and only the actual complex causal relationships between layers that conform to the causal topological structure are retained. Based on the actual complex causal relationships between layers, the conditional variational graph diffusion model is used to construct an inter-layer information transmission mechanism. The cross-layer propagation modeling of causal effects is realized through the inter-layer information transmission mechanism, and the optimized inter-layer information transmission mechanism model is obtained. The potential cascade fault propagation path during the operation of the virtual power plant is dynamically tracked through the optimized inter-layer information transmission mechanism model. Finally, comprehensive early warning information including fault location, impact range and evolution trend is generated through multi-indicator fusion.
[0009] Compared with the existing technology, the beneficial effect of this application is that by constructing a three-level hierarchical structure including the micro-device layer, the meso-system layer and the macro-market layer and establishing an interactive causal graph, the structured modeling of the cross-level complex system of the virtual power plant is realized, breaking through the traditional single-dimensional analysis of the separation of hierarchical associations. By modeling the global causal association of the Transformer encoder-decoder architecture based on the multi-resolution causal distillation mechanism, the temporal representation ability of cross-layer information fusion is significantly improved. By introducing the multi-level comparative causal attention mechanism, the automatic quantification of complex causal relationships between layers is realized, and the cascade of equipment failures to the system is effectively captured. The dynamic path-dependent characteristics of diffusion, after eliminating false correlations with the help of a high-order Markov random field framework, the constructed inter-layer information transmission mechanism model can accurately characterize the actual causal effect propagation chain, avoiding the misjudgment risk of traditional statistical correlation analysis. The information transmission mechanism optimized based on the conditional variational graph diffusion model dynamically simulates the fault propagation trajectory through the probabilistic graph structure, and ultimately improves the accuracy of identifying potential cascading fault paths, solving the problem of failure in modeling multi-level complex interactive relationships in existing technologies, and realizing full-link causal tracing from bottom-level equipment anomalies to market-level chain reactions, significantly improving the early warning accuracy of virtual power plants in complex fault scenarios.
[0010] Furthermore, the S1 comprises the following steps: S11. Based on the three-level hierarchical structure, clarify the upload path for transient device data from the micro-device layer to the meso-system layer, the logic for issuing control instructions from the meso-system layer to the micro-device layer, and the transmission rules for market signals from the macro-market layer to the meso-system layer. Establish the interactive relationship between the layers based on the upload path and transmission rules. S12. Based on the interaction relationship between the levels, a three-level hierarchical interactive causal diagram is established, in which nodes represent key variables at each level and edges represent the causal direction and causal strength between variables.
[0011] In this scheme, by clarifying the equipment transient data upload path, control command issuance logic and market signal transmission rules, a clear and standardized inter-level interaction relationship is established, ensuring the orderliness and efficiency of information interaction between various levels of the virtual power plant; by establishing a three-level hierarchical interaction causal diagram based on the inter-level interaction relationship, the causal relationship and strength between the key variables of each level can be intuitively presented, and thus an in-depth understanding of the complex causal mechanism within the three-level hierarchical structure can be achieved.
[0012] Furthermore, the step S2 includes the following steps: S21. Analyze the three-level hierarchical interaction causal graph to obtain the adjacency matrix mapping relationship; S22. Based on the adjacency matrix mapping relationship, a multi-resolution causal distillation mechanism is constructed. The low-resolution causal structure is topologically aligned and information is compressed into the high-resolution semantic space according to the multi-resolution causal distillation mechanism to obtain the distilled hierarchical embedding vector. The mathematical expression of the multi-resolution causal distillation mechanism is: ; Indicates the Tier Node pair Tier The attention weight of each node, represents the query parameter matrix, , A parameter matrix representing the key values, , Indicates the Tier The hidden state of each node, The dimension is , Indicates the Tier The hidden state of each node, The dimension is , represents the scaling factor of the attention mechanism, represents a sparse mask matrix based on causal strength, The value of is 0 or 1. Representation node and nodes The structured diffusion distance metric between represents the adaptive time-varying temperature parameter function, represents the structural similarity function, represents the verification function based on causal strength, where Representation node For Node The causal strength of represents the timing adaptive modulation function; S23. Inject the hierarchical embedding vectors into the Transformer encoder-decoder architecture to perform global causal relationship modeling and obtain the Transformer model.
[0013] In this scheme, the adjacency matrix mapping relationship is obtained by parsing the three-level hierarchical interactive causal graph. A multi-resolution causal distillation mechanism is constructed based on the adjacency matrix mapping relationship, and the low-resolution causal structure is topologically aligned and information is compressed into the high-resolution semantic space to obtain the distilled hierarchical embedding vector, which effectively retains the key causal information and improves the data representation efficiency. The hierarchical embedding vector is injected into the encoder-decoder architecture of the Transformer for global causal association modeling, resulting in a Transformer model that can comprehensively capture the complex causal relationships between nodes at all levels, enhancing the accuracy and reliability of the analysis.
[0014] Furthermore, the step S3 includes the following steps: S31. Construct a contrastive learning objective function based on the feature distribution of the original self-attention mechanism of the Transformer model, and obtain a latent space representation containing causal constraints based on the contrastive learning objective function. The mathematical expression of the contrastive learning objective function is: ;in, represents the value of the comparative learning objective function, is a set of positive sample pairs, Contains pairs of nodes that have causal relationships, Represents the set of positive sample pairs All node pairs in expectations, is a node The negative sample set, Contains nodes There is no causal relationship between nodes. Representation node The latent space features and nodes The latent space features The similarity measure between Representation node The latent space features and nodes The latent space features The similarity measure between represents the temperature parameter; S32. Design causal attention units through latent space representation, and construct a causal relationship matrix based on the causal attention units combined with the non-parametric causal discovery algorithm, where the causal attention units The mathematical expression is: ,in, Represents the causal relationship matrix, element Indicates the query location To key position The causal strength of represents element-wise multiplication, 、 、 represent the query matrix, key matrix and value matrix respectively, represents the embedding dimension, represents the similarity score between the query matrix and the key matrix; Cause and Effect Matrix The mathematical expression is: ,in, represents a causal discovery algorithm based on a nonlinear additive noise model, represents the time series causal discovery algorithm; S33. Based on the causal relationship matrix, we introduce sparsification constraints and dynamic threshold adjustment to quantify discrete causal relationships into continuous weight values, generating a dynamic causal relationship matrix that can be embedded in attention calculations. S34. Construct a top-down causal chain tracing path and a bottom-up feature fusion path based on the dynamic causal relationship matrix, and perform hierarchical aggregation on the dynamic causal relationship matrix based on the causal chain tracing path and the feature fusion path to obtain a hierarchical aggregation result, wherein the hierarchical aggregation The mathematical expression is: ,in, Indicates the The causal relationship matrix of the layer, is the layer weight, is the total number of layers; S35. Determine the comparative causal loss function based on the hierarchical aggregation results, and automatically quantify the complex causal relationship between layers based on the comparative causal loss function. The mathematical expression is: ,in, Represents the hierarchical aggregation result, represents the acyclic constraint loss, represents the sparsity constraint loss, represents the acyclic constraint loss The trade-off parameter, represents the sparsity constraint loss trade-off parameters.
[0015] In this scheme, by constructing a contrastive learning objective function based on the Transformer model, a latent space representation containing causal constraints is obtained, which effectively enhances the model's ability to capture causal relationships. By designing a causal attention unit and combining it with a non-parametric causal discovery algorithm to construct a causal relationship matrix, the causal strength between nodes can be more accurately characterized. By introducing sparsity constraints and dynamic threshold adjustment to generate a dynamic causal relationship matrix, continuous quantification of discrete causal relationships is achieved. By constructing a causal chain tracing path and a feature fusion path for hierarchical aggregation, a more structured hierarchical aggregation result is obtained. Finally, based on the hierarchical aggregation result, the contrastive causal loss function is determined, and the complex causal relationships between layers are automatically quantified and analyzed, thereby improving the accuracy of causal relationship analysis.
[0016] Furthermore, the micro-device layer uses the FlashAttention-4 structure to process device transient data, the meso-system layer introduces Performer-CDF to analyze system status data, and the macro-market layer combines Rotary PositionEmbedding and RWKV mechanisms to process market data.
[0017] In this solution, FlashAttention-4 refers to an algorithm that improves the efficiency of self-attention calculation in the Transformer model, Performer-CDF refers to an algorithm that analyzes the distribution characteristics of system state data through CDF to enhance the stability analysis and prediction of complex systems, Rotary Position Embedding refers to a position encoding method that embeds position information into a vector in a rotated manner, retains relative position information while adapting to long sequences, and the RWKV mechanism refers to a mechanism used to effectively capture long-term dependencies and local features in data while maintaining good computational efficiency. By using the FlashAttention-4 structure to process device transient data at the micro-device layer, it can efficiently process large-scale device data, reduce memory usage, increase computing speed, and quickly capture device transient features; by introducing Performer-CDF to analyze system state data at the meso-system layer, it can reduce computational complexity while ensuring accuracy and more accurately analyze system state changes; by combining Rotary Position Embedding and RWKV mechanisms to process market data at the macro-market layer, RotaryPosition Embedding can effectively encode temporal position information, while the RWKV mechanism improves the ability to process long sequences. The combination of the two can more accurately grasp market dynamics and trends.
[0018] Furthermore, the original self-attention mechanism includes: The attention of the micro-device layer is calculated, where the mathematical expression of the attention of the micro-device layer is: , represents the attention matrix at the micro-device level, 、 、 represent the query matrix, key matrix, and value matrix of the micro-device layer respectively; The attention of the meso-system layer is calculated, where the mathematical expression of the attention of the meso-system layer is: , represents the attention matrix of the meso-system layer, 、 、 They represent the query matrix, key matrix and value matrix of the meso-system layer respectively; The attention of the macro market layer is calculated, where the mathematical expression of the attention of the macro market layer is: , represents the attention matrix of the macro market layer, Represents market data at the macro market level, represents the rotation position encoding function, Indicates the RWKV mechanism.
[0019] In this scheme, by calculating the attention of the micro-device layer, the correlation information between the elements within the device layer can be accurately captured. The attention matrix constructed based on the query matrix, key matrix and value matrix of the micro-device layer helps to deeply understand the operating status of the device; by calculating the attention of the meso-system layer, the interaction relationship between different components of the system layer can be effectively analyzed. The attention matrix of the meso-system layer provides a key basis for system status evaluation; by calculating the attention of the macro-market layer, combining the rotational position encoding function and the RWKV mechanism to process market data, the temporal characteristics and long-term dependencies of market data can be better grasped.
[0020] Further, the S4 includes the following steps: S41. Construct a high-order Markov random field framework. The mathematical expression of the high-order Markov random field framework is: ,in, Indicates system status The joint probability distribution of represents the normalization constant, is the maximum clique set, Indicates that the definition is in the group The higher-order potential function on delegation A subset of random variables in ; S42. Define high-order potential function based on the high-order Markov random field framework. The mathematical expression is: ,in, represents the first-order characteristic function; represents the third-order interaction characteristic function, represents the first-order characteristic function The model parameters, represents the third-order interaction characteristic function The model parameters of , k represents the number of variables involved in the potential function; S43. determining a parameterized energy model based on the high-order potential function, and sampling the parameterized energy model according to the Hamiltonian Monte Carlo method to obtain a Markov chain sample sequence; S44. Analyze the Markov chain sample sequence according to the kernel density conditional probability engine to obtain the structured causal effect matrix. The mathematical expression is: ,in, Indicates except All variables except Indicates that it contains variables All regiments; S45. Optimize the pseudo-likelihood parameters and perform Bootstrap test on the structured causal effect matrix to obtain the complex causal relationship between the actual layers and the objective function of the pseudo-likelihood parameters. The mathematical expression is: ,in, Indicates the training samples, Indicates the In the training samples, except All other variables except represents the variable dimension, represents the set of model parameters, Represents the total number of training samples.
[0021] In this scheme, by constructing a high-order Markov random field framework, the complex dependencies between system states can be more comprehensively described. By defining a high-order potential function, the high-order interaction characteristics between variables can be accurately captured, thereby enhancing the accuracy of causal inference. By determining a parameterized energy model based on the high-order potential function and using the Hamiltonian Monte Carlo method to sample the Markov chain sample sequence, the diversity and representativeness of the samples are effectively improved. The sample sequence is analyzed through the kernel density conditional probability engine to obtain a structured causal effect matrix, which intuitively presents the causal relationship between variables. By performing pseudo-likelihood parameter optimization and Bootstrap test on the structured causal effect matrix, the complex causal relationship between actual layers is finally obtained, which significantly improves the reliability and accuracy of causal inference.
[0022] Furthermore, the step S5 includes the following steps: S51. Based on the complex causal relationship between actual layers, a conditional variational graph diffusion model is constructed. The mathematical expression of the conditional variational graph diffusion model is: ,in, Represents the time step The potential graph of Indicates conditional information. represents the potential state at a given final time step and condition information Under the condition of The conditional probability distribution of represents the number of diffusion steps, Represents the time step -1 potential graph, represents the latent state at the initial time step, represents the latent state at the final time step; S52. Based on the conditional variational graph diffusion model, the forward diffusion process is defined. The mathematical expression of the forward diffusion process is: ,in, represents the noise scheduling parameter, represents the identity matrix, Represents the time step The potential graph by is the mean, is the multivariate Gaussian distribution with covariance matrix, Indicates the forward diffusion process from time step -1 potential graph To time step The potential graph The conditional probability distribution of ; S53. Perform reverse denoising on the complex causal relationship between actual layers according to the forward diffusion process to obtain a reverse denoising result. The mathematical expression of the reverse denoising process is: , represents the mean parameterized by the neural network, represents the covariance function, Represents the time step -1 potential graph by is the mean, is the multivariate Gaussian distribution with covariance function, Indicates that at a given time step The potential graph and condition information Under the condition of -1 potential graph The conditional probability distribution of ; S54. Construct a diffusion model using a conditional embedding mechanism based on the inverse denoising result, and construct a fault feature extractor using the diffusion model using the conditional embedding mechanism. The mathematical expression of the diffusion model using the conditional embedding mechanism is: , represents the basic mean prediction network, represents the time modulation function, represents the conditional encoding function, represents element-wise multiplication, In a given latent graph , time step and condition information Under the condition of , the mean generated by the conditional embedding mechanism; Fault Feature Extractor The mathematical expression is: , represents the encoder network, Represents a feature filter; S55. Construct an inter-layer information transfer mechanism model based on the fault feature extractor and the variational learning objective function, wherein the variational learning objective function The mathematical expression is: , represents standard normal noise, represents the noise prediction network, express and The joint distribution of .
[0023] In this scheme, a conditional variational graph diffusion model is constructed according to the actual complex causal relationship between layers. By defining the forward diffusion process, the noise addition rules of the potential graph over time are clarified, which helps to capture the dynamic changes of the causal relationship; the reverse denoising result is obtained through the forward diffusion process, which can effectively restore the clear causal relationship structure; by constructing a conditional embedding mechanism to enable the diffusion model and constructing a fault feature extractor based on it, the key features related to the fault can be accurately extracted; finally, the inter-layer information transmission mechanism model is constructed through the fault feature extractor and the variational learning objective function, which realizes the effective transmission and fusion of inter-layer causal information and improves the accuracy of inter-layer relationship modeling of the virtual power plant.
[0024] Further, the S6 includes the following steps: S61. Construct an inter-layer information transfer function and a layered sparsification mechanism respectively. The mathematical expression of the inter-layer information transfer function is: , Indicates that from Tier Node to Tier The amount of information transmitted by each node, represents the parameterized transfer function, represents the symmetric mutual information measure; The mathematical expression of the layered sparsification mechanism is: , Representation node To Node The sparse connection strength, represents the sigmoid function, represents the sparsification control parameter, represents the thermodynamic mutual information, represents the dynamic threshold function; S62. Design a low-rank decomposition kernel function attention unit based on the inter-layer information transfer function and the layered sparsification mechanism. The expression of the kernel function attention unit is: , represents the kernel function attention matrix, 、 、 represent the query matrix, key matrix and value matrix respectively, Represents the query matrix The result after kernel function feature mapping, Represents the bond matrix Transpose the kernel function feature map; S63. Parallelize the kernel function attention unit to obtain a parallelized result, where the mathematical expression of the parallel computing strategy is: , represents the parallelized matrix, Indicates the The partial attention matrix calculated by the processing unit, Indicates the number of processing units, Represents matrix concatenation operation; S64. Perform model parameter decomposition on the parallelization result according to Tucker to obtain a model parameter decomposition result, and perform mixed precision quantization on the model parameter decomposition result to obtain a mixed precision quantization result. The mathematical expression of Tucker is: , represents the quantized weight, represents the original weight, represents the scaling factor, Represents a quantized operation; S65. Optimize the inter-layer information transmission mechanism model according to the mixed precision quantization result to obtain an optimized inter-layer information transmission mechanism model.
[0025] In this scheme, by constructing the inter-layer information transfer function and the layered sparsification mechanism respectively, the amount of inter-layer information transfer can be accurately quantified and the sparse connection strength of nodes can be controlled, effectively improving the model efficiency; by designing the low-rank decomposition kernel function attention unit, the model's ability to capture complex causal relationships is enhanced; by parallelizing the kernel function attention unit, the calculation speed is greatly improved and the processing time is shortened; by using Tucker to decompose the model parameters and quantize the mixed precision of the parallelization results, the model parameters are significantly reduced and the computing resource consumption is reduced; according to the mixed precision quantization results, the inter-layer information transfer mechanism model is optimized to obtain a more efficient and accurate optimized inter-layer information transfer mechanism model, which improves the overall performance of the virtual power plant inter-layer relationship modeling.
[0026] Further, the S7 includes the following steps: S71. Construct a fault propagation graph based on the optimized inter-layer information transmission mechanism model, and calculate the counterfactual intervention effect based on the fault propagation graph. The mathematical expression is: , Represents a collection of nodes, represents the edge set, represents the fault propagation probability matrix, with elements Represents a slave node The failure propagates to the node probability; The mathematical expression of the counterfactual intervention effect is: , Represents a slave node To Node The counterfactual effect of Indicates the budget for the intervention. Representation node The state variables, Representation node The state variables, Indicates the current state, Indicates a fault condition. Indicates that Set to Under the condition that the node State variables The probability of occurrence, Indicates that Set to Under the condition that the node State variables Probability of occurrence; S72. Construct a cascading failure path identification algorithm based on the fault propagation graph and the counterfactual intervention effect. The mathematical expression of the cascading failure path identification algorithm is: , Indicates the type of fault identified, Represents a given observation Next fault type The posterior probability of Indicates that among all possible fault types, we need to find the posterior probability The biggest failure type ; S73, identifying the potential cascading failure paths of the virtual power plant according to the cascading failure path identification algorithm, and obtaining comprehensive fault warning information, wherein the comprehensive fault warning information The mathematical expression is: , Represents the remaining time of the forecast, Indicates the severity of the fault. Indicates a cascade path.
[0027] In this scheme, by constructing a fault propagation graph based on the optimized inter-layer information transmission mechanism model and calculating the counterfactual intervention effect, the propagation path and probability of the fault between nodes, as well as the impact of intervention on fault propagation, can be intuitively presented; by constructing a cascading fault path identification algorithm, the fault type can be accurately identified based on the fault propagation graph and the counterfactual intervention effect; the algorithm is used to identify the potential cascading fault paths of the virtual power plant, and comprehensive fault warning information including the predicted remaining time, fault severity and cascade path is obtained, which can predict the fault development trend of the virtual power plant in advance, effectively reduce the losses caused by faults, and improve the stability of system operation. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] Figure 1 This is a flow chart of a virtual power plant fault warning method based on a hierarchical interactive causal graph Transformer according to an embodiment of the present invention. DETAILED DESCRIPTION
[0029] The following is further described in detail through specific implementation methods: See also Figure 1 As shown in FIG, it is a flow chart of a virtual power plant fault warning method based on a hierarchical interactive causal graph Transformer according to an embodiment of the present invention, which includes the following steps: S1. Based on the virtual power plant, a three-level hierarchical structure consisting of micro-equipment layer, meso-system layer and macro-market layer is constructed, and a three-level interactive causal diagram is established based on the three-level hierarchical structure; S2. Construct a multi-resolution causal distillation mechanism based on the three-level hierarchical interactive causal graph, and use the multi-resolution causal distillation mechanism to model the global causal relationship of the Transformer encoder-decoder architecture to obtain the Transformer model; S3. Introducing the multi-level contrastive causal attention mechanism into the Transformer model for automatic quantitative analysis to obtain complex causal relationships between layers; S4. Eliminate false correlations in the complex causal relationships between layers based on the high-order Markov random field framework to obtain the actual complex causal relationships between layers; S5. Construct an inter-layer information transmission mechanism model based on the conditional variational graph diffusion model and the actual complex causal relationship between layers; S6. Optimizing the inter-layer information transfer mechanism model based on the inter-layer information transfer function and the layered sparsification mechanism to obtain an optimized inter-layer information transfer mechanism model; S7. Based on the optimized inter-layer information transmission mechanism model, the potential cascading failure paths of the virtual power plant are identified to obtain comprehensive fault warning information.
[0030] Specifically, S1 includes the following steps: S11. Based on the three-level hierarchical structure, clarify the upload path for transient device data from the micro-device layer to the meso-system layer, the logic for issuing control instructions from the meso-system layer to the micro-device layer, and the transmission rules for market signals from the macro-market layer to the meso-system layer. Establish the interactive relationship between the layers based on the upload path and transmission rules. S12. Based on the interaction relationship between the levels, a three-level hierarchical interactive causal diagram is established, in which nodes represent key variables at each level and edges represent the causal direction and causal strength between variables.
[0031] In this embodiment, the micro-device layer processes microsecond-level device transient data, including voltage and current waveforms of power electronic equipment, switching state changes of power converters, and other high-frequency sampling data. The mathematical expression is: ,in, Representative A microscopic device, is the total number of micro-devices in the virtual power plant, and each micro-device A time series feature vector containing a set of transient data of microscopic devices ,in, Indicates at a point in time Microscopic devices The time series feature vector of Represents the characteristic dimension of the micro-level. The meso-system layer processes system status data at the second to minute level, including system-level information such as the topology of the power network, load distribution, and power flow. The mathematical expression is: ,in, Representative mesosystem components, is the total number of mesosystem components in the virtual power plant, and each mesosystem component A time series feature vector containing a set of system status data ,in, Indicates at a point in time System components The time series feature vector of The macro market layer processes market data from hourly to daily levels, including macro-environmental factors such as electricity prices, demand response signals, and market transaction information. The mathematical expression is: ,in, Representative Macro market factors, is the total number of macro market factors that affect the operation of virtual power plants. Contains a set of time series feature vectors ,in Indicates at a point in time Market factors The time series feature vector of It is the macro-level characteristic dimension.
[0032] Specifically, S2 includes the following steps: S21. Analyze the three-level hierarchical interaction causal graph to obtain the adjacency matrix mapping relationship; S22. Based on the adjacency matrix mapping relationship, a multi-resolution causal distillation mechanism is constructed. The low-resolution causal structure is topologically aligned and information is compressed into the high-resolution semantic space according to the multi-resolution causal distillation mechanism to obtain the distilled hierarchical embedding vector. The mathematical expression of the multi-resolution causal distillation mechanism is: ; Indicates the Tier Node pair Tier The attention weight of each node, represents the query parameter matrix, , A parameter matrix representing the key values, , Indicates the Tier The hidden state of each node, The dimension is , Indicates the Tier The hidden state of each node, The dimension is , represents the scaling factor of the attention mechanism, represents a sparse mask matrix based on causal strength, The value of is 0 or 1. Representation node and nodes The structured diffusion distance metric between represents the adaptive time-varying temperature parameter function, represents the structural similarity function, represents the verification function based on causal strength, where Representation node For Node The causal strength of represents the timing adaptive modulation function; S23. Inject the hierarchical embedding vectors into the Transformer encoder-decoder architecture to perform global causal relationship modeling and obtain the Transformer model.
[0033] In this embodiment, the mathematical expression of the sparse mask matrix is: ,in, Represents a slave node To Node The causal strength of is calculated by Granger causality test or structural equation model. Represents the causal relationship threshold, which is initially set to 0.3 and dynamically adjusted according to the system operation status. The mathematical expression is: ,in, It is feature extraction function, is the time-varying feature weight, is the total number of features, feature extraction function It is learned through the autoencoder network and can extract the key features and weights in the node representation. Dynamically adjust according to the importance of features in historical data. Adaptive time-varying temperature parameter function The mathematical expression is: ,in, is the basic temperature parameter, set to 2.5; is the information entropy of the current state of the system; is the reference information entropy, which is set as the average information entropy when the system is operating normally; is the parameter for adjusting the slope, set to 5.0. When the system information entropy increases (that is, the uncertainty of the system state increases), Reduce, make the attention distribution more uniform, and enhance the sensitivity to abnormal signals. Structural similarity function The mathematical expression is: ,in, is the feature mapping function, represents the covariance, represents the standard deviation, Represents the mean, which takes into account the correlation and scale difference of node representations. When two node representations are structurally similar, the larger value is taken. Verification function based on causal strength The mathematical expression is: ,in, is the sigmoid activation function, is the shape parameter, set to 10.0; is causal strength; It is a time-varying verification threshold that is dynamically adjusted with the system operation state. The verification function smoothes the boundary effect of the binary mask while retaining strong causal relationships. The mathematical expression is: ,in, is the modulation amplitude, set to 0.2; is the system period, determined according to the operating characteristics of the virtual power plant; is the phase parameter; is the current system global state vector; is the reference state vector; It is a normalization parameter. The timing modulation function can dynamically adjust the attention weight according to the periodic changes of the system and the degree of state deviation.
[0034] Specifically, S3 includes the following steps: S31. Construct a contrastive learning objective function based on the feature distribution of the original self-attention mechanism of the Transformer model, and obtain a latent space representation containing causal constraints based on the contrastive learning objective function. The mathematical expression of the contrastive learning objective function is: ;in, represents the value of the comparative learning objective function, is a set of positive sample pairs, Contains pairs of nodes that have causal relationships, Represents the set of positive sample pairs All node pairs in expectations, is a node The negative sample set, Contains nodes There is no causal relationship between nodes. Representation node The latent space features and nodes The latent space features The similarity measure between Representation node The latent space features and nodes The latent space features The similarity measure between represents the temperature parameter; S32. Design causal attention units through latent space representation, and construct a causal relationship matrix based on the causal attention units combined with the non-parametric causal discovery algorithm, where the causal attention units The mathematical expression is: ,in, Represents the causal relationship matrix, element Indicates the query location To key position The causal strength of represents element-wise multiplication, 、 、 represent the query matrix, key matrix and value matrix respectively, represents the embedding dimension, represents the similarity score between the query matrix and the key matrix; Cause and Effect Matrix The mathematical expression is: ,in, represents a causal discovery algorithm based on a nonlinear additive noise model, represents the time series causal discovery algorithm; S33. Based on the causal relationship matrix, we introduce sparsification constraints and dynamic threshold adjustment to quantify discrete causal relationships into continuous weight values, generating a dynamic causal relationship matrix that can be embedded in attention calculations. S34. Construct a top-down causal chain tracing path and a bottom-up feature fusion path based on the dynamic causal relationship matrix, and perform hierarchical aggregation on the dynamic causal relationship matrix based on the causal chain tracing path and the feature fusion path to obtain a hierarchical aggregation result, wherein the hierarchical aggregation The mathematical expression is: ,in, Indicates the The causal relationship matrix of the layer, is the layer weight, is the total number of layers; S35. Determine the comparative causal loss function based on the hierarchical aggregation results, and automatically quantify the complex causal relationship between layers based on the comparative causal loss function. The mathematical expression is: ,in, Represents the hierarchical aggregation result, represents the acyclic constraint loss, represents the sparsity constraint loss, represents the acyclic constraint loss The trade-off parameter will be Set to 0.5, represents the sparsity constraint loss The trade-off parameter will be Set to 0.3.
[0035] In this embodiment, the original self-attention mechanism includes: the micro-device layer uses the FlashAttention-4 structure to process device transient data, the meso-system layer introduces Performer-CDF to analyze system status data, and the macro-market layer combines Rotary Position Embedding and RWKV mechanism to process market data.
[0036] Specifically, implementing the FlashAttention-4 structure of the micro-device layer includes: calculating the attention of the micro-device layer, wherein the mathematical expression of the attention of the micro-device layer is: , represents the attention matrix at the micro-device level, 、 、 They represent the query matrix, key matrix and value matrix of the micro-device layer respectively. Block matrix multiplication optimization, block matrix multiplication is defined as: ,in, and are the number of row blocks and column blocks respectively, The query matrix blocks, The first The block size is optimized according to the hardware cache size and is usually set to 64 or 128. A fast normalization algorithm is implemented. The normalization algorithm is defined as: ,in, is the row maximum. By recording and updating the local maximum and local sum of each block, a normalized calculation with O(n) time complexity is achieved within O(1) space complexity. The fast normalization algorithm updates the local statistics incrementally to avoid repeated calculations. A block mechanism for transient waveform characteristic perception is designed. The block mechanism is defined as: ,in, is the basic block size, set to 64; is the scaling factor, set to 0.5; It's time The volatility of the data is measured, and smaller blocks are used in areas with high volatility to improve processing accuracy. FlashAttention-4 optimizes computational complexity, reducing the complexity of attention calculation from O(n²) to O(n·log(n) / d), where n is the sequence length and d is the feature dimension. An adaptive sparse attention strategy is designed for microsecond-level transient waveform features. : ,in, The operation retains the largest elements, and set the other elements to zero; It is adaptive sparsity, which is dynamically adjusted according to signal volatility, further reducing the computational burden while maintaining accuracy.
[0037] Implementing the Performer-CDF variant of the mesosystem layer includes calculating the attention of the mesosystem layer, where the mathematical expression of the attention of the mesosystem layer is: , represents the attention matrix of the meso-system layer, 、 、 Represent the query matrix, key matrix, and value matrix of the meso-system layer respectively. The Performer-CDF variant reduces the complexity of attention calculation by approximating the kernel function. The kernel function approximation is defined as: ,in, is a random feature mapping function, which is defined as: ,in, is the scaling factor, set to 1.0; is the projection dimension, set to 256; is the random projection matrix, ; is the bias vector, ; is the dimension of the key vector. A kernel function variant is designed to adapt to the system topology. The topology-aware kernel function is defined as: ,in, It is a basic kernel function, such as RBF kernel or polynomial kernel; It's a picture midpoint and The shortest path distance between them; is the distance attenuation parameter, set to 0.1, and the topology-aware kernel function enables the attention mechanism to consider the physical connection relationship of the system. The cumulative distribution function weighting is implemented, and the cumulative distribution function weighting mechanism is defined as: ,in, is the attention matrix The cumulative distribution function (CDF) of each row highlights the important attention weights in each row while maintaining the constraint that the row sum is 1. The Performer-CDF variant optimizes the computational complexity of the meso-system layer. It reduces the complexity of attention computation from O(n²) to O(n·d), where n is the number of system nodes and d is the state feature dimension. Furthermore, distributed computing further improves processing speed: ,in, is the number of processing units, It is The local attention matrix calculated by the processing units, It is a distributed reduction operation that enables the meso-system layer to process large-scale system state data in real time.
[0038] Implementing the improved Rotary Position Embedding and RWKV mechanism at the macro market layer includes calculating the attention of the macro market layer, where the mathematical expression of the attention of the macro market layer is: , represents the attention matrix of the macro market layer, Represents market data at the macro market level, represents the rotation position encoding function, Represents the RWKV mechanism. Implements improved rotational position encoding. The rotational position encoding is defined as: ,in, is the input feature vector, is the position index, is the frequency parameter, It is a feature dimension, introducing a time perception mechanism: ,in, is the seasonal adjustment parameter, set to 0.2; It's time The seasonal intensity is calculated by Fourier analysis, so that the position encoding can adapt to the cyclical changes of market data. An enhanced RWKV mechanism is designed, which is defined as: ,in, is the output vector, is the acceptance vector, is the key vector, is a value vector, is the output transformation matrix, is the mixing parameter, set to 0.5; is the adaptive time decay parameter: ,in, is the base decay parameter, set to 24 (corresponding to the number of data points in one day); is the fluctuation adjustment parameter, set to 0.3; It's a window The volatility of internal data and the adaptive decay parameter enable the model to dynamically adjust the memory length according to market volatility. The correlation analysis between market data and physical faults is realized, and the correlation analysis unit is defined as: ,in, It's market factors With physical equipment The strength of the association between is the bicorrelation analysis function, It's market factors Time series data, It is a physical device Time series data, is the time lag parameter. Optimizes the computational cost of long sequence processing. The RWKV mechanism avoids the calculation of the full attention matrix, reducing the complexity of long sequence processing from O(n²) to O(n), where n is the sequence length. Further optimization is achieved through incremental calculation: , ,in, is the accumulated state vector, the initial value Incremental computation enables the model to process market data of arbitrary length while requiring only O(1) additional memory.
[0039] Specifically, S4 includes the following steps: S41. Construct a high-order Markov random field framework. The mathematical expression of the high-order Markov random field framework is: ,in, Indicates system status The joint probability distribution of represents the normalization constant, is the maximum clique set, Indicates that the definition is in the group The higher-order potential function on delegation A subset of random variables in ; S42. Define high-order potential function based on the high-order Markov random field framework. The mathematical expression is: ,in, represents the first-order characteristic function; represents the third-order interaction characteristic function, represents the first-order characteristic function The model parameters, represents the third-order interaction characteristic function The model parameters of , k represents the number of variables involved in the potential function; S43. determining a parameterized energy model based on the high-order potential function, and sampling the parameterized energy model according to the Hamiltonian Monte Carlo method to obtain a Markov chain sample sequence; S44. Analyze the Markov chain sample sequence according to the kernel density conditional probability engine to obtain the structured causal effect matrix. The mathematical expression is: ,in, Indicates except All variables except Indicates that it contains variables All regiments; S45. Optimize the pseudo-likelihood parameters and perform Bootstrap test on the structured causal effect matrix to obtain the complex causal relationship between the actual layers and the objective function of the pseudo-likelihood parameters. The mathematical expression is: ,in, Indicates the training samples, Indicates the In the training samples, except All other variables except represents the variable dimension, represents the set of model parameters, Represents the total number of training samples.
[0040] Specifically, S5 includes the following steps: S51. Based on the complex causal relationship between actual layers, a conditional variational graph diffusion model is constructed. The mathematical expression of the conditional variational graph diffusion model is: ,in, Represents the time step The potential graph of Indicates conditional information. represents the potential state at a given final time step and condition information Under the condition of The conditional probability distribution of represents the number of diffusion steps, Set to 1000, Represents the time step -1 potential graph, represents the latent state at the initial time step, represents the latent state at the final time step; S52. Based on the conditional variational graph diffusion model, the forward diffusion process is defined. The mathematical expression of the forward diffusion process is: ,in, represents the noise scheduling parameter, represents the identity matrix, Represents the time step The potential graph by is the mean, is the multivariate Gaussian distribution with covariance matrix, Indicates the forward diffusion process from time step -1 potential graph To time step The potential graph Conditional probability distribution of noise scheduling parameters It is designed as a linear growth sequence from small to large, with an initial value of 0.0001 and a final value of 0.02.
[0041] S53. Perform reverse denoising on the complex causal relationship between actual layers according to the forward diffusion process to obtain a reverse denoising result. The mathematical expression of the reverse denoising process is: , represents the mean parameterized by the neural network, represents the covariance function, Represents the time step -1 potential graph by is the mean, is the multivariate Gaussian distribution with covariance function, Indicates that at a given time step The potential graph and condition information Under the condition of -1 potential graph The conditional probability distribution of ; S54. Construct a diffusion model using a conditional embedding mechanism based on the inverse denoising result, and construct a fault feature extractor using the diffusion model using the conditional embedding mechanism. The mathematical expression of the diffusion model using the conditional embedding mechanism is: , represents the basic mean prediction network, represents the time modulation function, represents the conditional encoding function, represents element-wise multiplication, In a given latent graph , time step and condition information Under the condition of , the mean generated by the conditional embedding mechanism; Fault Feature Extractor The mathematical expression is: , represents the encoder network, Represents a feature filter; S55. Construct an inter-layer information transfer mechanism model based on the fault feature extractor and the variational learning objective function, wherein the variational learning objective function The mathematical expression is: , represents standard normal noise, represents the noise prediction network, express and The joint distribution of .
[0042] Specifically, S6 includes the following steps: S61. Construct an inter-layer information transfer function and a layered sparsification mechanism respectively. The mathematical expression of the inter-layer information transfer function is: , Indicates that from Tier Node to Tier The amount of information transmitted by each node, represents the parameterized transfer function, represents the symmetric mutual information measure.
[0043] In this embodiment, Defined as: ,in, is mutual information, and They are and The symmetric mutual information metric is normalized to the interval [0,1] to facilitate the comparison of information transfer between different levels.
[0044] The mathematical expression of the layered sparsification mechanism is: , Representation node To Node The sparse connection strength, represents the sigmoid function, represents the sparsification control parameter, represents the thermodynamic mutual information, Represents a dynamic threshold function.
[0045] In this embodiment, Defined as: ,in, is the temperature parameter, set to 0.1, and They are conditional entropy, Defined as: ,in, is the basic threshold, set to 0.2; is the periodic fluctuation amplitude, set to 0.05; is the system period; is the abnormal response parameter, set to 0.1; It's time The dynamic threshold function enables the sparsification mechanism to adapt to the periodic changes and abnormal states of the system.
[0046] S62. Design a low-rank decomposition kernel function attention unit based on the inter-layer information transfer function and the layered sparsification mechanism. The expression of the kernel function attention unit is: , represents the kernel function attention matrix, 、 、 represent the query matrix, key matrix and value matrix respectively, Represents the query matrix The result after kernel function feature mapping, Represents the bond matrix Transpose the kernel feature map.
[0047] S63. Parallelize the kernel function attention unit to obtain a parallelized result, where the mathematical expression of the parallel computing strategy is: , represents the parallelized matrix, Indicates the The partial attention matrix calculated by the processing unit, Indicates the number of processing units, Represents a matrix concatenation operation.
[0048] S64. Perform model parameter decomposition on the parallelization result according to Tucker to obtain a model parameter decomposition result, and perform mixed precision quantization on the model parameter decomposition result to obtain a mixed precision quantization result. The mathematical expression of Tucker is: , represents the quantized weight, represents the original weight, represents the scaling factor, Represents quantization operations; different quantization accuracies are used according to the importance of different layers and different parameters. Key layers and sensitive parameters use 8-bit or 16-bit quantization, and secondary layers and insensitive parameters use 4-bit or 2-bit quantization. The mixed precision strategy minimizes the model size while maximizing the prediction accuracy.
[0049] S65. Optimize the inter-layer information transmission mechanism model according to the mixed precision quantization result to obtain an optimized inter-layer information transmission mechanism model.
[0050] Specifically, S7 includes the following steps: S71. Construct a fault propagation graph based on the optimized inter-layer information transmission mechanism model, and calculate the counterfactual intervention effect based on the fault propagation graph. The mathematical expression is: , Represents a collection of nodes, represents the edge set, represents the fault propagation probability matrix, with elements Represents a slave node The failure propagates to the node probability.
[0051] In this embodiment, Defined as: ,in, is the weight of historical data, set to 0.7; is the probability of fault propagation calculated based on historical data, is the fault propagation probability based on expert knowledge.
[0052] The mathematical expression of the counterfactual intervention effect is: , Represents a slave node To Node The counterfactual effect of Indicates the budget for the intervention. Representation node The state variables, Representation node The state variables, Indicates the current state, Indicates a fault condition. Indicates that Set to Under the condition that the node State variables The probability of occurrence, Indicates that Set to Under the condition that the node State variables Probability of occurrence.
[0053] S72. Construct a cascading failure path identification algorithm based on the fault propagation graph and the counterfactual intervention effect. The mathematical expression of the cascading failure path identification algorithm is: , Indicates the type of fault identified, Represents a given observation Next fault type The posterior probability of Indicates that among all possible fault types, we need to find the posterior probability The biggest failure type .
[0054] In this implementation, Defined as: ,in, is the likelihood function, is the prior probability of the fault type.
[0055] S73, identifying the potential cascading failure paths of the virtual power plant according to the cascading failure path identification algorithm, and obtaining comprehensive fault warning information, wherein the comprehensive fault warning information The mathematical expression is: , Represents the remaining time of the forecast, Indicates the severity of the fault. Indicates a cascade path.
[0056] In this embodiment, the fault occurrence time prediction is defined as: ,in, is the predicted time of failure, is the current time, is the predicted remaining time, Defined as: ,in, is a random variable of the time of failure occurrence, is the observation time, is the currently observed system state. Defined as: ,in, is the severity of the fault, is the evaluation function, is the fault characteristic, It is the system status context. The evaluation function comprehensively considers the impact of the fault on system performance, maintenance cost, safety risk and potential cascading failure range, and gives a severity score between 0 and 10.
[0057] Based on the above experimental analysis: This experiment was conducted on a large-scale virtual power plant test platform of the State Grid, which integrates three different types of power generation units, including photovoltaic power generation units, wind power generation units and energy storage units, as well as 12 distributed power load units and a complete distribution network structure. The hardware environment uses an Intel Xeon Gold 6248R processor with a main frequency of 3.0GHz, equipped with a 24-core processing unit, 128GB of system memory RAM, and an NVIDIA A100 GPU as the graphics processing unit. The software environment is based on the Python 3.9 programming language, the deep learning framework uses PyTorch 1.12, the network analysis tool uses Network X 2.8, and is combined with the independently developed power system simulation platform ElectroSim v4.3 for experimental verification. The test dataset contains actual operation data from January 2021 to December 2022, with a total data volume of approximately 4.3TB, covering 374 labeled fault types and 2,867 complete fault sequences, providing rich data support for the experiment.
[0058] Comparison plan 1 description: Comparison Option 1 employs a traditional deep learning approach, with a core architecture that combines LSTM and CNN. This approach uses a long short-term memory network to capture temporal features, combined with a convolutional neural network to extract spatial features, and applies a standard attention mechanism to enhance the model's ability to identify key information. However, this approach has significant limitations: it processes data only at a single time scale and cannot capture the complex interactions between information at different time scales; it fails to incorporate causal reasoning, resulting in a poor understanding of the causal links within the system; it lacks a hierarchical structure, making it unable to distinguish and process system data at different levels; and it uses linear thinking to analyze fault modes, limiting its ability to identify nonlinear and complex fault patterns. This model has a total of 127M parameters and takes approximately 48 hours to train. It requires high computing resources and is not suitable for deployment on edge devices.
[0059] Comparison plan 2 description: Comparison Option 2 employs a modified Transformer approach, which optimizes the standard Transformer architecture to a certain extent. This approach introduces a simple two-layer structure, dividing the system into a device layer and a system layer, and uses a basic causal diagram to assist in constructing internal system relationships. Compared to Comparison Option 1, this approach is able to partially capture the interactions between systems and provide preliminary identification of fault propagation paths. However, this approach still has some shortcomings: its computational complexity remains at the O(n²) level, making it inefficient for processing long sequences of data; it lacks an adaptive inter-layer information transfer mechanism, making it difficult to accurately describe the information flow between different layers; its causal diagram construction is relatively static and cannot adapt to dynamic system changes; and its overly simplistic hierarchical division ignores the potential impact of market-level factors on system failures. This model has 82M parameters and takes approximately 36 hours to train. While this model represents an improvement over Comparison Option 1, it still leaves much room for improvement in computational efficiency and real-time performance.
[0060] First, data preprocessing is performed, and the raw data is layered according to different time scales. The micro-device layer processes microsecond-level device transient data, the meso-system layer processes minute-level system status data, and the macro-market layer processes hourly and daily market data. All data are standardized to eliminate the influence of different dimensions, and specific feature extraction methods are designed for different levels to lay the foundation for subsequent model construction. Then, three comparative models are constructed: the LSTM+CNN hybrid architecture of comparative solution 1, the improved Transformer structure of comparative solution 2, and the hierarchical interactive causal graph Transformer model of the present invention. After the model is built, data from January 2021 to June 2022 (approximately 70% of the total data volume) is used for training. The Adam optimizer is used, the initial learning rate is set to 1e-4, and the cosine decay strategy is used to dynamically adjust the learning rate. During the training process, the batch size is set to 128, and a total of 300 epochs of iterative training are performed. After training, data from July to September 2022 (approximately 15% of the total data volume) was used for validation and hyperparameter tuning. A grid search method was used to optimize key hyperparameters, including the number of attention heads, inter-layer transfer strength parameters, and sparsification control parameters. Finally, data from October to December 2022 (approximately 15% of the total data volume) was used as an independent test set to comprehensively evaluate the performance of the three methods. Furthermore, specially designed system load tests and anti-interference tests were conducted. The system load test tested the inference speed, resource utilization, and deployment difficulty of the three methods under different computing resource scales (ranging from high-performance servers to edge computing devices). The anti-interference test tested the robustness of the three methods by injecting varying degrees of noise and simulating extreme operating conditions (such as sudden changes in device load and communication interruptions).
[0061] To comprehensively evaluate the performance of the three methods, a comprehensive testing method and standards were designed. Warning accuracy was calculated using a standard classification evaluation metric: Accuracy = (TP + TN) / (TP + TN + FP + FN), where TP represents true positives (correct warnings), TN represents true negatives (correct no warnings), FP represents false positives (false warnings), and FN represents false negatives (missed warnings). Warning lead time is defined as the interval (in minutes) between the issuance of a warning and the actual occurrence of a fault. This metric reflects the response time provided by the warning system to operations and maintenance personnel. The weak fault signal detection rate was determined by artificially injecting 100 low-intensity fault signals of varying types and intensities into the system and calculating the proportion of successful detections. The false alarm rate, FalseAlarmRate, was calculated using the formula: FalseAlarmRate = FP / (FP + TN). This metric represents the proportion of false alarms issued by the system when there are no actual faults. This metric significantly impacts the practical application value of the system. Computational efficiency was assessed by measuring the average time (in milliseconds) per inference. The average value was obtained from multiple repeated tests conducted under different computing environments. The model compression rate is defined as the ratio of the compressed model size to the original model size, expressed as a percentage. The cascading fault path recognition rate is calculated by comparing the consistency of the cascading fault path identified by the system with the actual fault propagation path, reflecting the depth of the system's understanding of the fault propagation mechanism. Edge device compatibility is evaluated by deploying the model on typical edge computing devices (Jetson Nano, Raspberry Pi 4, etc.) and testing the running status, including frame rate, memory usage, power consumption and other indicators for comprehensive evaluation. The adaptability to complex working conditions is scored on a 0-10 point scale, and an expert panel scores the system based on its performance under various complex working conditions. The scoring criteria include adaptability to nonlinear system behavior, stability under abnormal working conditions, and versatility in scenarios of different complexity. The experimental results are shown in Table 1: Table 1 Comparison of experimental results
[0062] As shown in Table 1, the hierarchical interactive causal graph Transformer method of the present invention significantly outperforms the comparative solution across multiple key performance metrics. In terms of early warning accuracy, the present invention achieves a high accuracy of 99.85% ± 0.12%, an improvement of 1.15 percentage points compared to the 98.70% ± 0.36% achieved by Comparative Solution 2. The error fluctuation is also significantly reduced, demonstrating more stable and reliable prediction results. This performance improvement is primarily due to the multi-resolution causal distillation mechanism, which dynamically adjusts the information flow between different layers, accurately capturing the complex interactions between the micro-device layer, the meso-system layer, and the macro-market layer. Warning lead time is a key metric for fault early warning systems. The present invention achieves a warning lead time of 92.0 ± 4.8 minutes, 19 minutes earlier than Comparative Solution 2, a 26.0% improvement. This significant improvement stems from the present invention's multi-level comparative causal attention mechanism and conditional variational graph diffusion model, which are able to identify potential fault modes even when fault signatures are still subtle, providing operators with more time to react. The weak fault signal detection rate is a key metric for evaluating the sensitivity of early warning systems. The present invention achieved a high level of 99.6% ± 0.3% in this metric, a 7.1 percentage point improvement over the 92.5% ± 2.4% of Comparative Solution 2. Simultaneously, the false alarm rate was significantly reduced to 0.30% ± 0.08%, a 62.5% reduction compared to the 0.80% ± 0.18% of Comparative Solution 2. The simultaneous optimization of these two metrics demonstrates that the present invention effectively reduces the risk of false alarms while improving early warning sensitivity. This is due to the ability of the high-order Markov random field framework to eliminate spurious correlations. In terms of computational performance, the present invention achieves significant improvements through a multi-kernel function attention mechanism and low-rank decomposition techniques. The inference speed is only 1.45 ± 0.21 milliseconds, 8.7 times faster than Comparative Solution 2. The end-to-end latency is reduced to 5.8 ± 0.38 milliseconds, a reduction of 86.6%. The model size is only 24.9 MB, a 92.0% reduction compared to the 312.3 MB of Comparative Solution 2. The computing resource requirement is reduced to 12.8 GFLOPS, a reduction of 84.8%. These performance optimizations enable the present invention to run efficiently on edge computing devices. The measured frame rate on a Raspberry Pi 4 reached 42.5 FPS, an increase of 2,261.1% over Comparative Solution 2, making real-time early warning on the edge possible. The recognition rate of cascading fault paths reached 99.3% ± 0.4%, an increase of 17.4 percentage points over Comparative Solution 2. This is because the probabilistic counterfactual reasoning mechanism of the present invention can accurately identify potential fault propagation paths, generate accurate risk propagation maps, and significantly enhance the interpretability of the system. The adaptability score for complex working conditions reached 9.7 ± 0.2 points (out of 10 points), an increase of 32.9% over Comparative Solution 2, indicating that the present invention performs well under various unconventional working conditions and has strong generalization capabilities.The fault type identification accuracy reached 99.2%±0.5%, which is 6.2 percentage points higher than that of comparison scheme 2. This shows that the present invention can not only predict the occurrence of faults, but also accurately identify the fault type, providing more targeted guidance for fault handling.
[0063] To further demonstrate the applicability and scalability of the proposed method in virtual power plant scenarios of varying types and sizes, this example details the application of a virtual power plant fault warning method based on a hierarchical interactive causal graph (Transformer) in a large hybrid energy virtual power plant. This hybrid energy virtual power plant comprises a 150MW photovoltaic power station, a 120MW wind farm, an 80MW cascade hydropower station, a 60MWh electrochemical energy storage system, and a 40MW demand response resource. These are distributed across four different geographic locations and coordinated through a regional energy internet.
[0064] In this extended application scenario, the complexity of the micro-device layer is significantly increased, including 12,600 photovoltaic panels, 45 wind turbines, 18 hydroelectric generators, 6 sets of energy storage converter equipment, and 820 industrial load control units. In order to handle such a complex device layer structure, this embodiment expands the FlashAttention-4 structure of the micro-device layer and introduces a multi-head grouping processing mechanism. Specifically, all micro-devices are divided into five main categories according to their functional characteristics. Each type of device is processed using an independent attention head, and then the outputs of each head are merged through a cross-category attention integration unit. This grouping processing mechanism can be expressed as: A_micro=MultiHeadIntegration(FlashAttention-4_PV(Q_PV,K_PV,V_PV),FlashAttention-4_Wind(Q_Wind,K_Wind,V_Wind),FlashAtt ention-4_Hydro(Q_Hydro,K_Hydro,V_Hydro),FlashAttention-4_ESS(Q_ESS,K_ESS,V_ESS),FlashAttention-4_DR(Q_DR,K_DR,V_DR)); Among them, A_micro represents the final output result obtained after the micro-device layer passes through the multi-head grouping processing mechanism, MultiHeadIntegration (∙) represents the multi-head integration function, FlashAttention-4_PV(Q_PV,K_PV,V_PV) represents the FlashAttention-4 structure processing function for photovoltaic modules, FlashAttention-4_Wind(Q_Wind,K_Wind,V_Wind) represents the FlashAttention-4 structure processing function for wind turbines, and FlashAttention-4_Hydro (Q_Hydro, K_Hydro, V_Hydro) represents the FlashAttention-4 structure processing function for hydro-generators, FlashAttention-4_ESS (Q_ESS, K_ESS, V_ESS) represents the FlashAttention-4 structure processing function for energy storage converters, and FlashAttention-4_DR (Q_DR, K_DR, V_DR) represents the FlashAttention-4 structure processing function for industrial load control units. Different attention window sizes are used for each device type, dynamically adjusted based on their respective signal characteristics. For example, photovoltaic devices use a 256ms window to capture rapid irradiance variations, while wind turbines use a 512ms window to capture wind speed fluctuations. This heterogeneous window design further reduces the overall computational complexity to O(∑(n_i·log(n_i) / d_i)), where n_i represents the sequence length for the i-th device and d_i represents the feature dimension for the i-th device.
[0065] To address the unique characteristics of distributed virtual power plants, the meso-system layer introduces a geo-topology-aware Performer-CDF variant to enhance system state analysis by integrating spatial distance information. To adapt to the multi-energy flow characteristics of hybrid energy virtual power plants, the macro-market layer designs a multi-energy flow market data analysis unit, extending the existing Rotary Position Embedding and RWKV mechanisms. This unit introduces an energy type encoding, E_type, to differentiate the representations of different types of energy market data while maintaining the ability to analyze their associations: E_combined = RotaryEmbed(X_macro) + a × E_type; where a represents the type encoding strength factor, balancing the weight of temporal location information with energy type information. Furthermore, a specific energy price association analysis network layer is designed to capture the complex dynamic relationships between electricity prices, natural gas prices, carbon emission rights prices, and ancillary service prices, enabling early identification of operational risks caused by market anomalies.
[0066] In terms of inter-layer information transmission, this embodiment introduces a dynamic routing mechanism (DynamicRouting) to enhance the adaptability of the hierarchical interactive causal graph, which is particularly suitable for scenarios where the virtual power plant operation mode frequently switches. The dynamic routing mechanism is implemented by iteratively calculating and updating the routing probability matrix P_route, which is specifically expressed as: P_route^(t+1)=softmax(P_route^t+b_ij·cos(T_i^l,H_j^(L-1)));where b_ij represents the routing coefficient, and cos(T_i^l,H_j^(L-1))) represents the cosine similarity between the transformed features of the L layer and the hidden state of the L-1 layer. The iterative update mechanism allows the model to dynamically adjust the information flow between layers, so that key information can be adaptively transmitted between layers according to the current operating state, significantly improving the model's adaptability to operating mode switching.
[0067] In terms of fault feature extraction, combined with the multi-source heterogeneous data characteristics of the virtual power plant, this embodiment expands the conditional variational graph diffusion model and introduces a multimodal conditional encoder. The encoder can simultaneously process time series data X_time, image data X_image (such as infrared thermal imaging) and text data X_text (such as operation records), and integrate multiple modal information for fault feature extraction through a unified latent space representation: c_multimodal=Encoder_time(X_time)⊕Encoder_image(X_image)⊕Encoder_text(X_text), where ⊕ represents a feature fusion operation, which is specifically implemented as an attention-weighted feature connection. Improve the sensitivity of fault detection while maintaining a low computational burden, especially for equipment aging faults with complex characterization methods. To meet the deployment requirements of edge computing devices, this embodiment further optimizes the model compression technology and designs a quantization-aware training strategy for different levels. By simulating quantization noise during training, the model can proactively adapt to the accuracy loss L_total during deployment: L_total = L_prediction + β × L_quantization, where L_prediction represents the prediction loss, L_quantization represents the quantization loss, and β represents the trade-off coefficient. When reducing the precision from FP32 to INT8, the prediction accuracy loss is kept within 0.3%, while the model size is further reduced to 5.6% of the original size, enabling millisecond-level inference on edge devices with the ARM Cortex-A72 architecture.
[0068] In terms of fault warning output, this embodiment designs a hierarchical warning mechanism to automatically generate differentiated warning information according to different user roles (such as operation and maintenance personnel, dispatchers, and management decision-makers). For cascading fault analysis, this embodiment expands the counterfactual reasoning framework and introduces a time window rolling prediction mechanism, which can predict the fault development path at different time points in the future. This mechanism is implemented by constructing nested counterfactual models on multiple time scales, supporting prediction spans from minutes to days, and providing warning information and handling time window estimates of different granularities for various types of faults. Compared with the traditional single time point prediction method, this mechanism increases the advance time of key node fault prediction by an average of 28 minutes, which provides operation and maintenance personnel with more sufficient handling time.
[0069] The above are only embodiments of the present invention. Common knowledge such as the known specific structures and characteristics in the scheme are not described in detail here. Ordinary technicians in the field are aware of all common technical knowledge in the technical field of the invention before the application date or priority date, can obtain all existing technologies in the field, and have the ability to apply conventional experimental means before that date. Ordinary technicians in the field can improve and implement this scheme in combination with their own abilities under the inspiration given by this application. Some typical known structures or known methods should not become obstacles for ordinary technicians in the field to implement this application. It should be pointed out that for those skilled in the art, without departing from the structure of the present invention, several variations and improvements can be made, which should also be regarded as the scope of protection of the present invention. These will not affect the effect of the implementation of the present invention and the practicality of the patent. The scope of protection required by this application shall be based on the content of its claims, and the specific implementation methods and other records in the specification can be used to interpret the content of the claims.
Claims
1. A virtual power plant fault warning method based on hierarchical interactive causal graph Transformer, characterized by: The following steps are involved: S1. Based on the virtual power plant, a three-level hierarchical structure consisting of micro-equipment layer, meso-system layer and macro-market layer is constructed, and a three-level interactive causal diagram is established based on the three-level hierarchical structure; S2. Construct a multi-resolution causal distillation mechanism based on the three-level hierarchical interactive causal graph, and use the multi-resolution causal distillation mechanism to model the global causal relationship of the Transformer encoder-decoder architecture to obtain the Transformer model; S3. Introducing the multi-level contrastive causal attention mechanism into the Transformer model for automatic quantitative analysis to obtain complex causal relationships between layers; S4. Eliminate false correlations in the complex causal relationships between layers based on the high-order Markov random field framework to obtain the actual complex causal relationships between layers; S5. Construct an inter-layer information transmission mechanism model based on the conditional variational graph diffusion model and the actual complex causal relationship between layers; S6. Optimizing the inter-layer information transfer mechanism model based on the inter-layer information transfer function and the layered sparsification mechanism to obtain an optimized inter-layer information transfer mechanism model; S7. Based on the optimized inter-layer information transmission mechanism model, the potential cascading failure paths of the virtual power plant are identified to obtain comprehensive fault warning information.
2. The virtual power plant fault warning method based on hierarchical interactive causal graph Transformer according to claim 1 is characterized by: Said S1 comprises the following steps: S11. Based on the three-level hierarchical structure, clarify the upload path for transient device data from the micro-device layer to the meso-system layer, the logic for issuing control instructions from the meso-system layer to the micro-device layer, and the transmission rules for market signals from the macro-market layer to the meso-system layer. Establish the interactive relationship between the layers based on the upload path and transmission rules. S12. Based on the interaction relationship between the levels, a three-level hierarchical interactive causal diagram is established, in which nodes represent key variables at each level and edges represent the causal direction and causal strength between variables.
3. The virtual power plant fault warning method based on hierarchical interactive causal graph Transformer according to claim 1 is characterized by: The S2 comprises the following steps: S21. Analyze the three-level hierarchical interaction causal graph to obtain the adjacency matrix mapping relationship; S22. Based on the adjacency matrix mapping relationship, a multi-resolution causal distillation mechanism is constructed. The low-resolution causal structure is topologically aligned and information is compressed into the high-resolution semantic space according to the multi-resolution causal distillation mechanism to obtain the distilled hierarchical embedding vector. The mathematical expression of the multi-resolution causal distillation mechanism is: ; Indicates the Tier Node pair Tier The attention weight of each node, represents the query parameter matrix, , A parameter matrix representing the key values, , Indicates the Tier The hidden state of each node, The dimension is , Indicates the Tier The hidden state of each node, The dimension is , represents the scaling factor of the attention mechanism, represents a sparse mask matrix based on causal strength, The value of is 0 or 1. Representation node and nodes The structured diffusion distance metric between represents the adaptive time-varying temperature parameter function, represents the structural similarity function, represents the verification function based on causal strength, where Representation node For Node The causal strength of represents the timing adaptive modulation function; S23. Inject the hierarchical embedding vectors into the Transformer encoder-decoder architecture to perform global causal relationship modeling and obtain the Transformer model.
4. The virtual power plant fault warning method based on hierarchical interactive causal graph Transformer according to claim 1 is characterized by: The S3 includes the following steps: S31. Construct a contrastive learning objective function based on the feature distribution of the original self-attention mechanism of the Transformer model, and obtain a latent space representation containing causal constraints based on the contrastive learning objective function. The mathematical expression of the contrastive learning objective function is: ;in, represents the comparative learning objective function value, is a set of positive sample pairs, Contains pairs of nodes that have causal relationships, Represents the set of positive sample pairs All node pairs in expectations, is a node The negative sample set, Contains nodes There is no causal relationship between nodes. Representation node The latent space features and nodes The latent space features The similarity measure between Representation node The latent space features and nodes The latent space features The similarity measure between represents the temperature parameter; S32. Design causal attention units through latent space representation, and construct a causal relationship matrix based on the causal attention units combined with the non-parametric causal discovery algorithm, where the causal attention units The mathematical expression is: ,in, Represents the causal relationship matrix, element Indicates the query location To key position The causal strength of represents element-wise multiplication, 、 、 represent the query matrix, key matrix and value matrix respectively, represents the embedding dimension, represents the similarity score between the query matrix and the key matrix; Cause and Effect Matrix The mathematical expression is: ,in, represents a causal discovery algorithm based on a nonlinear additive noise model, represents the time series causal discovery algorithm; S33. Based on the causal relationship matrix, we introduce sparsification constraints and dynamic threshold adjustment to quantify discrete causal relationships into continuous weight values, generating a dynamic causal relationship matrix that can be embedded in attention calculations. S34. Construct a top-down causal chain tracing path and a bottom-up feature fusion path based on the dynamic causal relationship matrix, and perform hierarchical aggregation on the dynamic causal relationship matrix based on the causal chain tracing path and the feature fusion path to obtain a hierarchical aggregation result, wherein the hierarchical aggregation The mathematical expression is: ,in, Indicates the The causal relationship matrix of the layer, is the layer weight, is the total number of layers; S35. Determine the comparative causal loss function based on the hierarchical aggregation results, and automatically quantify the complex causal relationship between layers based on the comparative causal loss function. The mathematical expression is: ,in, Represents the hierarchical aggregation result, represents the acyclic constraint loss, represents the sparsity constraint loss, represents the acyclic constraint loss The trade-off parameter, represents the sparsity constraint loss trade-off parameters.
5. The virtual power plant fault early warning method based on hierarchical interactive causal graph Transformer according to claim 4 is characterized by: The micro-device layer uses the FlashAttention-4 structure to process device transient data, the meso-system layer introduces Performer-CDF to analyze system status data, and the macro-market layer combines Rotary PositionEmbedding and RWKV mechanisms to process market data.
6. The virtual power plant fault warning method based on hierarchical interactive causal graph Transformer according to claim 5 is characterized by: The original self-attention mechanism includes: The attention of the micro-device layer is calculated, where the mathematical expression of the attention of the micro-device layer is: , represents the attention matrix at the micro-device level, 、 、 represent the query matrix, key matrix, and value matrix of the micro-device layer respectively; The attention of the meso-system layer is calculated, where the mathematical expression of the attention of the meso-system layer is: , represents the attention matrix of the meso-system layer, 、 、 They represent the query matrix, key matrix and value matrix of the meso-system layer respectively; The attention of the macro market layer is calculated, where the mathematical expression of the attention of the macro market layer is: , represents the attention matrix of the macro market layer, Represents market data at the macro market level, represents the rotation position encoding function, Indicates the RWKV mechanism.
7. The virtual power plant fault warning method based on hierarchical interactive causal graph Transformer according to claim 1 is characterized by: The S4 comprises the following steps: S41. Construct a high-order Markov random field framework. The mathematical expression of the high-order Markov random field framework is: ,in, Indicates system status The joint probability distribution of represents the normalization constant, is the maximum clique set, Indicates that the definition is in the group The higher-order potential function on delegation A subset of random variables in ; S42. Define high-order potential function based on the high-order Markov random field framework. The mathematical expression is: ,in, represents the first-order characteristic function; represents the third-order interaction characteristic function, represents the first-order characteristic function The model parameters, represents the third-order interaction characteristic function The model parameters of , k represents the number of variables involved in the potential function; S43. determining a parameterized energy model based on the high-order potential function, and sampling the parameterized energy model according to the Hamiltonian Monte Carlo method to obtain a Markov chain sample sequence; S44. Analyze the Markov chain sample sequence according to the kernel density conditional probability engine to obtain the structured causal effect matrix. The mathematical expression is: ,in, Indicates except All variables except Indicates that it contains variables All regiments; S45. Optimize the pseudo-likelihood parameters and perform Bootstrap test on the structured causal effect matrix to obtain the complex causal relationship between the actual layers and the objective function of the pseudo-likelihood parameters. The mathematical expression is: ,in, Indicates the training samples, Indicates the In the training samples, except All other variables except represents the variable dimension, represents the set of model parameters, Represents the total number of training samples.
8. The virtual power plant fault warning method based on hierarchical interactive causal graph Transformer according to claim 1 is characterized by: The S5 comprises the following steps: S51. Based on the complex causal relationship between actual layers, a conditional variational graph diffusion model is constructed. The mathematical expression of the conditional variational graph diffusion model is: ,in, Represents the time step The potential graph of Indicates conditional information. represents the potential state at a given final time step and condition information Under the condition of The conditional probability distribution of represents the number of diffusion steps, Represents the time step -1 potential graph, represents the latent state at the initial time step, represents the latent state at the final time step; S52. Based on the conditional variational graph diffusion model, the forward diffusion process is defined. The mathematical expression of the forward diffusion process is: ,in, represents the noise scheduling parameter, represents the identity matrix, Represents the time step The potential graph by is the mean, is the multivariate Gaussian distribution with covariance matrix, Indicates the forward diffusion process from time step -1 potential graph To time step The potential graph The conditional probability distribution of ; S53. Perform reverse denoising on the complex causal relationship between actual layers according to the forward diffusion process to obtain a reverse denoising result. The mathematical expression of the reverse denoising process is: , represents the mean parameterized by the neural network, represents the covariance function, Represents the time step -1 potential graph by is the mean, is the multivariate Gaussian distribution with covariance function, Indicates that at a given time step The potential graph and condition information Under the condition of -1 potential graph The conditional probability distribution of ; S54. Construct a diffusion model using a conditional embedding mechanism based on the inverse denoising result, and construct a fault feature extractor using the diffusion model using the conditional embedding mechanism. The mathematical expression of the diffusion model using the conditional embedding mechanism is: , represents the basic mean prediction network, represents the time modulation function, represents the conditional encoding function, represents element-wise multiplication, In a given latent graph , time step and condition information Under the condition of , the mean generated by the conditional embedding mechanism; Fault Feature Extractor The mathematical expression is: , represents the encoder network, Represents a feature filter; S55. Construct an inter-layer information transfer mechanism model based on the fault feature extractor and the variational learning objective function, wherein the variational learning objective function The mathematical expression is: , represents standard normal noise, represents the noise prediction network, express and The joint distribution of .
9. The virtual power plant fault warning method based on hierarchical interactive causal graph Transformer according to claim 1 is characterized by: The S6 comprises the following steps: S61. Construct an inter-layer information transfer function and a layered sparsification mechanism respectively. The mathematical expression of the inter-layer information transfer function is: , Indicates that from Tier Node to Tier The amount of information transmitted by each node, represents the parameterized transfer function, represents the symmetric mutual information measure; The mathematical expression of the layered sparsification mechanism is: , Representation node To Node The sparse connection strength, represents the sigmoid function, represents the sparsification control parameter, represents the thermodynamic mutual information, represents the dynamic threshold function; S62. Design a low-rank decomposition kernel function attention unit based on the inter-layer information transfer function and the layered sparsification mechanism. The expression of the kernel function attention unit is: , represents the kernel function attention matrix, 、 、 represent the query matrix, key matrix and value matrix respectively, Represents the query matrix The result after kernel function feature mapping, Represents the bond matrix Transpose the kernel function feature map; S63. Parallelize the kernel function attention unit to obtain a parallelized result, where the mathematical expression of the parallel computing strategy is: , represents the parallelized matrix, Indicates the The partial attention matrix calculated by the processing unit, Indicates the number of processing units, Represents matrix concatenation operation; S64. Perform model parameter decomposition on the parallelization result according to Tucker to obtain a model parameter decomposition result, and perform mixed precision quantization on the model parameter decomposition result to obtain a mixed precision quantization result. The mathematical expression of Tucker is: , represents the quantized weight, represents the original weight, represents the scaling factor, Represents a quantized operation; S65. Optimize the inter-layer information transmission mechanism model according to the mixed precision quantization result to obtain an optimized inter-layer information transmission mechanism model.
10. The virtual power plant fault warning method based on hierarchical interactive causal graph Transformer according to claim 1 is characterized by: The S7 comprises the following steps: S71. Construct a fault propagation graph based on the optimized inter-layer information transmission mechanism model, and calculate the counterfactual intervention effect based on the fault propagation graph. The mathematical expression is: , Represents a collection of nodes, represents the edge set, represents the fault propagation probability matrix, with elements Represents a slave node The failure propagates to the node probability; The mathematical expression of the counterfactual intervention effect is: , Represents a slave node To Node The counterfactual effect of Indicates the budget for the intervention. Representation node The state variables, Representation node The state variables, Indicates the current state, Indicates a fault condition. Indicates that Set to Under the condition that the node State variables The probability of occurrence, Indicates that Set to Under the condition that the node State variables Probability of occurrence; S72. Construct a cascading failure path identification algorithm based on the fault propagation graph and the counterfactual intervention effect. The mathematical expression of the cascading failure path identification algorithm is: , Indicates the type of fault identified, Represents a given observation Next fault type The posterior probability of Indicates that among all possible fault types, we need to find the posterior probability The biggest failure type ; S73, identifying the potential cascading failure paths of the virtual power plant according to the cascading failure path identification algorithm, and obtaining comprehensive fault warning information, wherein the comprehensive fault warning information The mathematical expression is: , Represents the remaining time of the forecast, Indicates the severity of the fault. Indicates a cascade path.
Citation Information
Patent Citations
Virtual power plant energy state sensing method
CN117951577A
Cited By
Semiconductor anomaly detection method and system based on physical causal relationship modeling
CN120781269A
Semiconductor anomaly detection method and system based on physical causal relationship modeling
CN120781269B
Genetic algorithm-based polycaprolactone polyol synthesis path optimization method
CN120808929A
Industrial process remaining time prediction method and device and storage medium
CN120930074A
Industrial process fault detection method based on space-time causal graph auto-encoder
CN120974245A