Multi-source heterogeneous data fusion power grid whole process scene chain risk early warning method

By integrating multi-source heterogeneous data and using dynamic early warning threshold algorithms, the problem of false alarms and missed alarms in traditional power grid risk early warning technologies under conditions of high renewable energy penetration and the entire life cycle of equipment has been solved, thus achieving accurate early warning and full-cycle management of power grid risks.

CN121458480APending Publication Date: 2026-02-03LUOYANG MENGJIN POWER SUPPLY CO OF STATE GRID HENAN ELECTRIC POWER CO
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202511303410.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-12
Publication Date
2026-02-03

AI Technical Summary

Technical Problem

Traditional power grid risk early warning technologies have shortcomings in multi-source heterogeneous data fusion, scenario chain construction, and risk threshold design. They cannot adapt to complex scenarios with high new energy penetration and the entire life cycle of equipment, resulting in high false alarm and false alarm rates, and failing to meet the needs of full-cycle risk management.

Method used

A risk early warning method for the entire power grid scenario chain is adopted by fusing multi-source heterogeneous data. Through data preprocessing, multi-layer fusion architecture, coupling correlation algorithm and dynamic early warning threshold algorithm, combined with equipment health status, environmental interference and real-time operating conditions, a precise risk scenario chain and dynamic early warning mechanism are constructed.

Benefits of technology

It improves the accuracy and adaptability of risk warnings, can dynamically respond to changes in complex power grid scenarios, reduce false alarms and missed alarms, and provide full-cycle risk management support.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121458480A_ABST
    Figure CN121458480A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-source heterogeneous data fusion power grid whole process scene chain risk early warning method. The method comprises the following steps: S1, collecting and preprocessing power grid multi-source heterogeneous original data; s2, carrying out data fusion; s3, constructing a whole-process scene chain of the power grid; s4, triggering a graded early warning response; according to the method, the adaptive fusion algorithm is used in the feature layer fusion stage, so that feature deviation caused by single-dimensional fusion is avoided, more accurate basic data support is provided for subsequent risk scene recognition and correlation analysis, and the initial accuracy of overall risk early warning is improved; a coupling correlation algorithm is adopted, so that calculation of correlation strength between nodes is not limited to a static statistical rule any more, but changes of key node characteristics and real-time operation conditions of a power grid can be dynamically responded; a dynamic early warning threshold algorithm is adopted, the neglect of the traditional threshold design which only considers working conditions and risk accumulation on the influence of new energy characteristics and equipment aging is made up, and the scene adaptability of the early warning threshold is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of power grid risk early warning technology, specifically a method for risk early warning of the entire power grid scenario chain by fusing multi-source heterogeneous data. Background Technology

[0002] As power systems evolve towards higher penetration rates of renewable energy, greater equipment complexity, and full lifecycle management, power grid operation scenarios are becoming increasingly complex. On the one hand, the large-scale grid connection of renewable energy sources such as photovoltaics and wind power disrupts the traditional power balance stability of the power grid due to the randomness and volatility of their output, introducing new risks such as sudden drops in renewable energy output and curtailment of solar and wind power. On the other hand, the number of power grid equipment is surging and covers the entire lifecycle from planning, construction, operation, maintenance, to decommissioning. Risks at different stages exhibit cascading propagation characteristics, placing higher demands on the full-cycle, precise, and dynamic nature of risk early warning. Furthermore, power grid data exhibits multi-source heterogeneous characteristics, encompassing SCADA real-time data, equipment status monitoring data, environmental interference data, and renewable energy output data. How to effectively integrate these data and transform them into effective support for risk early warning has become a core technical requirement for ensuring the safe and stable operation of the power grid. Traditional risk early warning technologies have certain shortcomings in addressing these issues.

[0003] Firstly, in the field of multi-source heterogeneous data fusion in power grids, existing technologies mostly focus on designing fusion algorithms based on two dimensions: temporal correlation and spatial topology. For example, they use covariance analysis to analyze the temporal correlation of data and calculate spatial weights based on the geographical location of equipment to integrate different data sources. However, most of these technologies neglect the impact of equipment health status on data reliability. Furthermore, environmental interference has a significant impact on data acquisition, and traditional fusion algorithms do not differentiate weight allocation for data under different levels of environmental interference, which can easily lead to a decrease in the accuracy of the fusion results. These limitations make it difficult for the core features after fusion to accurately reflect the actual operating status of the power grid, thus creating potential accuracy risks for subsequent risk identification and early warning.

[0004] Secondly, existing technologies for constructing power grid risk scenario chains are mostly based on statistical analysis of historical fault data. They determine the correlation between nodes and form scenario chains by calculating the propagation probability and propagation time between risk scenario nodes. However, they ignore the differences in node importance. Key nodes in the power grid, such as hub substations and renewable energy grid connection points, have high load proportions and dense topological connections. Their risk propagation range and impact are far greater than those of ordinary nodes. Traditional methods do not quantify the impact of node importance on correlation strength, and are prone to missing risk propagation paths dominated by key nodes. Furthermore, they do not dynamically adjust correlation relationships in conjunction with real-time operating conditions: changes in real-time power grid load rate, renewable energy output fluctuations, and other operating conditions can accelerate or slow down risk propagation. Static correlation analysis cannot respond to changes in operating conditions, making it difficult for scenario chains to adapt to the real-time operating status of the power grid. The characterization of risk correlations in non-operational phases such as planning and construction is also one-sided, and it is impossible to form a complete risk propagation view covering the entire life cycle.

[0005] Finally, current power grid risk warning thresholds mostly adopt fixed thresholds or simple dynamic thresholds that only consider operating conditions and risk accumulation. Fixed thresholds are set based on power grid safety guidelines and cannot respond to changes in different operating scenarios. Although simple dynamic thresholds adjust the thresholds in combination with real-time load factor and risk duration, they do not adapt to the risk characteristics brought about by the increase in renewable energy penetration. Fluctuations in renewable energy output will disrupt the power balance of the power grid. Traditional thresholds do not include this factor in the adjustment dimension, which can easily lead to missed or false warnings of renewable energy-related risks. Moreover, the impact of equipment aging on thresholds is not considered. As the years of operation of power grid equipment increase, its risk resistance will decrease significantly. If the warning thresholds of new equipment are still used, the potential faults of aging equipment will not be identified in advance, and the best maintenance time will be missed. The rigidity of the above threshold design results in a high false alarm rate and missed alarm rate in the early warning system under scenarios such as high renewable energy penetration and full life cycle operation of equipment, which makes it difficult to meet the needs of full life cycle risk management of the power grid.

[0006] Therefore, it is essential to design a risk early warning method for the entire power grid scenario chain that integrates multi-source heterogeneous data. Summary of the Invention

[0007] The purpose of this invention is to provide a method for risk early warning of the entire power grid scenario chain through multi-source heterogeneous data fusion. This addresses the problems mentioned in the background section where traditional multi-source heterogeneous data fusion technologies focus solely on spatiotemporal dimensions in algorithm design, neglecting the impact of equipment health status on data reliability and environmental interference such as strong winds and icing on data acquisition accuracy. This results in the core features of the fused data deviating from the actual operating state of the power grid, failing to provide accurate data support for risk early warning. Furthermore, existing power grid scenario chain construction relies on static statistical analysis of historical fault data, determining node relationships solely through risk propagation probability and propagation time, without quantifying node importance indicators such as load share and topological centrality of hub nodes. The current risk warning threshold design is rigid, with fixed thresholds or simple dynamic thresholds that only consider the accumulation of operating conditions and risks. This fails to adapt to the power output fluctuations caused by the increasing penetration of new energy sources and does not quantify the impact of aging by combining the ratio of equipment operating years to design life. As a result, new energy-related risks are easily missed in early warning, and potential failures of aging equipment are difficult to identify in advance. The false alarm rate and false alarm rate are high, making it difficult to meet the needs of full-cycle risk management of the power grid.

[0008] To achieve the above objectives, the present invention provides the following technical solution: a method for risk early warning of the entire power grid scenario chain through multi-source heterogeneous data fusion, comprising the following steps:

[0009] S1: Collect multi-source heterogeneous raw data of the power grid, and preprocess the multi-source heterogeneous raw data by missing value imputation, noise filtering and dimension unification to obtain standardized processed data;

[0010] S2: The standardized processed data is fused using a three-layer fusion architecture consisting of a data layer, a feature layer, and a decision layer to extract the core features of the power grid operation status and generate a unified feature dataset;

[0011] S3: Based on the unified feature dataset, risk scenario nodes are divided into each stage of the power grid's entire life cycle. Coupled correlation algorithm is used to mine the correlation between nodes and construct a power grid full-process scenario chain.

[0012] S4: Based on the power grid full-process scenario chain, the real-time risk value of each node is calculated through the risk warning model, and the risk value is compared with the dynamic warning threshold to trigger a graded warning response, thereby realizing full-cycle control of power grid risks.

[0013] As a further technical solution of the present invention, in S1, the multi-source heterogeneous raw data of the power grid specifically includes:

[0014] Real-time operating data: voltage, current, and power data of the SCADA system, and PMU synchronization phasor data;

[0015] Equipment status data: dissolved gas concentration in transformer oil, infrared temperature measurement data of the line, partial discharge of switchgear, and equipment health index; among which, the equipment health index is calculated based on the dissolved gas concentration in oil and partial discharge, and the value range is [0,1]. The closer the value is to 1, the better the equipment health status.

[0016] Environmental interference data: regional wind speed, rainfall, ice thickness, ambient temperature, and environmental interference level; among which, regional wind speed, rainfall, ice thickness, and ambient temperature are divided into high interference and low interference to assess the reliability of data collection.

[0017] Management support data: daily electricity load curve, photovoltaic or wind power output data, grid topology parameters, and output prediction curves of new energy grid-connected nodes.

[0018] As a further technical solution of the present invention, the specific process of preprocessing multi-source heterogeneous raw data in S1 is as follows:

[0019] S11: Missing value imputation:

[0020] If the missing rate is ≤5%, linear interpolation is used, and the formula is:

[0021]

[0022] Where, x t The missing value at time t, t t For the corresponding timestamp, x t-1 x t+1 The values ​​are valid values ​​at adjacent time points;

[0023] If the missing rate is >5%, the K-nearest neighbor imputation method is used, where K = 3-5. Similar samples are selected based on Euclidean distance, and similar samples must meet the requirement of consistent environmental interference levels to avoid imputation deviation caused by environmental differences.

[0024] S12: Noise Filtering

[0025] Wavelet transform algorithm is used for basic denoising, with the db4 wavelet basis function selected for three-level decomposition. Level 1 decomposition separates high-frequency random noise, and levels 2-3 decomposition separate low-frequency interference generated by equipment operation. Matching the characteristics of the power grid data signal, soft thresholding is applied to the decomposed high-frequency coefficients. The threshold calculation formula is as follows:

[0026]

[0027] Where σ is the noise standard deviation and N is the data length; random noise is removed after reconstruction; Kalman filtering is added to the new energy output data, where the process noise covariance Q = 0.01 and the observation noise covariance R = 0.05 are set. Through the iterative process from state prediction to observation update, the output fluctuation noise is further smoothed to ensure that the noise level of new energy data is consistent with that of other power grid data and to suppress the systematic noise caused by output fluctuation.

[0028] S13: Dimensional uniformity:

[0029] Using Min-Max normalization, the formula is as follows:

[0030]

[0031] Where x represents the multi-source heterogeneous original data, x max x min These represent the maximum and minimum values ​​in the data sample, respectively; the output is standardized data with values ​​ranging from [0,1] to eliminate the influence of dimensional differences on the fusion result;

[0032] By matching the environmental interference level during missing value filling, the filling bias caused by environmental differences is effectively avoided, improving the accuracy of missing data completion. Kalman filtering is additionally applied to the new energy output data to specifically smooth the output fluctuation noise, ensuring that the noise level of new energy data is consistent with that of other power grid data and eliminating systematic interference. Min-Max normalization completely eliminates the difference in dimensions, enabling multi-source heterogeneous data to have a unified comparison and fusion benchmark, laying a high-quality data foundation for the efficient and accurate operation of the subsequent three-layer fusion architecture.

[0033] As a further technical solution of the present invention, the three-layer fusion architecture in S2 includes:

[0034] 1) Data layer fusion: For sensor data of the same type, a weighted average method is used.

[0035]

[0036] Among them, w data,i For the weight, σ i Let σ be the standard deviation of the data from the i-th sensor. avg The average standard deviation of all data from the same type of sensor is used. The smaller the standard deviation, the higher the weight. Stable data is given priority.

[0037] Additional prediction accuracy weights are introduced for new energy output data:

[0038]

[0039] Where, φ iThe prediction accuracy of the i-th group of new energy output data is defined as follows, with a value range of [0,1], and the final weight is w. data,i ×w p Strengthen the contribution of high-precision prediction data;

[0040] 2) Feature layer fusion: An adaptive fusion algorithm is adopted to achieve effective integration of heterogeneous features through multi-dimensional weight allocation;

[0041] 3) Decision-level fusion: Preliminary risk results from various data sources are synthesized using the DS evidence theory.

[0042] 31) Divide the data source into 4 categories of evidence E = {E1, E2, E3, E4}, where:

[0043] Device status class E1: Based on the device health feature vector F output by the feature layer h The initial risk probabilities P1(θ1) and P1(θ2) of normal equipment (θ1) and abnormal equipment (θ2) are generated by the SVM classifier, where the SVM kernel function is RBF, the penalty coefficient C = 10, and the gamma parameter = 0.1.

[0044] Real-time running class E2: Based on real-time running feature vector F r Calculate the deviation rate of voltage, current, and power from the rated values. The probability of mapping to normal operation (θ1) and operation exceeding the limit (θ2) is P2(θ1) = 1 - min(δ / 0.2,1) and P2(θ2) = min(δ / 0.2,1), where P2(θ2) = 1 when the deviation rate exceeds 20%.

[0045] New Energy Category E3: Based on New Energy Feature Vector F new Combined with power output volatility Based on the inverter status parameters, the probabilities of generating normal (θ1) and abnormal (θ2) renewable energy output are P3(θ1) = 1 - (min(ΔP / 0.15,1)×0.7 + inverter abnormality coefficient×0.3) and P3(θ2) = 1 - P3(θ1), where the inverter abnormality coefficient is calculated from the capacitor temperature and IGBT on-state voltage drop, and takes a value of 0-1;

[0046] Environment class E4: Based on environment feature vector F e The probability mappings based on environmental interference levels are as follows: for high interference, P4(θ1) = 0.3 and P4(θ2) = 0.7; for low interference, P4(θ1) = 0.8 and P4(θ2) = 0.2.

[0047] 32) For each type of evidence, introduce the uncertainty proposition Θ = {θ1, θ2}, where θ1 is "risk-free", θ2 is "risky", and the BPA function mi (A) The definition is as follows:

[0048] Single-point proposition:

[0049] m i ({θ1})=P i (θ1)×λ i

[0050] m i ({θ2})=P i (θ2)×λ i

[0051] Where, λ i The credibility of the evidence is calculated based on the signal-to-noise ratio after data fusion at the data layer. Values ​​range from 0.6 to 0.95;

[0052] Uncertainty propositions:

[0053] m i (Θ)=1-m i ({θ1})-m i ({θ2})

[0054] This reflects the uncertainty caused by the lack of evidence, such as ambiguity due to incomplete data sampling;

[0055] 33) Calculate the conflict coefficient K between pieces of evidence to determine the degree of conflict:

[0056]

[0057] If K ≤ 0.3 indicates low conflict: directly apply the DS composition rule, that is, for any non-empty proposition... The synthesized BPA is:

[0058]

[0059] If K > 0.3 indicates high conflict: introduce an evidence discount factor α. i =1-0.2×K, the greater the conflict, the stronger the discount, and the BPA of each piece of evidence is adjusted to m. i ′({θ1})=m i ({θ1})×α i m i ′({θ2})=m i ({θ2})×α i m i ′(Θ)=1-m i ′({θ1})-m i Then perform the above DS synthesis again;

[0060] 34) Based on the synthesized BPA, calculate the trust function Bel(θ1) = m({θ1}) and the likelihood function Pl(θ1) = 1 - m({θ2}) for the risk-free θ1, and output the comprehensive risk result according to the following rules:

[0061] If Bel(θ2)≥0.7: it is judged as high risk, triggering early warning preparation;

[0062] If 0.4 ≤ Bel(θ2) < 0.7: it is judged as medium risk, and a risk warning is output;

[0063] If Bel(θ2) < 0.4: it is judged as low risk, and only the risk trend is recorded;

[0064] If Pl(θ1)-Bel(θ1)>0.5: it is judged as high uncertainty: automatically increase the sampling frequency of the corresponding data source, re-collect data and make a second fusion decision;

[0065] Data sources are categorized into four types of evidence: equipment, operation, new energy, and environment, achieving comprehensive coverage of multi-dimensional risks and avoiding the one-sidedness of assessments based on a single data source; the credibility of the evidence is assessed using λ. i The BPA function is modified so that data reliability directly affects risk assessment weights, thus improving the objectivity of the assessment; a discount factor α is introduced for highly conflicting evidence. i It solves the problem of distortion in synthesis results of traditional DS evidence theory in high-conflict scenarios; it automatically increases the sampling frequency and performs secondary fusion under high uncertainty, effectively avoiding decision-making errors caused by fuzzy data. It is especially suitable for the complex needs of power grid risk assessment after the increase in the penetration rate of new energy, and significantly improves the accuracy and reliability of risk decision-making.

[0066] As a further technical solution of the present invention, the adaptive fusion algorithm used in the feature layer fusion is specifically as follows:

[0067] Let the feature vector of the i-th class of standardized data be x. i Let i = 1, 2, ..., n, where n is the number of data source types, and the core feature vector after fusion be F. Then:

[0068]

[0069] Wherein, weight w i By time weight w t,i Spatial weight w s,i Equipment health weight w h,i Environmental interference weight w e,i Adaptive calculation, the formula is:

[0070] w i =α·w t,i +β·w s,i +γ·wh,i +(1-α-β-γ)·w e,i

[0071] Where α is the time weight coefficient, β is the spatial weight coefficient, and γ is the health weight coefficient, satisfying 0.2≤α≤0.4, 0.2≤β≤0.4, 0.1≤γ≤0.3, and α+β+γ≤1. Grid search optimization is performed with the goal of maximizing the Pearson comprehensive correlation between the fused features and the grid fault label and the new energy output anomaly label. The specific calculation methods for each weight are as follows:

[0072] Time weight w t,i The formula reflects the temporal correlation between data and the real-time status of the power grid:

[0073]

[0074] Where y is the real-time frequency / voltage vector of the power grid, p is the real-time power output vector of new energy sources, and Cov is the covariance. The larger the absolute value of the covariance, the stronger the time-series correlation and the higher the weight.

[0075] Spatial weight w s,i The formula reflects the impact of spatial topological associations of devices on data importance:

[0076]

[0077] Where, d i This represents the shortest topological distance from the corresponding data device to the hub substation; 0.1 is a correction term to avoid a denominator of 0. i S represents the load capacity of the node where the equipment is located. max The maximum load capacity of a node in the entire network is used; the closer the node is, the greater its load capacity, and the higher its weight.

[0078] Device health weight w h,i The formula reflects the impact of equipment health status on data reliability:

[0079]

[0080] Among them, HI i The health index of the device corresponding to the i-th data type is calculated as follows: c k,i c represents the actual concentration of dissolved gases in the oil. k,max The higher the health index, the stronger the data reliability and the higher the weight, corresponding to the safety threshold of the gas.

[0081] Environmental interference weight w e,i The formula reflects the impact of environmental interference on the reliability of data acquisition:

[0082]

[0083] Sensor data is prone to distortion in high-interference environments, so the weighting is reduced; data reliability is high in low-interference environments, so the weighting is increased.

[0084] As a further technical solution of the present invention, in S3, the nodes of each stage of the power grid's entire life cycle and the corresponding risk scenarios include:

[0085] Planning phase: Nodes with insufficient topology redundancy, nodes with equipment selection errors, and nodes with insufficient grid-connected capacity of new energy sources; the corresponding risks are: power flow blockage, overload burnout, and curtailment of solar or wind power.

[0086] Construction phase: Key points include substandard construction techniques, installation deviations, and incorrect selection of new energy access cables; the corresponding risks are: line short circuit, partial discharge, and cable overheating, respectively.

[0087] Operational phases include: overload nodes, voltage sag nodes, frequency offset nodes, and sudden drop in renewable energy output nodes; the corresponding risks are: equipment overheating, sensitive load outages, grid instability, and power deficit, respectively.

[0088] Maintenance phase: Overlooked maintenance points, delayed maintenance points, and missing maintenance points for new energy inverters; the corresponding risks are: fault expansion, accelerated equipment aging, and inverter failure, respectively.

[0089] The decommissioning phase includes addressing safety hazards, asset waste, and oversights in the environmental treatment of decommissioned equipment; the corresponding risks are personal injury, resource depletion, and environmental pollution.

[0090] As a further technical solution of the present invention, the specific steps in S3 for constructing the power grid full-process scenario chain are as follows:

[0091] S31: Node Initialization: Use the risk scenario nodes as the initial node set V = {v1, v2, ..., v...} m}, where m is the total number of nodes, and each node is associated with a corresponding risk type, scope of impact, and characteristic indicators;

[0092] S32: Association Mining: Using a coupled association algorithm, the association strength between any two nodes is calculated to identify risk propagation paths;

[0093] S33: Chain network generation: When the correlation strength is ≥0.65, a directed edge is established between nodes, with the arrow direction indicating the risk propagation direction; high-priority correlation edges are additionally marked for new energy-related nodes, and an additional 10% correction coefficient is added when calculating the correlation strength, ultimately forming a full-process scenario chain for the power grid, including planning, construction, operation, maintenance, and decommissioning.

[0094] As a further technical solution of the present invention, the coupling correlation algorithm is specifically as follows:

[0095] Let the preceding node be the risk propagation starting point v. a To subsequent nodes: Risk propagation endpoint v b The correlation strength is R a→b ,but:

[0096]

[0097] Wherein, β1 is the probability weight coefficient, β2 is the time weight coefficient, and β3 is the importance and operating condition coupling weight coefficient, satisfying 0.3≤β1≤0.5, 0.2≤β2≤0.4, 0.1≤β3≤0.3, and β1+β2+β3=1, calibrated through cascading failure cases; the specific calculation methods for each parameter are as follows:

[0098] Conditional probability P(v) b |v a ): Reflects v a After it happened v b The trigger probability is given by the formula:

[0099]

[0100] Where, N a∩b For v a With v b The number of co-occurring samples, N a For v a The total number of samples, m is the total number of nodes, +1 is Laplace smoothing to avoid a probability of 0; N a∩b∩p For v a v b The number of co-occurrence samples of abnormal power output events with new energy sources, N a∩p For v a The number of co-occurrence samples with abnormal power output from new energy sources strengthens the correlation probability of new energy-related nodes;

[0101] Propagation time T a→b : Reflects v a Risk spread to v b The average time is calculated based on historical fault time-series data; for nodes related to sudden drops in new energy output, an additional correction for output fluctuation response time is introduced, and the correction formula is:

[0102] T a→b ′=T a→b ×(1+0.2×ΔP)

[0103] Wherein, ΔP is the power output volatility of new energy sources, with a value range of [0,1]. The greater the volatility, the longer the response time.

[0104] Maximum tolerance time T max The maximum tolerance time for risk propagation in the power grid is adjusted according to the penetration rate of new energy sources.

[0105] Node Importance I a : Reflects the preceding node v a The influence of the power grid's position on the correlation strength is expressed by the following formula:

[0106]

[0107] Among them, L a For v a The proportion of the load of the node to the total network load, L max The highest percentage of load across the entire network; C a For v a Degree centrality is the ratio of the number of other nodes connected to a node to the total number of nodes in the network. The value ranges from [0,1]. The higher the load ratio and the denser the topology connections, the higher the importance.

[0108] Real-time operating condition factor G a : Reflects v a The formula for the accelerating effect of the real-time operating conditions of the node on risk propagation is:

[0109]

[0110] Where, ρ a For v a The real-time load factor of the node, with a value range of [0,1]; P a,new For v a The associated renewable energy nodes are generating power in real time, P a,new,max The rated output of the new energy node is defined as [0,1]. The higher the load factor, the closer the new energy output is to the rated value, the more intense the operating conditions, the faster the risk propagation, and the higher the factor value.

[0111] As a further technical solution of the present invention, in S4, the risk warning model adopts a dynamic warning threshold algorithm, specifically as follows:

[0112] Let the real-time risk value of the k-th node in the scenario chain be R. k :

[0113]

[0114] Where w is the feature weight vector, b is the bias term, and the training samples cover power grid faults and new energy anomaly cases;

[0115] The dynamic early warning threshold is Th k :

[0116] Th k =Th0k ·[1+γ·D k +η·C k +θ·ΔP k +μ·A k ]

[0117] Among them, Th 0k Based on the basic threshold, the definitions and calculation methods of the remaining parameters are as follows:

[0118] Operating condition influence coefficient γ: reflects the adjustment range of the threshold due to real-time operating conditions, with a value range of 0.2-0.5, and adjusts linearly with the node load rate ρ. The formula is:

[0119] γ = 0.2 + 0.3ρ

[0120] Where ρ∈[0,1], the higher the load rate, the more intense the working conditions, the larger the coefficient, and the higher the threshold to avoid false alarms;

[0121] Operating condition deviation D k This reflects the degree of deviation between the current operating condition and the rated operating condition; the formula is:

[0122]

[0123] Where, x ck This is the current operating condition vector of the node, containing voltage, current, and power; x 0k x is the node's rated operating condition vector; ck,new x is the current output vector of the new energy node; 0k,new The rated output vector of the new energy node; ||·|| is the Euclidean distance, the greater the deviation, the greater the threshold adjustment range;

[0124] Risk accumulation coefficient η: reflects the impact of risk duration on the threshold, with a value ranging from 0.1 to 0.3, adjusted according to the node risk duration t, and the formula is:

[0125] η = 0.1 + 0.2 × min(t / 60, 1)

[0126] Among them, the t of the new energy-related nodes needs to be superimposed with the duration of the power output fluctuation, that is, t = t 风险 +t 波动 The longer the duration, the larger the coefficient, and the lower the threshold to trigger an alert;

[0127] New energy output fluctuation factor θ: reflects the impact of new energy output fluctuation on the threshold, with a value range of 0.1-0.3, and varies with the new energy output fluctuation rate ΔP. k Adjust the formula as follows:

[0128] θ = 0.1 + 0.2 × min(ΔP) k / 0.15,1)

[0129] in, The value range is [0,1]. The greater the volatility, the larger the factor and the lower the threshold.

[0130] Equipment aging factor μ: reflects the impact of equipment aging degree on the threshold, with a value range of 0.1-0.25, varying with the equipment's operating years Y. k With design life Y k,设计 The ratio adjustment is calculated using the following formula:

[0131] μ = 0.1 + 0.15 × min(Y) k / Y k,设计 ,1)

[0132] Among them, the larger the ratio, the more severe the aging; the larger the factor, the lower the threshold.

[0133] Warning triggering and classification rules: When R k ≥Th k When an alert is triggered, press The risk levels are categorized as follows: 0-0.2 (low level), 0.2-0.5 (medium level), and >0.5 (high level). The early warning level for new energy-related nodes is raised by one level, while the high level remains unchanged, thus strengthening the response priority for new energy risks.

[0134] The dynamic early warning threshold incorporates four factors: operating conditions, risk accumulation, fluctuations in new energy output, and equipment aging. The threshold adjustment is linked in real time with the actual operating status of the power grid. Each factor is dynamically quantified and adjusted according to actual operating parameters, so that the threshold can accurately match the risk perception needs under different scenarios.

[0135] As a further technical solution of the present invention, the specific measures for the graded early warning response in S4 are as follows:

[0136] Low-level warning: Generate equipment tracking instructions, increase the sampling frequency of conventional sensors and the monitoring frequency of new energy output, and upload data to the power grid monitoring platform in real time. Maintenance personnel need to check the tracking data every hour.

[0137] Medium-level early warning: Activate the joint plan for load transfer and new energy power output adjustment, use the Newton-Raphson method to calculate the power flow of the grid and screen the load transfer path; simultaneously calculate the adjustment amount of new energy power output, push the plan to the grid dispatch terminal and the new energy power station control terminal, and dispatch personnel complete the execution of the plan;

[0138] High-level early warning: Automatically triggers emergency control measures, links circuit breakers to cut off power to the fault area, and additionally links inverters at new energy grid-connected nodes to cut off power in an emergency; at the same time, it generates a maintenance work order, including GPS coordinates of the fault location, equipment model, recommended tool list, and temporary alternative power output plan for new energy, pushes it to the mobile terminal of maintenance personnel, and simultaneously reports it to the power grid supervision platform and new energy supervision department, and maintenance personnel arrive at the site to handle the situation.

[0139] In low-level early warnings, the monitoring frequency of renewable energy output is increased, enabling real-time tracking of renewable energy output fluctuation risks and providing data support for early risk intervention. In medium-level early warnings, load transfer and renewable energy output adjustment are combined to achieve synergy between traditional grid control and renewable energy management, avoiding the problem that single control cannot balance grid supply and demand, and ensuring stable grid operation. In high-level early warnings, renewable energy grid-connected nodes are linked to the emergency disconnection of inverters, which can quickly isolate renewable energy side faults and prevent risk spread. Maintenance work orders include temporary alternative solutions for renewable energy output, ensuring grid power supply stability during fault handling and fully adapting to the graded management and control needs of grid risks after the increase in renewable energy penetration.

[0140] Compared with existing technologies, the beneficial effects of this multi-source heterogeneous data fusion method for risk early warning across the entire power grid process are:

[0141] By employing an adaptive fusion algorithm in the feature layer fusion stage, this approach overcomes the limitations of traditional fusion methods that rely solely on spatiotemporal dimensions. It incorporates equipment health status and environmental interference factors into the weight allocation system, effectively addressing issues such as uneven data reliability due to differences in equipment health and data distortion caused by environmental interference during power grid data acquisition. The algorithm dynamically adjusts the weights of corresponding data based on the equipment health index, prioritizing the use of data from equipment with good health status. Simultaneously, it adjusts data contribution by incorporating environmental interference levels, ensuring that the fused core features better reflect the actual operating state of the power grid. This avoids feature bias caused by single-dimensional fusion, providing more accurate basic data support for subsequent risk scenario identification and correlation analysis, and improving the initial accuracy of overall risk warning.

[0142] In the construction of the power grid full-process scenario chain, a coupled correlation algorithm is adopted, adding the coupled calculation of node importance and real-time operating conditions. By quantifying the load ratio and topological centrality of hub nodes to reflect node importance, and combining real-time load rate and renewable energy output status to characterize the impact of operating conditions, the calculation of the correlation strength between nodes is no longer limited to static statistical laws, but can dynamically respond to changes in the characteristics of key power grid nodes and real-time operating conditions. This effectively avoids the omission or misjudgment of key risk paths by static correlation analysis, ensuring that the constructed full-process scenario chain can accurately identify the core risk propagation paths at each stage from planning to decommissioning, and providing a more targeted scenario carrier for the full-cycle management of power grid risks.

[0143] In designing risk warning thresholds, a dynamic warning threshold algorithm is adopted to compensate for the neglect of the characteristics of new energy sources and the impact of equipment aging by traditional threshold designs that only consider operating conditions and risk accumulation. This significantly improves the scenario adaptability of the warning thresholds. Incorporating new energy output fluctuations into the threshold adjustment dimension adapts to the special characteristics of power grid operation under high new energy penetration rates, preventing new risks such as sudden drops in new energy output from being missed due to rigid thresholds. Simultaneously, the aging impact is quantified by combining the ratio of equipment operating years to design life, making the warning thresholds for aging equipment more closely match its actual risk resistance capabilities and identifying potential failure risks of aging equipment in advance. This multi-dimensional adjustment mechanism of dynamic thresholds effectively reduces false alarms and missed alarms of fixed thresholds under different operating conditions, new energy fluctuations, and equipment states. This makes the warning response more aligned with the operating characteristics of the entire power grid lifecycle, providing more accurate risk warning support for the safe and stable operation of the power grid. Attached Figure Description

[0144] Figure 1 This is a schematic diagram of the method flow of the present invention. Detailed Implementation

[0145] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0146] Please see the appendix Figure 1 Embodiment 1 of the present invention provides a method for risk early warning of the entire power grid scenario chain through multi-source heterogeneous data fusion, comprising the following steps:

[0147] S1: Collect multi-source heterogeneous raw data of the power grid, and preprocess the multi-source heterogeneous raw data by missing value imputation, noise filtering and dimension unification to obtain standardized processed data;

[0148] The specific raw data of the multi-source heterogeneous power grid includes:

[0149] Real-time operating data: voltage, current, and power data of the SCADA system, and PMU synchronization phasor data;

[0150] Equipment status data: dissolved gas concentration in transformer oil, infrared temperature measurement data of the line, partial discharge of switchgear, and equipment health index; among which, the equipment health index is calculated based on the dissolved gas concentration in oil and partial discharge, and the value range is [0,1]. The closer the value is to 1, the better the equipment health status.

[0151] Environmental interference data: regional wind speed, rainfall, ice thickness, ambient temperature, and environmental interference level; among which, regional wind speed, rainfall, ice thickness, and ambient temperature are divided into high interference and low interference to assess the reliability of data collection.

[0152] Management support data: daily electricity load curve, photovoltaic or wind power output data, grid topology parameters, and output prediction curves of new energy grid-connected nodes;

[0153] The specific process for preprocessing multi-source heterogeneous raw data is as follows:

[0154] S11: Missing value imputation:

[0155] If the missing rate is ≤5%, linear interpolation is used, and the formula is:

[0156]

[0157] Where, x t The missing value at time t, t t For the corresponding timestamp, x t-1 x t+1 The values ​​are valid values ​​at adjacent time points;

[0158] If the missing rate is >5%, the K-nearest neighbor imputation method is used, where K = 3-5. Similar samples are selected based on Euclidean distance, and similar samples must meet the requirement of consistent environmental interference levels to avoid imputation deviation caused by environmental differences.

[0159] S12: Noise Filtering

[0160] Wavelet transform algorithm is used for basic denoising, with the db4 wavelet basis function selected for three-level decomposition. Level 1 decomposition separates high-frequency random noise, and levels 2-3 decomposition separate low-frequency interference generated by equipment operation. Matching the characteristics of the power grid data signal, soft thresholding is applied to the decomposed high-frequency coefficients. The threshold calculation formula is as follows:

[0161]

[0162] Where σ is the noise standard deviation and N is the data length; random noise is removed after reconstruction; Kalman filtering is added to the new energy output data, where the process noise covariance Q = 0.01 and the observation noise covariance R = 0.05 are set. Through the iterative process from state prediction to observation update, the output fluctuation noise is further smoothed to ensure that the noise level of new energy data is consistent with that of other power grid data and to suppress the systematic noise caused by output fluctuation.

[0163] S13: Dimensional uniformity:

[0164] Using Min-Max normalization, the formula is as follows:

[0165]

[0166] Where x represents the multi-source heterogeneous original data, x max x min These represent the maximum and minimum values ​​in the data sample, respectively; the output is standardized data with values ​​ranging from [0,1] to eliminate the influence of dimensional differences on the fusion result;

[0167] S2: A three-layer fusion architecture consisting of a data layer, a feature layer, and a decision layer is adopted to fuse standardized processed data, extract the core features of the power grid operation status, and generate a unified feature dataset;

[0168] The three-tier converged architecture includes:

[0169] 1) Data layer fusion: For sensor data of the same type, a weighted average method is used.

[0170]

[0171] Among them, w data,i For the weight, σ i Let σ be the standard deviation of the data from the i-th sensor. avg The average standard deviation of all data from the same type of sensor is used. The smaller the standard deviation, the higher the weight. Stable data is given priority.

[0172] Additional prediction accuracy weights are introduced for new energy output data:

[0173]

[0174] Where, φ i The prediction accuracy of the i-th group of new energy output data is defined as follows, with a value range of [0,1], and the final weight is w. data,i ×w p Strengthen the contribution of high-precision prediction data;

[0175] 2) Feature Layer Fusion: An adaptive fusion algorithm is adopted to effectively integrate heterogeneous features through multi-dimensional weight allocation. The adaptive fusion algorithm is as follows:

[0176] Let the feature vector of the i-th class of standardized data be x. i Let i = 1, 2, ..., n, where n is the number of data source types, and the core feature vector after fusion be F. Then:

[0177]

[0178] Wherein, weight w i By time weight w t,i Spatial weight w s,i Equipment health weight w h,i Environmental interference weight we,i Adaptive calculation, the formula is:

[0179] w i =α·w t,i +β·w s,i +γ·w h,i +(1-α-β-γ)·w e,i

[0180] Where α is the time weight coefficient, β is the spatial weight coefficient, and γ is the health weight coefficient, satisfying 0.2≤α≤0.4, 0.2≤β≤0.4, 0.1≤γ≤0.3, and α+β+γ≤1. Grid search optimization is performed with the goal of maximizing the Pearson comprehensive correlation between the fused features and the grid fault label and the new energy output anomaly label. The specific calculation methods for each weight are as follows:

[0181] Time weight w t,i The formula reflects the temporal correlation between data and the real-time status of the power grid:

[0182]

[0183] Where y is the real-time frequency / voltage vector of the power grid, p is the real-time power output vector of new energy sources, and Cov is the covariance. The larger the absolute value of the covariance, the stronger the time-series correlation and the higher the weight.

[0184] Spatial weight w s,i The formula reflects the impact of spatial topological associations of devices on data importance:

[0185]

[0186] Where, d i This represents the shortest topological distance from the corresponding data device to the hub substation; 0.1 is a correction term to avoid a denominator of 0. i S represents the load capacity of the node where the equipment is located. max The maximum load capacity of a node in the entire network is used; the closer the node is, the greater its load capacity, and the higher its weight.

[0187] Device health weight w h,i The formula reflects the impact of equipment health status on data reliability:

[0188]

[0189] Among them, HI i The health index of the device corresponding to the i-th data type is calculated as follows: c k,i c represents the actual concentration of dissolved gases in the oil. k,maxThe higher the health index, the stronger the data reliability and the higher the weight, corresponding to the safety threshold of the gas.

[0190] Environmental interference weight w e,i The formula reflects the impact of environmental interference on the reliability of data acquisition:

[0191]

[0192] Sensor data is prone to distortion in high-interference environments, so the weight is reduced; data reliability is high in low-interference environments, so the weight is increased.

[0193] 3) Decision-level fusion: Preliminary risk results from various data sources are synthesized using the DS evidence theory.

[0194] 31) Divide the data source into 4 categories of evidence E = {E1, E2, E3, E4}, where:

[0195] Device status class E1: Based on the device health feature vector F output by the feature layer h The initial risk probabilities P1(θ1) and P1(θ2) of normal equipment (θ1) and abnormal equipment (θ2) are generated by the SVM classifier, where the SVM kernel function is RBF, the penalty coefficient C = 10, and the gamma parameter = 0.1.

[0196] Real-time running class E2: Based on real-time running feature vector F r Calculate the deviation rate of voltage, current, and power from the rated values. The probability of mapping to normal operation (θ1) and operation exceeding the limit (θ2) is P2(θ1) = 1 - min(δ / 0.2,1) and P2(θ2) = min(δ / 0.2,1), where P2(θ2) = 1 when the deviation rate exceeds 20%.

[0197] New Energy Category E3: Based on New Energy Feature Vector F new Combined with power output volatility Based on the inverter status parameters, the probabilities of generating normal (θ1) and abnormal (θ2) renewable energy output are P3(θ1) = 1 - (min(ΔP / 0.15,1)×0.7 + inverter abnormality coefficient×0.3) and P3(θ2) = 1 - P3(θ1), where the inverter abnormality coefficient is calculated from the capacitor temperature and IGBT on-state voltage drop, and takes a value of 0-1;

[0198] Environment class E4: Based on environment feature vector F e The probability mappings based on environmental interference levels are as follows: for high interference, P4(θ1) = 0.3 and P4(θ2) = 0.7; for low interference, P4(θ1) = 0.8 and P4(θ2) = 0.2.

[0199] 32) For each type of evidence, introduce the uncertainty proposition Θ = {θ1, θ2}, where θ1 is "risk-free", θ2 is "risky", and the BPA function m i (A) The definition is as follows:

[0200] Single-point proposition:

[0201] m i ({θ1})=P i (θ1)×λ i

[0202] m i ({θ2})=P i (θ2)×λ i

[0203] Where, λ i The credibility of the evidence is calculated based on the signal-to-noise ratio after data fusion at the data layer. Values ​​range from 0.6 to 0.95;

[0204] Uncertainty propositions:

[0205] m i (Θ)=1-m i ({θ1})-m i ({θ2})

[0206] This reflects the uncertainty caused by the lack of evidence, such as ambiguity due to incomplete data sampling;

[0207] 33) Calculate the conflict coefficient K between pieces of evidence to determine the degree of conflict:

[0208]

[0209] If K ≤ 0.3 indicates low conflict: directly apply the DS composition rule, that is, for any non-empty proposition... The synthesized BPA is:

[0210]

[0211] If K > 0.3 indicates high conflict: introduce an evidence discount factor α. i =1-0.2×K, the greater the conflict, the stronger the discount, and the BPA of each piece of evidence is adjusted to m. i ′({θ1})=m i ({θ1})×α i m i ′({θ2})=m i ({θ2})×α i m i ′(Θ)=1-m i ′({θ1})-mi Then perform the above DS synthesis again;

[0212] 34) Based on the synthesized BPA, calculate the trust function Bel(θ1) = m({θ1}) and the likelihood function Pl(θ1) = 1 - m({θ2}) for the risk-free θ1, and output the comprehensive risk result according to the following rules:

[0213] If Bel(θ2)≥0.7: it is judged as high risk, triggering early warning preparation;

[0214] If 0.4 ≤ Bel(θ2) < 0.7: it is judged as medium risk, and a risk warning is output;

[0215] If Bel(θ2) < 0.4: it is judged as low risk, and only the risk trend is recorded;

[0216] If Pl(θ1)-Bel(θ1)>0.5: it is judged as high uncertainty: automatically increase the sampling frequency of the corresponding data source, re-collect data and make a second fusion decision;

[0217] S3: Based on a unified feature dataset, risk scenario nodes are divided into each stage of the power grid's entire life cycle. Coupled correlation algorithms are used to mine the correlation relationships between nodes and construct a power grid full-process scenario chain.

[0218] The stages of the power grid's entire lifecycle and their corresponding risk scenarios include:

[0219] Planning phase: Nodes with insufficient topology redundancy, nodes with equipment selection errors, and nodes with insufficient grid-connected capacity of new energy sources; the corresponding risks are: power flow blockage, overload burnout, and curtailment of solar or wind power.

[0220] Construction phase: Key points include substandard construction techniques, installation deviations, and incorrect selection of new energy access cables; the corresponding risks are: line short circuit, partial discharge, and cable overheating, respectively.

[0221] Operational phases include: overload nodes, voltage sag nodes, frequency offset nodes, and sudden drop in renewable energy output nodes; the corresponding risks are: equipment overheating, sensitive load outages, grid instability, and power deficit, respectively.

[0222] Maintenance phase: Overlooked maintenance points, delayed maintenance points, and missing maintenance points for new energy inverters; the corresponding risks are: fault expansion, accelerated equipment aging, and inverter failure, respectively.

[0223] Decommissioning phase: Dismantling safety hazards, asset waste, and oversights in the environmental treatment of decommissioned equipment; the corresponding risks are: personal injury, resource depletion, and environmental pollution, respectively.

[0224] The specific steps for constructing a full-process scenario chain for the power grid are as follows:

[0225] S31: Node Initialization: Use the risk scenario nodes as the initial node set V = {v1, v2, ..., v...} m}, where m is the total number of nodes, and each node is associated with a corresponding risk type, scope of impact, and characteristic indicators;

[0226] S32: Association Mining: A coupled association algorithm is used to calculate the association strength between any two nodes and identify risk propagation paths. The coupled association algorithm is as follows:

[0227] Let the preceding node be the risk propagation starting point v. a To subsequent nodes: Risk propagation endpoint v b The correlation strength is R a→b ,but:

[0228]

[0229] Wherein, β1 is the probability weight coefficient, β2 is the time weight coefficient, and β3 is the importance and operating condition coupling weight coefficient, satisfying 0.3≤β1≤0.5, 0.2≤β2≤0.4, 0.1≤β3≤0.3, and β1+β2+β3=1, calibrated through cascading failure cases; the specific calculation methods for each parameter are as follows:

[0230] Conditional probability P(v) b |v a ): Reflects v a After it happened v b The trigger probability is given by the formula:

[0231]

[0232] Where, N a∩b For v a With v b The number of co-occurring samples, N a For v a The total number of samples, m is the total number of nodes, +1 is Laplace smoothing to avoid a probability of 0; N a∩b∩p For v a v b The number of co-occurrence samples of abnormal power output events with new energy sources, N a∩p For v a The number of co-occurrence samples with abnormal power output from new energy sources strengthens the correlation probability of new energy-related nodes;

[0233] Propagation time T a→b : Reflects v a Risk spread to v bThe average time is calculated based on historical fault time-series data; for nodes related to sudden drops in new energy output, an additional correction for output fluctuation response time is introduced, and the correction formula is:

[0234] T a→b ′=T a→b ×(1+0.2×ΔP)

[0235] Wherein, ΔP is the power output volatility of new energy sources, with a value range of [0,1]. The greater the volatility, the longer the response time.

[0236] Maximum tolerance time T max The maximum tolerance time for risk propagation in the power grid is adjusted according to the penetration rate of new energy sources.

[0237] Node Importance I a : Reflects the preceding node v a The influence of the power grid's position on the correlation strength is expressed by the following formula:

[0238]

[0239] Among them, L a For v a The proportion of the load of the node to the total network load, L max The highest percentage of load across the entire network; C a For v a Degree centrality is the ratio of the number of other nodes connected to a node to the total number of nodes in the network. The value ranges from [0,1]. The higher the load ratio and the denser the topology connections, the higher the importance.

[0240] Real-time operating condition factor G a : Reflects v a The formula for the accelerating effect of the real-time operating conditions of the node on risk propagation is:

[0241]

[0242] Where, ρ a For v a The real-time load factor of the node, with a value range of [0,1]; P a,new For v a The associated renewable energy nodes are generating power in real time, P a,new,max The rated output of the new energy node is [0,1]. The higher the load factor and the closer the new energy output is to the rated value, the more intense the working conditions and the faster the risk propagation, the higher the factor value.

[0243] S33: Chain network generation: When the correlation strength is ≥0.65, a directed edge is established between nodes, with the arrow direction indicating the risk propagation direction; high-priority correlation edges are additionally marked for new energy-related nodes, and an additional 10% correction coefficient is added when calculating the correlation strength, ultimately forming a full-process scenario chain of the power grid, including planning, construction, operation, maintenance, and decommissioning;

[0244] The node initialization covers all stages of the power grid's entire life cycle and adds new energy-related risk nodes to address the limitations of traditional scenario chains that only focus on traditional power grid stages and ignore new energy risks. High-priority associated edges are marked for new energy nodes and a 10% association strength correction coefficient is added to highlight the propagation characteristics of new energy risks, making the scenario chain more adaptable to new energy grid-connected power grids.

[0245] S4: Based on the entire power grid scenario chain, the real-time risk value of each node is calculated through the risk warning model. After comparing it with the dynamic warning threshold, a graded warning response is triggered to achieve full-cycle control of power grid risks.

[0246] The risk warning model employs a dynamic warning threshold algorithm, specifically:

[0247] Let the real-time risk value of the k-th node in the scenario chain be R. k :

[0248]

[0249] Where w is the feature weight vector, b is the bias term, and the training samples cover power grid faults and new energy anomaly cases;

[0250] The dynamic early warning threshold is Th k :

[0251] Th k =Th 0k ·[1+γ·D k +η·C k +θ·ΔP k +μ·A k ]

[0252] Among them, Th 0k Based on the basic threshold, the definitions and calculation methods of the remaining parameters are as follows:

[0253] Operating condition influence coefficient γ: reflects the adjustment range of the threshold due to real-time operating conditions, with a value range of 0.2-0.5, and adjusts linearly with the node load rate ρ. The formula is:

[0254] γ = 0.2 + 0.3ρ

[0255] Where ρ∈[0,1], the higher the load rate, the more intense the working conditions, the larger the coefficient, and the higher the threshold to avoid false alarms;

[0256] Operating condition deviation D k This reflects the degree of deviation between the current operating condition and the rated operating condition; the formula is:

[0257]

[0258] Where, x ck This is the current operating condition vector of the node, containing voltage, current, and power; x 0k x is the node's rated operating condition vector; ck,new x is the current output vector of the new energy node; 0k,new The rated output vector of the new energy node; ||·|| is the Euclidean distance, the greater the deviation, the greater the threshold adjustment range;

[0259] Risk accumulation coefficient η: reflects the impact of risk duration on the threshold, with a value ranging from 0.1 to 0.3, adjusted according to the node risk duration t, and the formula is:

[0260] η = 0.1 + 0.2 × min(t / 60, 1)

[0261] Among them, the t of the new energy-related nodes needs to be superimposed with the duration of the power output fluctuation, that is, t = t 风险 +t 波动 The longer the duration, the larger the coefficient, and the lower the threshold to trigger an alert;

[0262] New energy output fluctuation factor θ: reflects the impact of new energy output fluctuation on the threshold, with a value range of 0.1-0.3, and varies with the new energy output fluctuation rate ΔP. k Adjust the formula as follows:

[0263] θ = 0.1 + 0.2 × min(ΔP) k / 0.15,1)

[0264] in, The value range is [0,1]. The greater the volatility, the larger the factor and the lower the threshold.

[0265] Equipment aging factor μ: reflects the impact of equipment aging degree on the threshold, with a value range of 0.1-0.25, varying with the equipment's operating years Y. k With design life Y k,设计 The ratio adjustment is calculated using the following formula:

[0266] μ = 0.1 + 0.15 × min(Y) k / Y k,设计 ,1)

[0267] Among them, the larger the ratio, the more severe the aging; the larger the factor, the lower the threshold.

[0268] Warning triggering and classification rules: When R k ≥Thk When an alert is triggered, press The risk levels are categorized as follows: 0-0.2 (low level), 0.2-0.5 (medium level), and >0.5 (high level). The early warning level for new energy-related nodes is raised by one level, while the high level remains unchanged, thus strengthening the response priority for new energy risks.

[0269] The dynamic early warning threshold incorporates four factors: operating conditions, risk accumulation, fluctuations in new energy output, and equipment aging. The threshold adjustment is linked in real time with the actual operating status of the power grid. Each factor is dynamically quantified and adjusted according to actual operating parameters, so that the threshold can accurately match the risk perception needs under different scenarios.

[0270] The specific measures for tiered early warning response are as follows:

[0271] Low-level warning: Generate equipment tracking instructions, increase the sampling frequency of conventional sensors and the monitoring frequency of new energy output, and upload data to the power grid monitoring platform in real time. Maintenance personnel need to check the tracking data every hour.

[0272] Medium-level early warning: Activate the joint plan for load transfer and new energy power output adjustment, use the Newton-Raphson method to calculate the power flow of the grid and screen the load transfer path; simultaneously calculate the adjustment amount of new energy power output, push the plan to the grid dispatch terminal and the new energy power station control terminal, and dispatch personnel complete the execution of the plan;

[0273] High-level early warning: Automatically triggers emergency control measures, links circuit breakers to cut off power to the fault area, and additionally links inverters at new energy grid-connected nodes to cut off power in an emergency; at the same time, it generates a maintenance work order, including GPS coordinates of the fault location, equipment model, recommended tool list, and temporary alternative power output plan for new energy, pushes it to the mobile terminal of maintenance personnel, and simultaneously reports it to the power grid supervision platform and new energy supervision department, and maintenance personnel arrive at the site to handle the situation.

[0274] Example 2: A 220kV regional power grid covers three industrial parks and two residential areas. It includes one 220kV hub substation, three 220kV transmission lines, and two main transformers. The main transformers have a design life of 20 years and have been in operation for 15 years. They are connected to a 100MW photovoltaic power station and a 50MW wind farm, serving the industrial and residential loads in the area. The power grid needs to cover the entire life cycle of planning, construction, operation, maintenance, and decommissioning, and faces typical risks such as fluctuations in renewable energy output, equipment aging, and environmental interference.

[0275] Data Acquisition: Real-time Operational Data: The SCADA system acquires voltage 220±5kV, current 0-800A, and power 0-200MW at a sampling interval of 1s. PMU synchronizes phasor data at a sampling frequency of 100Hz.

[0276] Equipment status data: Dissolved gases in main transformer oil: H2: 15μL / L, CH4: 20μL / L, C2H2: 0.5μL / L, line infrared temperature measurement 45℃, main transformer health index calculated based on gas concentration is 0.65, indicating a moderate health status;

[0277] Environmental interference data: Regional wind speed 18m / s, low interference; ambient temperature 32℃; typhoon forecast for the next 24 hours, maximum wind speed 25m / s, high interference.

[0278] New energy data: Photovoltaic power plant real-time output 40MW, predicted output 45MW, volatility 11%; wind farm real-time output 20MW, predicted output 22MW.

[0279] Preprocessing:

[0280] Missing value imputation: There were 3 missing values ​​in the photovoltaic output data, with a missing rate of 3%. Linear interpolation was used to imput the missing values.

[0281] Noise filtering: The wind power output data is denoised using db4 wavelet transform and then superimposed with Kalman filtering, Q=0.01, R=0.05, to suppress fluctuations;

[0282] Dimensional unification: Normalize all data to the range [0,1] using Min-Max to eliminate dimensional differences;

[0283] Calculate the weights of each data point:

[0284] Time weighting: The absolute value of the covariance between photovoltaic power output data and grid frequency is large, indicating a strong time-series correlation. t,光伏 =0.3;

[0285] Spatial weight: The main transformer is only 2km away from the hub substation, and its load capacity accounts for 40% of the area. s,主变 =0.25;

[0286] Equipment health weight: The main transformer health index is 0.65, lower than the 0.9 for new equipment. h,主变 =0.2, lower than the 0.3 of health equipment;

[0287] Environmental weights: Current wind speed 18 m / s, low interference, environmental data w e =0.7; Typhoon forecast, high interference, environmental weight adjusted to 0.3 after 2 hours;

[0288] After fusion, a core feature vector F is generated, which includes key information such as the health status of the main transformer, photovoltaic power output fluctuations, and line load rate.

[0289] Node initialization: Identify core risk scenario nodes, including 8 nodes such as insufficient grid-connected photovoltaic capacity, aging and overloaded main transformer, and line icing;

[0290] Correlation analysis: Calculation of the correlation strength between a sudden drop in photovoltaic power output and main transformer overload:

[0291] Conditional probability: 32 historical co-occurrences, P = 0.35;

[0292] Transmission time: average 5 minutes, adjusted to 3 minutes during typhoon days;

[0293] Node importance: Main transformer load accounts for 40%, topology centrality is 0.8, I a =0.6×0.4+0.4×0.8=0.56;

[0294] Real-time operating conditions: Current load factor 85%, photovoltaic output 40MW, rated 45%, G a =0.7×0.85+0.3×0.89=0.84;

[0295] The correlation strength R = 0.4 × 0.35 + 0.3 × exp(-3 / 20) + 0.3 × 0.56 × 0.84 ≈ 0.72, ≥ 0.65 threshold, establish directed edges;

[0296] Scenario chain generation: forming a high-priority scenario chain of insufficient grid-connected photovoltaic capacity → sudden drop in photovoltaic output → main transformer overload → line tripping;

[0297] Risk value calculation: Real-time risk value R of the main transformer overload node based on fused feature F. k =0.82;

[0298] Dynamic threshold calculation:

[0299] Basic threshold: Standard Th for main transformer overload node 0k =0.8;

[0300] Operating condition deviation: Current load rate 85%, rated 70%, D k =0.2;

[0301] New energy volatility factor: Photovoltaic volatility 11%, θ=0.1+0.2×(0.11 / 0.15)=0.25;

[0302] Equipment aging factor: Main transformer operating for 15 years, designed for 20 years, μ=0.1+0.15×(15 / 20)=0.21;

[0303] Dynamic threshold Th k =0.8×[1+0.4×0.2+0.2×0.1+0.25×0.11+0.21×0.1]≈0.79;

[0304] Warning triggered: R k =0.82≥Th k=0.79, according to Initially classified as a low-level warning; due to its association with a new energy node, the level has been upgraded to a medium-level warning.

[0305] Response measures:

[0306] Contingency plan activated: Calculate power flow using the Newton-Raphson method and transfer the 10MW load from the industrial park to an adjacent line, increasing network losses by ≤2%;

[0307] The photovoltaic power station was instructed to adjust its output to 45MW with a fluctuation rate of ≤5%, and simultaneously increase the main transformer monitoring frequency to 0.5s / time.

[0308] Maintenance recommendations: Based on a health index of 0.65, arrange for main transformer oil sample testing two days in advance, and reinforce the line icing monitoring device before the typhoon.

[0309] In summary, this invention, by employing an adaptive fusion algorithm in the feature layer fusion stage, overcomes the limitations of traditional fusion methods that rely solely on spatiotemporal dimensions. It incorporates equipment health status and environmental interference factors into the weight allocation system, effectively addressing issues such as uneven data reliability due to differences in equipment health and data distortion caused by environmental interference during power grid data acquisition. The invention dynamically adjusts the weights of corresponding data based on the equipment health index, prioritizing the use of data from equipment with good health status. Simultaneously, it adjusts data contribution based on environmental interference levels, ensuring that the fused core features better reflect the actual operating state of the power grid. This avoids feature bias caused by single-dimensional fusion, providing more accurate basic data support for subsequent risk scenario identification and correlation analysis, and improving the initial accuracy of overall risk warning.

[0310] In the construction of the power grid full-process scenario chain, a coupled correlation algorithm is adopted, adding the coupled calculation of node importance and real-time operating conditions. By quantifying the load ratio and topological centrality of hub nodes to reflect node importance, and combining real-time load rate and renewable energy output status to characterize the impact of operating conditions, the calculation of the correlation strength between nodes is no longer limited to static statistical laws, but can dynamically respond to changes in the characteristics of key power grid nodes and real-time operating conditions. This effectively avoids the omission or misjudgment of key risk paths by static correlation analysis, ensuring that the constructed full-process scenario chain can accurately identify the core risk propagation paths at each stage from planning to decommissioning, and providing a more targeted scenario carrier for the full-cycle management of power grid risks.

[0311] In designing risk warning thresholds, a dynamic warning threshold algorithm is adopted to compensate for the neglect of the characteristics of new energy sources and the impact of equipment aging by traditional threshold designs that only consider operating conditions and risk accumulation. This significantly improves the scenario adaptability of the warning thresholds. Incorporating new energy output fluctuations into the threshold adjustment dimension adapts to the special characteristics of power grid operation under high new energy penetration rates, preventing new risks such as sudden drops in new energy output from being missed due to rigid thresholds. Simultaneously, the aging impact is quantified by combining the ratio of equipment operating years to design life, making the warning thresholds for aging equipment more closely match its actual risk resistance capabilities and identifying potential failure risks of aging equipment in advance. This multi-dimensional adjustment mechanism of dynamic thresholds effectively reduces false alarms and missed alarms of fixed thresholds under different operating conditions, new energy fluctuations, and equipment states. This makes the warning response more aligned with the operating characteristics of the entire power grid lifecycle, providing more accurate risk warning support for the safe and stable operation of the power grid.

[0312] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the invention can be implemented in other specific forms without departing from its spirit or essential characteristics. Therefore, the embodiments should be considered in all respects as exemplary and non-limiting, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be included within the present invention. No reference numerals in the claims should be construed as limiting the scope of the claims.

Claims

1. A method for risk early warning of the entire power grid scenario chain through multi-source heterogeneous data fusion, characterized by: Includes the following steps: S1: Collect multi-source heterogeneous raw data of the power grid, and preprocess the multi-source heterogeneous raw data by missing value imputation, noise filtering and dimension unification to obtain standardized processed data; S2: The standardized processed data is fused using a three-layer fusion architecture consisting of a data layer, a feature layer, and a decision layer to extract the core features of the power grid operation status and generate a unified feature dataset; S3: Based on the unified feature dataset, risk scenario nodes are divided into each stage of the power grid's entire life cycle. Coupled correlation algorithm is used to mine the correlation between nodes and construct a power grid full-process scenario chain. S4: Based on the power grid full-process scenario chain, the real-time risk value of each node is calculated through the risk warning model, and the risk value is compared with the dynamic warning threshold to trigger a graded warning response, thereby realizing full-cycle control of power grid risks.

2. The method for risk early warning of the entire power grid scenario chain through multi-source heterogeneous data fusion as described in claim 1, characterized in that: In S1, the multi-source heterogeneous raw data of the power grid specifically includes: Real-time operating data: voltage, current, and power data of the SCADA system, and PMU synchronization phasor data; Equipment status data: dissolved gas concentration in transformer oil, infrared temperature measurement data of the line, partial discharge of switchgear, and equipment health index; among which, the equipment health index is calculated based on the dissolved gas concentration in oil and partial discharge, and the value range is [0,1]. The closer the value is to 1, the better the equipment health status. Environmental interference data: regional wind speed, rainfall, ice thickness, ambient temperature, and environmental interference level; among which, regional wind speed, rainfall, ice thickness, and ambient temperature are divided into high interference and low interference to assess the reliability of data collection. Management support data: daily electricity load curve, photovoltaic or wind power output data, grid topology parameters, and output prediction curves of new energy grid-connected nodes.

3. The method for risk early warning of the entire power grid scenario chain through multi-source heterogeneous data fusion as described in claim 1, characterized in that: In step S1, the specific process for preprocessing multi-source heterogeneous raw data is as follows: S11: Missing value imputation: If the missing rate is ≤5%, linear interpolation is used, and the formula is: Where, x t The missing value at time t, t t For the corresponding timestamp, x t-1 x t+1 The values ​​are valid values ​​at adjacent time points; If the missing rate is >5%, the K-nearest neighbor imputation method is used, where K = 3-5. Similar samples are selected based on Euclidean distance, and similar samples must meet the requirement of consistent environmental interference levels to avoid imputation deviation caused by environmental differences. S12: Noise Filtering Wavelet transform algorithm is used for basic denoising, with the db4 wavelet basis function selected for three-level decomposition. Level 1 decomposition separates high-frequency random noise, and levels 2-3 decomposition separate low-frequency interference generated by equipment operation. Matching the characteristics of the power grid data signal, soft thresholding is applied to the decomposed high-frequency coefficients. The threshold calculation formula is as follows: Where σ is the noise standard deviation and N is the data length; random noise is removed after reconstruction; Kalman filtering is added to the new energy output data, where the process noise covariance Q = 0.01 and the observation noise covariance R = 0.05 are set. Through the iterative process from state prediction to observation update, the output fluctuation noise is further smoothed to ensure that the noise level of new energy data is consistent with that of other power grid data and to suppress the systematic noise caused by output fluctuation. S13: Dimensional uniformity: Using Min-Max normalization, the formula is as follows: Where x represents the multi-source heterogeneous original data, x max x min These represent the maximum and minimum values ​​in the data sample, respectively; the output is standardized data with a value range of [0,1] to eliminate the influence of dimensional differences on the fusion result.

4. The method for risk early warning of the entire power grid scenario chain through multi-source heterogeneous data fusion as described in claim 1, characterized in that: In S2, the three-layer converged architecture includes: 1) Data layer fusion: For sensor data of the same type, a weighted average method is used. Among them, w data,i For the weight, σ i Let σ be the standard deviation of the data from the i-th sensor. avg The average standard deviation of all data from the same type of sensor is used. The smaller the standard deviation, the higher the weight. Stable data is given priority. Additional prediction accuracy weights are introduced for new energy output data: Where, φ i The prediction accuracy of the i-th group of new energy output data is defined as follows, with a value range of [0,1], and the final weight is w. data,i ×w p Strengthen the contribution of high-precision prediction data; 2) Feature layer fusion: An adaptive fusion algorithm is adopted to achieve effective integration of heterogeneous features through multi-dimensional weight allocation; 3) Decision-level fusion: Preliminary risk results from various data sources are synthesized using the DS evidence theory. 31) Divide the data source into 4 categories of evidence E = {E1, E2, E3, E4}, where: Device status class E1: Based on the device health feature vector F output by the feature layer h The initial risk probabilities P1(θ1) and P1(θ2) of normal equipment (θ1) and abnormal equipment (θ2) are generated by the SVM classifier, where the SVM kernel function is RBF, the penalty coefficient C = 10, and the gamma parameter = 0.

1. Real-time running class E2: Based on real-time running feature vector F r Calculate the deviation rate of voltage, current, and power from the rated values. The probability of mapping to normal operation (θ1) and operation exceeding the limit (θ2) is P2(θ1) = 1 - min(δ / 0.2,1) and P2(θ2) = min(δ / 0.2,1), where P2(θ2) = 1 when the deviation rate exceeds 20%. New Energy Category E3: Based on New Energy Feature Vector F new Combined with power output volatility Based on the inverter status parameters, the probabilities of generating normal (θ1) and abnormal (θ2) renewable energy output are P3(θ1) = 1 - (min(ΔP / 0.15,1)×0.7 + inverter abnormality coefficient×0.3) and P3(θ2) = 1 - P3(θ1), where the inverter abnormality coefficient is calculated from the capacitor temperature and IGBT on-state voltage drop, and takes a value of 0-1; Environment class E4: Based on environment feature vector F e The probability mappings based on environmental interference levels are as follows: for high interference, P4(θ1) = 0.3 and P4(θ2) = 0.7; for low interference, P4(θ1) = 0.8 and P4(θ2) = 0.

2. 32) For each type of evidence, introduce the uncertainty proposition Θ = {θ1, θ2}, where θ1 is "risk-free", θ2 is "risky", and the BPA function m i (A) The definition is as follows: Single-point proposition: m i ({θ1})=P i (θ1)×λ i m i ({θ2})=P i (θ2)×λ i Where, λ i The credibility of the evidence is calculated based on the signal-to-noise ratio after data fusion at the data layer. Values ​​range from 0.6 to 0.95; Uncertainty propositions: m i (Θ)=1-m i ({θ1})-m i ({θ2}) This reflects the uncertainty caused by the lack of evidence, such as ambiguity due to incomplete data sampling; 33) Calculate the conflict coefficient K between pieces of evidence to determine the degree of conflict: If K ≤ 0.3 indicates low conflict: directly apply the DS composition rule, that is, for any non-empty proposition... The synthesized BPA is: If K > 0.3 indicates high conflict: introduce an evidence discount factor α. i =1-0.2×K, the greater the conflict, the stronger the discount, and the BPA of each piece of evidence is adjusted to m. i ′({θ1})=m i ({θ1})×α i m i ′({θ2})=m i ({θ2})×α i m i ′(Θ)=1-m i ′({θ1})-m i Then perform the above DS synthesis again; 34) Based on the synthesized BPA, calculate the trust function Bel(θ1) = m({θ1}) and the likelihood function Pl(θ1) = 1 - m({θ2}) for the risk-free θ1, and output the comprehensive risk result according to the following rules: If Bel(θ2)≥0.7: it is judged as high risk, triggering early warning preparation; If 0.4 ≤ Bel(θ2) < 0.7: it is judged as medium risk, and a risk warning is output; If Bel(θ2) < 0.4: it is judged as low risk, and only the risk trend is recorded; If Pl(θ1)-Bel(θ1)>0.5: it is judged as high uncertainty: automatically increase the sampling frequency of the corresponding data source, re-collect data and make a secondary fusion decision.

5. The method for risk early warning of the entire power grid scenario chain based on multi-source heterogeneous data fusion according to claim 4, characterized in that: The adaptive fusion algorithm used in the feature layer fusion is as follows: Let the feature vector of the i-th class of standardized data be x. i Let i = 1, 2, ..., n, where n is the number of data source types, and the core feature vector after fusion be F. Then: Wherein, weight w i By time weight w t,i Spatial weight w s,i Equipment health weight w h,i Environmental interference weight w e,i Adaptive calculation, the formula is: w i =α·w t,i +β·w s,i +γ·w h,i +(1-a-b-c)·w e,i Where α is the time weight coefficient, β is the spatial weight coefficient, and γ is the health weight coefficient, satisfying 0.2≤α≤0.4, 0.2≤β≤0.4, 0.1≤γ≤0.3, and α+β+γ≤1. Grid search optimization is performed with the goal of maximizing the Pearson comprehensive correlation between the fused features and the grid fault label and the new energy output anomaly label. The specific calculation methods for each weight are as follows: Time weight w t,i The formula reflects the temporal correlation between data and the real-time status of the power grid: Where y is the real-time frequency / voltage vector of the power grid, p is the real-time power output vector of new energy sources, and Cov is the covariance. The larger the absolute value of the covariance, the stronger the time-series correlation and the higher the weight. Spatial weight w s,i The formula reflects the impact of spatial topological associations of devices on data importance: Where, d i This represents the shortest topological distance from the corresponding data device to the hub substation; 0.1 is a correction term to avoid a denominator of 0. i S represents the load capacity of the node where the equipment is located. max The maximum load capacity of a node in the entire network is used; the closer the node is, the greater its load capacity, and the higher its weight. Device health weight w h,i The formula reflects the impact of equipment health status on data reliability: Among them, HI i The health index of the device corresponding to the i-th data type is calculated as follows: c k,i c represents the actual concentration of dissolved gases in the oil. k,max The higher the health index, the stronger the data reliability and the higher the weight, corresponding to the safety threshold of the gas. Environmental interference weight w e,i The formula reflects the impact of environmental interference on the reliability of data acquisition: Sensor data is prone to distortion in high-interference environments, so the weighting is reduced; data reliability is high in low-interference environments, so the weighting is increased.

6. The method for risk early warning of the entire power grid scenario chain through multi-source heterogeneous data fusion as described in claim 1, characterized in that: In S3, the stages of the entire power grid life cycle and the corresponding risk scenario nodes include: Planning phase: Nodes with insufficient topology redundancy, nodes with equipment selection errors, and nodes with insufficient grid-connected capacity of new energy sources; the corresponding risks are: power flow blockage, overload burnout, and curtailment of solar or wind power. Construction phase: Key points include substandard construction techniques, installation deviations, and incorrect selection of new energy access cables; the corresponding risks are: line short circuit, partial discharge, and cable overheating, respectively. Operational phases include: overload nodes, voltage sag nodes, frequency offset nodes, and sudden drop in renewable energy output nodes; the corresponding risks are: equipment overheating, sensitive load outages, grid instability, and power deficit, respectively. Maintenance phase: Overlooked maintenance points, delayed maintenance points, and missing maintenance points for new energy inverters; the corresponding risks are: fault expansion, accelerated equipment aging, and inverter failure, respectively. The decommissioning phase includes addressing safety hazards, asset waste, and oversights in the environmental treatment of decommissioned equipment; the corresponding risks are personal injury, resource depletion, and environmental pollution.

7. The method for risk early warning of the entire power grid scenario chain based on multi-source heterogeneous data fusion according to claim 1, characterized in that: In S3, the specific steps for constructing the entire power grid scenario chain are as follows: S31: Node Initialization: Use the risk scenario nodes as the initial node set V = {v1, v2, ..., v...} m }, where m is the total number of nodes, and each node is associated with a corresponding risk type, scope of impact, and characteristic indicators; S32: Association Mining: Using a coupled association algorithm, the association strength between any two nodes is calculated to identify risk propagation paths; S33: Chain network generation: When the correlation strength is ≥0.65, a directed edge is established between nodes, with the arrow direction indicating the risk propagation direction; high-priority correlation edges are additionally marked for new energy-related nodes, and an additional 10% correction coefficient is added when calculating the correlation strength, ultimately forming a full-process scenario chain for the power grid, including planning, construction, operation, maintenance, and decommissioning.

8. The method for risk early warning of the entire power grid scenario chain based on multi-source heterogeneous data fusion according to claim 7, characterized in that: The coupling correlation algorithm is specifically as follows: Let the preceding node be the risk propagation starting point v. a To subsequent nodes: Risk propagation endpoint v b The correlation strength is R a→b ,but: Wherein, β1 is the probability weight coefficient, β2 is the time weight coefficient, and β3 is the importance and operating condition coupling weight coefficient, satisfying 0.3≤β1≤0.5, 0.2≤β2≤0.4, 0.1≤β3≤0.3, and β1+β2+β3=1, calibrated through cascading failure cases; the specific calculation methods for each parameter are as follows: Conditional probability P(v) b |v a ): Reflects v a After it happened v b The trigger probability is given by the formula: Where, N a∩b For v a With v b The number of co-occurring samples, N a For v a The total number of samples, m is the total number of nodes, +1 is Laplace smoothing to avoid a probability of 0; N a∩b∩p For v a v b The number of co-occurrence samples of abnormal power output events with new energy sources, N a∩p For v a The number of co-occurrence samples with abnormal power output from new energy sources strengthens the correlation probability of new energy-related nodes; Propagation time T a→b : Reflects v a Risk spread to v b The average time is calculated based on historical fault time-series data; for nodes related to sudden drops in new energy output, an additional correction for output fluctuation response time is introduced, and the correction formula is: T a→b ′=T a→b ×(1+0.2×ΔP) Wherein, ΔP is the power output volatility of new energy sources, with a value range of [0,1]. The greater the volatility, the longer the response time. Maximum tolerance time T max The maximum tolerance time for risk propagation in the power grid is adjusted according to the penetration rate of new energy sources. Node Importance I a : Reflects the preceding node v a The influence of the power grid status on the correlation strength is expressed by the following formula: Among them, L a For v a The proportion of the load of the node to the total network load, L max The highest percentage of load across the entire network; C a For v a Degree centrality is the ratio of the number of other nodes connected to a node to the total number of nodes in the network. The value ranges from [0,1]. The higher the load ratio and the denser the topology connections, the higher the importance. Real-time operating condition factor G a : Reflects v a The formula for the accelerating effect of the real-time operating conditions of the node on risk propagation is: Where, ρ a For v a The real-time load factor of the node, with a value range of [0,1]; P a,new For v a The associated renewable energy nodes are generating power in real time, P a,new,max The rated output of the new energy node is defined as [0,1]. The higher the load factor, the closer the new energy output is to the rated value, the more intense the operating conditions, the faster the risk propagation, and the higher the factor value.

9. The method for risk early warning of the entire power grid scenario chain based on multi-source heterogeneous data fusion according to claim 1, characterized in that: In S4, the risk warning model adopts a dynamic warning threshold algorithm, specifically as follows: Let the real-time risk value of the k-th node in the scenario chain be R. k : Where w is the feature weight vector, b is the bias term, and the training samples cover power grid faults and new energy anomaly cases; The dynamic early warning threshold is Th k : Th k =Th 0k ·[1+γ·D k +η·C k +θ·ΔP k +μ·A k ] Among them, Th 0k Based on the basic threshold, the definitions and calculation methods of the remaining parameters are as follows: Operating condition influence coefficient γ: reflects the adjustment range of the threshold due to real-time operating conditions, with a value range of 0.2-0.5, and adjusts linearly with the node load rate ρ. The formula is: γ = 0.2 + 0.3ρ Where ρ∈[0,1], the higher the load rate, the more intense the working conditions, the larger the coefficient, and the higher the threshold to avoid false alarms; Operating condition deviation D k : Reflects the degree of deviation between the current operating condition and the rated operating condition, the formula is: Where, x ck This is the current operating condition vector of the node, containing voltage, current, and power; x 0k x is the node's rated operating condition vector; ck,new x is the current output vector of the new energy node; 0k,new The rated output vector of the new energy node; ||·|| is the Euclidean distance, the greater the deviation, the greater the threshold adjustment range; Risk accumulation coefficient η: reflects the impact of risk duration on the threshold, with a value ranging from 0.1 to 0.3, adjusted according to the node risk duration t, and the formula is: η = 0.1 + 0.2 × min(t / 60, 1) Among them, the t of the new energy-related nodes needs to be superimposed with the duration of the power output fluctuation, that is, t = t 风险 +t 波动 The longer the duration, the larger the coefficient, and the lower the threshold to trigger an alert; New energy output fluctuation factor θ: reflects the impact of new energy output fluctuation on the threshold, with a value range of 0.1-0.3, varying with the new energy output fluctuation rate ΔP. k Adjustment, the formula is: θ=0.1+0.2×min(ΔP k / 0.15,1) in, The value range is [0,1]. The greater the volatility, the larger the factor and the lower the threshold. Equipment aging factor μ: reflects the impact of equipment aging degree on the threshold, with a value range of 0.1-0.25, varying with the equipment's operating years Y. k With design life Y k,设计 The ratio adjustment is calculated using the following formula: μ=0.1+0.15×min(Y k / AND k,设计 ,1) Among them, the larger the ratio, the more severe the aging; the larger the factor, the lower the threshold. Warning triggering and classification rules: When R k ≥Th k When an alert is triggered, press The risk levels are categorized as follows: 0-0.2 (low level), 0.2-0.5 (medium level), and >0.5 (high level). The early warning level for new energy-related nodes is raised by one level, while the high level remains unchanged, thus strengthening the response priority for new energy risks.

10. The method for risk early warning of the entire power grid scenario chain based on multi-source heterogeneous data fusion according to claim 1, characterized in that: The specific measures for the graded early warning response in S4 are as follows: Low-level warning: Generate equipment tracking instructions, increase the sampling frequency of conventional sensors and the monitoring frequency of new energy output, and upload data to the power grid monitoring platform in real time. Maintenance personnel need to check the tracking data every hour. Medium-level early warning: Activate the joint plan for load transfer and new energy power output adjustment, use the Newton-Raphson method to calculate the power flow of the grid and screen the load transfer path; simultaneously calculate the adjustment amount of new energy power output, push the plan to the grid dispatch terminal and the new energy power station control terminal, and dispatch personnel complete the execution of the plan; High-level early warning: Automatically triggers emergency control measures, linking circuit breakers to cut off power to the fault area, and additionally linking inverters at new energy grid-connected nodes to cut off power in an emergency; at the same time, a maintenance work order is generated, including GPS coordinates of the fault location, equipment model, recommended tool list, and temporary alternative power output solutions for new energy, which is pushed to the mobile terminal of maintenance personnel and simultaneously reported to the power grid supervision platform and new energy supervision department, and maintenance personnel arrive at the site to handle the situation.

Citation Information

Cited By

  • Intelligent simulation system and method based on power distribution digital system

    CN121638073A

  • Electric power measurement data security risk assessment method and device

    CN122022503A