Power failure risk assessment method for multi-source data
By calculating the time difference and historical deviation of the data source to generate a time-varying credibility set and establishing an adaptive weight configuration, the timeliness and accuracy problems in multi-source data evaluation are solved, and more reliable power outage risk assessment and early warning are achieved.
Patent Information
- Application Number
- CN202511497811.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-20
- Publication Date
- 2026-01-16
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing technologies fail to effectively address the differences in timeliness and accuracy of data sources when assessing power outage risks from multiple data sources, leading to distorted assessment results.
By calculating the difference between the generation timestamp of each data source and the current time, and combining the deviation between historical predictions and actual records, a multi-source data time-varying reliability set is generated. Based on this, an adaptive data source weight configuration is established, data weighting calculation is performed, and an initial value of power grid unit fusion risk is generated. Finally, the risk threshold is used to make step-by-step judgments and determine the early warning level.
Ensuring that the information used for assessment is novel and historically validated helps to suppress interference from low-quality data, provides a more reliable basis for decision-making, and improves the practicality and accuracy of fault prediction.
Smart Images

Figure CN121350700A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of fault prediction technology, and in particular to a method for assessing power outage risks based on multi-source data. Background Technology
[0002] Fault prediction is used for real-time monitoring and analysis of the operational status of industrial equipment, systems, or infrastructure, identifying potential anomalies, degradation trends, or precursors to failure before functional failures occur. Current technologies assume all connected data sources have equal reliability and timeliness, directly inputting raw data into the model for analysis. This ignores the fact that in actual operation, the data sources themselves can experience quality fluctuations due to sensor aging, communication delays, model drift, etc. For example, a weather station's sensor may fail to update data for several hours, while the accuracy of another load forecasting model may decline significantly within a week. Treating these outdated or inaccurate data the same as normal data contaminates the risk assessment model, leading to distorted prediction results. Therefore, improvements are needed. Summary of the Invention
[0003] The purpose of this invention is to address the shortcomings of existing technologies by proposing a power outage risk assessment method based on multi-source data.
[0004] To achieve the above objectives, the present invention adopts the following technical solution: a power outage risk assessment method based on multi-source data, comprising the following steps: The difference between the generation timestamp of each data source and the current system time is obtained to calculate the timeliness score. Then, the deviation between the predicted values of meteorological data and load data in the previous cycle and the actual state records of the power grid is retrieved to calculate and generate a multi-source data time-varying reliability set. Based on the multi-source data time-varying credibility set, the credibility values of all data sources are accumulated to obtain the total credibility value. The credibility value of each data source in the multi-source data time-varying credibility set is then calculated with the total credibility value to establish an adaptive data source weight configuration. Based on the adaptive data source weight configuration, the current temperature, humidity, wind speed, and transformer load rate are extracted to obtain a pairing list. The initial value of the power grid unit fusion risk is obtained based on the pairing list. Based on the initial risk value of the power grid unit, a step-by-step judgment is made against the multi-level risk thresholds preset for different power grid units such as switching equipment to obtain the risk range attribution determination. Based on the result of the risk range attribution determination, the corresponding level identifier is selected from the predefined risk level library to establish a power outage risk warning level.
[0005] Preferably, the steps for obtaining the multi-source data time-varying reliability set are as follows: The generation timestamp of each data source is analyzed and compared with the current system time. The time difference of the first data source is calculated. The predicted values of meteorological data and load data in the previous cycle are paired with the actual power grid status records under the same time index to obtain the error sequence. The time difference and error sequence are then organized according to the data source identifier. The time-varying reliability value of the data source is calculated based on the time difference and error sequence. Based on the time-varying confidence values, the time-varying confidence values of each data source are collected in the order of data source identifiers, outliers and missing items are removed, and a multi-source data time-varying confidence set is generated.
[0006] Preferably, the step of obtaining the total credibility value is as follows: Based on the multi-source data time-varying confidence set, the time-varying confidence value of each data source is read source by source. Traversal pointers are established according to the fixed order of the data source identifiers. Each time-varying confidence value is written into the cumulative register area in turn, and the index position of the data source identifier is recorded at the same time. After the end pointer verification is completed, the sum of all written values is obtained to get the total confidence value.
[0007] Preferably, the step of obtaining the adaptive data source weight configuration is as follows: Based on the total confidence value, the multi-source data time-varying confidence set is traversed again in a fixed order according to the data source identifier, the corresponding time-varying confidence value is read, a division operation is performed to obtain the normalized value, and each normalized value is written into the same-order result sequence according to the index position to generate a normalized weight sequence for transmission line fault prediction. Based on the normalized weight sequence for transmission line fault prediction, the one-to-one correspondence between the sequence items and the data source identifier is checked, missing records are removed and index gaps are filled, and the data is encapsulated into a key-value pair structure in a fixed order to form an adaptive data source weight configuration.
[0008] Preferably, the step of obtaining the pairing list is as follows: Based on the adaptive data source weight configuration, extract the current temperature value, current humidity value, current wind speed value, and current transformer load rate value according to the data source identifier, verify the consistency of the time index, and complete the unit conversion to obtain the pairing list.
[0009] Preferably, the steps for obtaining the initial value of the grid unit fusion risk are as follows: Based on the pairing list, calculate the initial value of the grid unit fusion risk.
[0010] Preferably, the steps for determining the risk range attribution are as follows: Based on the initial risk value of the power grid unit integration, a multi-level risk threshold table is retrieved according to the power grid unit type. The thresholds are verified to be ordered and without repetition, and the unified boundary is left closed and right open, forming a multi-level risk threshold sequence for the power grid unit. Based on the multi-level risk threshold sequence of the power grid unit, the initial risk value of the power grid unit is compared step by step from the smallest to the largest. When the lower limit is hit, it is recorded as the interval to which it belongs. When it is equal to the upper limit, it is moved up one level. When it exceeds the limit, it is assigned to the lowest or highest interval respectively, thus generating a risk interval assignment determination.
[0011] Preferably, the step of obtaining the power outage risk warning level is as follows: Based on the risk range attribution determination, a predefined risk level library is retrieved, and the corresponding level identifier is mapped according to the range code to form a power outage risk warning level.
[0012] Compared with the prior art, the advantages and positive effects of the present invention are as follows: This invention generates a data credibility set that changes over time by quantifying the difference between the generation time and the current time of each data source and combining it with the deviation between its historical predictions and actual records. This ensures that the information used for evaluation is not only novel but also verified by historical performance. This credibility set is then transformed into a data source weight configuration, allowing higher-quality and more reliable data sources to have a greater influence in subsequent calculations and suppressing noise interference from low-quality data. When aggregating real-time features of multiple dimensions such as temperature, humidity, wind speed, and load rate, a weighted calculation is performed to generate a fusion risk initial value that reflects the current pressure state of the power grid unit. This initial value is then compared with preset multi-level risk thresholds, providing operation and maintenance personnel with an intuitive and operable decision-making basis, improving the practicality of fault prediction, and thus shifting from passive response to proactive intervention. Attached Figure Description
[0013] Figure 1 This is a schematic diagram of the steps of the present invention. Detailed Implementation
[0014] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.
[0015] Please see Figure 1 This invention provides a technical solution, a method for power outage risk assessment based on multi-source data, comprising the following steps: The difference between the generation timestamp of each data source and the current system time is obtained to calculate the timeliness score. Then, the deviation between the predicted values of meteorological data and load data in the previous cycle and the actual state records of the power grid is retrieved to calculate and generate a multi-source data time-varying reliability set. Based on the time-varying credibility set of multi-source data, the credibility values of all data sources are accumulated to obtain the total credibility value. The credibility value of each data source in the time-varying credibility set of multi-source data is then calculated with the total credibility value to establish an adaptive data source weight configuration. Based on the adaptive data source weight configuration, the current temperature, humidity, wind speed, and transformer load rate are extracted to obtain a pairing list. The initial value of the power grid unit fusion risk is obtained based on the pairing list. Based on the initial risk value of the integrated power grid unit, the risk range is determined by stepwise judgment against the multi-level risk thresholds preset for different power grid units such as switching equipment. Based on the result of the risk range determination, the corresponding level identifier is selected from the predefined risk level library to establish the power outage risk warning level.
[0016] The steps for obtaining the time-varying confidence set of multi-source data are as follows: Parse the generation timestamp of each data source with the current system time, and calculate the first... Time difference of each data source The predicted values of meteorological data and load data from the previous cycle are combined. Recorded values of actual power grid status Pairing is performed under the same time index to obtain the error sequence. And organize them into time difference and error sequences according to the data source identifier; The time-varying reliability value of the data source is calculated based on the time difference and error sequence. The calculation formula is as follows: ; in, For the first Time-varying confidence values of each data source. For the first The difference between the generation timestamp of each data source and the current system time. Based on the base time decay rate, For error sensitivity factor, For the first Data sources at time The difference between the predicted value and the actual value, Let be the number of terms in the error sequence. For the first The historical standard deviation of each data source corresponds to the actual state record of the power grid; Based on the time-varying confidence values, the time-varying confidence values of each data source are collected in the order of the data source identifier, outliers and missing items are removed, and a multi-source data time-varying confidence set is generated.
[0017] Specifically, the latest data generation timestamps are retrieved from the interfaces of various data source systems. For example, the timestamp "2023-10-26 14:00:00" is obtained from the meteorological information service interface, and the timestamp "2023-10-26 14:15:00" is obtained from the power grid load forecasting system. Simultaneously, the current system standard time, such as "2023-10-26 14:30:00", is obtained. By calculating the time difference between these two timestamps, the time difference for each data source is obtained. For example, the time difference of the meteorological data source is 30 minutes, and the time difference of the load data source is 15 minutes. Then, the actual status records of the previous complete cycle, such as the past 24 hours, are extracted from the power grid historical status database. The recorded values include the actual line temperature, humidity, wind speed, and transformer load rate, recorded every 15 minutes. Additionally, predicted values for the same time period are extracted from historical forecast archives of various data sources. To align each time index point, for example, the actual temperature record of 30.5℃ at 10:00 AM is paired with the predicted temperature of 31.0℃ published by the meteorological data source at that time. If the update frequency of the data source is inconsistent with the actual recording frequency, for example, the meteorological data is updated once an hour, while the actual power grid status is recorded once every 15 minutes, then linear interpolation is used to fill in the missing predicted values. Specifically, the predicted values at 9:00 and 10:00 are used to calculate the predicted values at 9:15, 9:30, and 9:45, ensuring that all time points have corresponding predicted and actual values. Then, the difference between each paired data point is calculated one by one to form an error sequence. For example, for temperature prediction from a meteorological data source, a temperature error sequence containing 96 data points (24 hours x 4 points / hour) is obtained, and its numerical sequence example is as follows: Finally, the calculated time difference With the corresponding error sequence Based on the unique identifier of the data source, such as "WeatherSource-A" or "LoadForecast-B", structured association storage is performed to form time difference and error sequences.
[0018] In the formula for calculating the time-varying confidence value, the basic time decay rate is used. and time difference The product is used to construct the basic timeliness penalty, ensuring that the credibility of any data decreases over time, while introducing a dynamic adjustment term driven by historical prediction errors. This item standardizes the root mean square error of the data source over a period of time and then applies it using an error sensitivity factor. By amplifying or reducing its impact and adding it to the base decay rate, the data source with the worst historical prediction record will have its credibility decay faster over time. This avoids applying a one-size-fits-all fixed weight or decay rate to all data sources, thus enabling a smarter focus on data that is both timely and accurate when assessing the risk of power outages.
[0019] For the first The difference between the generation timestamp of a data source and the current system time reflects the timeliness of the data; the larger the value, the older the data. It is obtained by directly reading the current time from the system clock and parsing the generation timestamp of the data source from the received data packet metadata. The difference between the two is then calculated. To ensure consistency of the formula's dimensions... The unit must be consistent with the base time decay rate. The calculation units are matched. In this example, the hour (h) is used as the time unit. For example, if the current system time is 15:00 and the received meteorological data is generated at 13:00, then... The calculated value is 2.0 h.
[0020] The base time decay rate is a quantity that is the reciprocal of time (e.g., h). It defines the inherent rate at which data reliability decays over time, without considering prediction errors. This parameter is set based on the "half-life" of a specific type of data. This refers to the time required for the data value to decrease by half. The half-life is determined based on domain experience. For example, for rapidly changing load forecast data, the half-life might be set to 4 hours, while for relatively stable meteorological data, the half-life could be set to 8 hours. The calculation formula is as follows: ,in The unit must be with The units should be consistent; for example, for meteorological data sources, their information half-life should be set. If it is 8 hours, then .
[0021] The error sensitivity factor is a dimensionless adjustment coefficient used to control the impact of historical prediction errors on the rate of reliability decay. Its value is set to balance the weight of timeliness and accuracy in reliability assessment. The setting process requires optimization based on historical data. Specifically, it involves selecting power outage event records for a historical time period and the corresponding multi-source data, and setting a set of... Candidate values, for example The reliability of each candidate value across all data sources is calculated and then substituted into the subsequent risk assessment model to obtain a series of risk prediction results. These prediction results are compared with actual power outage event records to calculate performance metrics such as prediction accuracy, recall, or F1 score. The model that optimizes the performance is then selected. The value is used as the final parameter, for example, by testing to find that when At that time, the risk warning score was the highest in F1, so it was ultimately determined. The value is 0.5.
[0022] This represents the number of terms in the error sequence. For example, for load forecast data updated every 15 minutes, if you want to assess its accuracy over the past 24 hours, then... The calculated value is .
[0023] For the first Data sources at time The difference between the predicted and actual values is obtained directly from the error sequence formed in the preceding steps. Its dimensions are consistent with the predicted physical quantity. For example, for temperature prediction, its unit is degrees Celsius (°C).
[0024] For the first Each data source corresponds to the historical standard deviation of the actual power grid status records. This parameter is used to standardize the prediction error, and its dimensions must be consistent with the corresponding error sequence. The dimensions are exactly the same, and the calculation method is to extract the actual state records of the power grid over a long time span (e.g., the past year). For example, collect temperature records for all moments in the past year to form a large dataset, and then apply the standard deviation formula to calculate it. For example, by calculating the power grid temperature records for the past year, its historical standard deviation can be obtained. It is 8.5℃.
[0025] Calculations based on parameters: Using a certain meteorological data source (set as the first data source), Taking (e.g.,) time-varying confidence values as an example. The calculations ensure consistency in the units of all parameters.
[0026] The parameters obtained are as follows: Time difference h.
[0027] Basic time decay rate h .
[0028] Error sensitivity factor (Dimensionless).
[0029] Number of terms in the error sequence (Dimensionless).
[0030] Obtain the temperature prediction error sequence for the past 24 hours from this data source using time difference and error sequences. Its unit is ℃. Calculate its root mean square error (RMSE): ; The calculated root mean square error is 1.6 ℃.
[0031] The first data source corresponds to the historical standard deviation of the actual state records of the power grid. ℃.
[0032] First, calculate the dimensionless error penalty term: ; Substituting the above values into the main formula, calculate the dimensionless parameter of the exponent: ; ; ; ; ; Finally, calculate the credibility value. : ; The results indicate that the time-varying reliability of this meteorological data source at the current moment is 0.8274. The closer the value is to 1, the higher the reliability, and vice versa. This result of 0.8274 reflects that the data source was generated 2 hours ago (with a certain time delay), and its prediction accuracy over the past 24 hours is a root mean square error of 1.6℃ (which is acceptable compared to the historical fluctuation of 8.5℃).
[0033] Based on the time-varying confidence values of each data source calculated in the previous step, we first iterate through all active data source channels and read the time-varying confidence value calculated for each data source one by one. For example, a confidence value of 0.8274 is obtained from weather source A, 0.9135 from load forecast source B, and 0.7550 from another weather source C. These confidence values are bound to their unique data source identifiers (e.g., "WeatherSource-A", "LoadForecast-B", "WeatherSource-C") to form a temporary key-value pair list. Next, outlier and missing entries are removed from this list. Missing entries are determined by data sources that failed to calculate a confidence value within a predetermined calculation period. For example, if data source D fails to provide data due to network interruption, its confidence value is empty, and the record will be directly removed. Outlier determination is based on a preset... The effective confidence interval has a lower threshold of 0.01. This threshold is set based on the fact that a confidence level below this value indicates extremely high latency or significant historical errors in the data source, rendering the data unreliable and potentially introducing noise into the calculation. The upper threshold is set to 1.0. Any calculation result exceeding 1.0 is considered a program error or data anomaly and must be removed. For example, if data source E calculates a confidence value of -0.1, the record is removed because it is below the lower threshold. After the removal operation, the remaining valid confidence values and their corresponding data source identifiers are sorted in a predefined fixed order, such as [load forecast source, meteorological source A, meteorological source C], to generate a multi-source data time-varying confidence set.
[0034] The steps to obtain the total credibility score are as follows: Based on the time-varying confidence set of multi-source data, the time-varying confidence value of each data source is read one by one. Traversal pointers are established according to the fixed order of the data source identifiers. Each time-varying confidence value is written into the cumulative register area in turn, and the index position of the data source identifier is recorded at the same time. After the end pointer verification is completed, the sum of all written values is obtained to get the total confidence value.
[0035] Specifically, based on the multi-source data time-varying confidence set generated in the aforementioned steps, this set contains the identifiers of each data source and their corresponding time-varying confidence values. For example, a set containing three data sources: {("Load Forecasting Source B", 0.9135), ("Weather Source A", 0.8274), ("Weather Source C", 0.7550)}, a floating-point register for accumulation is initialized, with its initial value set to 0.0. Simultaneously, a fixed-order list of data source processing pre-configured in the system is used, for example, ["Load Forecasting Source B", "Weather Source A", ...]. ["Weather Source C"] is used to set the order of the traversal operations. A traversal pointer is set to point to the first data source identifier "Load Forecasting Source B" in the order list. Based on this identifier, the corresponding time-varying confidence value of 0.9135 is searched and read in the multi-source data time-varying confidence set. This value is added to the current value of 0.0 in the accumulation register, and the value of the register is updated to 0.9135. At the same time, the index position of "Load Forecasting Source B" in the fixed order list is recorded as 0. Then, the traversal pointer is moved to the next data source identifier "Weather Source A" in the list, and the above search and accumulation operations are repeated. The confidence value of 0.8274 is read and added to the current value of 0.9135 in the accumulation register. The current value of the register is added to 0.9135, updating the register value to 1.7409. The index position of "Weather Source A" is recorded as 1. The pointer is then moved to "Weather Source C", its confidence value of 0.7550 is read, and added to the current value of 1.7409 in the register to obtain a new cumulative value of 2.4959. Its index position is recorded as 2. When the traversal pointer moves to the end of the fixed-order list, the end pointer is checked to confirm that all data source identifiers in the list have been traversed and calculated without any omissions. After confirmation, the final value of 2.4959 stored in the cumulative register is output as the calculation result to obtain the total confidence value.
[0036] The steps to obtain the adaptive data source weight configuration are as follows: Based on the total confidence value, the multi-source data time-varying confidence set is traversed again in a fixed order according to the data source identifier, the corresponding time-varying confidence value is read, and a division operation is performed to obtain the normalized value. Each normalized value is written into the same-order result sequence according to the index position to generate a normalized weight sequence for transmission line fault prediction. Based on the normalized weight sequence for transmission line fault prediction, the one-to-one correspondence between the sequence items and the data source identifier is checked, missing records are removed and index gaps are filled, and the data is encapsulated into a key-value pair structure in a fixed order to form an adaptive data source weight configuration.
[0037] Specifically, using the total confidence value calculated in the previous step, which is 2.4959, the multi-source data time-varying confidence set is called again. This set contains {("Load Forecast Source B", 0.9135), ("Weather Source A", 0.8274), ("Weather Source C", 0.7550)}. Simultaneously, an empty sequence structure is initialized to store the normalized values obtained in subsequent calculations. The entire calculation process strictly follows the previously defined fixed order of data source identifiers, i.e., ["Load Forecast Source B", "Weather Source A", "Weather Source C"]. First, the first data source in the sequence list, "Load Forecast Source B", is processed. Its corresponding time-varying confidence value of 0.9135 is read from the set. Then, a normalized division operation is performed, dividing this confidence value by the total confidence value. The calculation process is as follows: The normalized value is approximately 0.3660. Next, according to the index position 0 of "Load Forecast Source B" in the fixed-order list, the calculated normalized value of 0.3660 is stored in the 0th position of the result sequence. Then, the second data source in the list, "Meteorological Source A," is processed, its time-varying confidence value of 0.8274 is read, and the same division operation is performed. The result is approximately 0.3315, and this value is stored in the first position of the result sequence. Finally, the third data source, "Meteorological Source C," is processed, and its confidence value of 0.7550 is read. The calculation is then performed. The result is approximately 0.3025, which is stored in the second position of the result sequence. After all data sources have been processed, the content of that result sequence is [0.3660, 0.3315, 0.3025]. This sequence is the normalized weight sequence for transmission line fault prediction.
[0038] Based on the normalized weight sequence for transmission line fault prediction generated in the previous step, with the content [0.3660, 0.3315, 0.3025], and combined with the fixed-order list of data source identifiers ["Load Forecasting Source B", "Weather Source A", "Weather Source C"] used as the calculation benchmark, the final configuration generation process is initiated. First, a consistency check is performed to verify that the length of the normalized weight sequence is equal to the length of the data source identifier list. In this example, the sequence length is 3, and the list length is also 3, so they match and the check passes. Next, the system will traverse this identifier list and weight sequence, performing pairing and encapsulation operations. This process aims to construct a key-value data structure, where the data source identifier is the key and its corresponding normalized weight is the value. During the traversal, possible anomalies are handled. For example, if a non-numeric (NaN) or infinite (Infinite) value is found at a certain position in the normalized weight sequence... The record for (ty) indicates an error in the previous calculation. Therefore, the entire data source entry corresponding to that position, including its identifier and the erroneous weight value, will be removed and not included in the final configuration. Meanwhile, the index hole completion operation here refers to the fact that if an intermediate entry is removed, the final key-value pair structure itself, due to its unordered nature or key-based access characteristics, naturally "fills in" the discontinuities in the index. This is because it does not rely on continuous numeric indexes. After confirming that all weight values are valid floating-point numbers, encapsulation begins. The first element of the identifier list, "Load Prediction Source B," and the first element of the weight sequence, 0.3660, are combined to form the first key-value pair: "Load Prediction Source B." 0.3660, then take the second element "Weather Source A" and 0.3315, combine them to form: "Weather Source A": 0.3315, and finally take the third element "Weather Source C" and 0.3025, combine them to form: "Weather Source C": 0.3025. All these key-value pairs are put together to form a complete data object, which is the adaptive data source weight configuration.
[0039] The steps to obtain the pairing list are as follows: Based on the adaptive data source weight configuration, extract the current temperature value, current humidity value, current wind speed value, and current transformer load rate value according to the data source identifier, verify the consistency of the time index, and complete the unit conversion to obtain the pairing list.
[0040] Specifically, based on the adaptive data source weight configuration formed in the previous step, which includes the identifiers of each data source and their corresponding normalized weights, such as {"Load Forecasting Source B": 0.3660, "Meteorological Source A": 0.3315, "Meteorological Source C": 0.3025}, the feature data extraction process for the current moment is initiated. This process iterates through each data source identifier in the configuration and sends real-time data requests to the corresponding data interfaces. For example, it requests the latest meteorological observation data from the application programming interfaces (APIs) of "Meteorological Source A" and "Meteorological Source C," while simultaneously requesting the real-time load rate of the currently monitored transformer from the system interface of "Load Forecasting Source B." The obtained raw data might be {"source": "Meteorological Source A", "timestamp": "1698303600", "temp_f": 95.36, "humidity_pct": 78.5, "wind_kph": 15.3} and {"source": "Load Forecasting Source B", "timestamp": "1698303595", "load_mw": 0.3660, "timestamp": 0.3315, "load_mw": 0.3025}. 8.89, "capacity_mw": 10.0}, After collecting all data, a time index consistency check is performed. A current system timestamp is set as the baseline, and a time tolerance threshold of 60 seconds is defined. This threshold is based on the real-time requirement of power outage risk assessment, which requires that all input data must be generated within one minute before the assessment is initiated. For each returned data, its timestamp is compared with the baseline timestamp. If the absolute value of the time difference is less than or equal to 60 seconds, the data is considered valid; otherwise, it is discarded. After verification, a unified unit conversion is performed, converting all data to internal standard units. For example, the calculation for converting Fahrenheit temperature (°F) to Celsius (°C) is... The calculation of converting wind speed in kilometers per hour (kph) to meters per second (m / s) is as follows: The actual load of the transformer is calculated as a percentage of the rated capacity. After conversion, the data from meteorological source A becomes {temperature: 35.2°C, humidity: 78.5%, wind speed: 4.25 m / s}, and the data from load forecasting source B becomes {load rate: 88.9%}. Finally, all the verified and converted valid data, along with their data source identifiers and the time-varying confidence values obtained from previous steps, are structurally combined to obtain a pairing list.
[0041] The steps for obtaining the initial value of the risk of grid unit integration are as follows: Based on the pairing list, the initial value of the grid unit integration risk is calculated using the following formula: ; in, This is the initial value for the risk of grid unit integration. The total number of feature types, with a value of four. For the first The feature risk contribution parameter of the item. For the first The number of data sources corresponding to each feature For the first Time-varying confidence values of each data source. For the first Item feature in the first The current time feature values of each data source. For the first The nonlinear risk mapping function of the feature is expressed as follows: , For the first Risk sensitivity parameter of the feature, For the first Risk threshold of the feature and its relationship with Same dimension.
[0042] Specifically, in the formula for calculating the initial value of grid unit integration risk, the weighted average term... Through a nonlinear mapping function of Sigmoid form The original physical quantity (e.g., temperature value) is converted into a risk metric between 0 and 1. This simulates the reality that the impact of most physical factors on risk is not linear, but rather increases sharply near a certain critical point. The time-varying confidence value calculated in the previous steps is then applied. Weighted averaging, based on these factors, ensures that information sources with more accurate historical predictions and more timely data updates dominate risk assessments for single characteristics. Finally, the outer summation... The fusion risk values of various characteristics are ranked according to their overall contribution to the power outage event. Perform linear combinations.
[0043] To account for the total number of feature types, four environmental and operational features most closely associated with power outage risk were selected: temperature, humidity, wind speed, and transformer load rate. The value of is a fixed integer 4.
[0044] For the first The feature risk contribution parameter is a dimensionless weighting coefficient, and all... The sum of these values is 1, quantifying the relative importance of each feature to the final power outage risk. This parameter is determined based on statistical analysis and machine learning modeling of historical power outage data. The specific steps are as follows: collect a historical dataset containing the above four features and whether a power outage occurred at the corresponding time point (0 or 1); use this dataset to train a logistic regression model, the model taking the form of… ,in This represents the probability of a power outage. The coefficients learned by the model are the feature risk contribution parameters. This is obtained by normalizing the absolute values of these coefficients, using the following formula: For example, by training a model using data from the power grid of a certain region over the past three years, the model coefficients obtained are as follows: Then the corresponding contribution parameter is calculated as follows: Similarly, we can obtain , , .
[0045] For the first The number of data sources corresponding to a feature. For example, if two weather stations currently provide valid temperature data, then for the temperature feature ( ), .
[0046] For the first The time-varying reliability values of each data source are derived from the calculation results of the "multi-source data time-varying reliability set" in the previous steps, reflecting the timeliness and historical accuracy of each data source. In this example, the previously calculated values are used.
[0047] For the first Item feature in the first The current moment characteristic value of each data source is extracted from the "pairing list" generated in the first stage of this step. Its dimension is consistent with the original physical quantity, such as temperature in degrees Celsius (°C).
[0048] For the first The risk sensitivity parameter of a feature, the dimension of which is the reciprocal of the physical dimension of the corresponding feature (e.g., °C). ,% This is used to control the steepness of the risk mapping function curve. The higher the value, the more drastic the transition from a risk that has never occurred to one that has definitely occurred. Its setting is based on expert knowledge and the physical model of equipment failure. Specifically, domain experts define two risk points and one low-risk point. and a high-risk point ,in Represent the risk probability, then solve the simultaneous equations. Determine by a system of equations and For example, regarding transformer load factor, experts define the risk as 10% when the load factor reaches 85%. When it reaches 95%, the risk is 90% ( Then a system of equations can be set up. Solving for % and .
[0049] For the first The risk threshold of a feature, its dimensions are... The same value represents a 50% risk probability threshold. This value is mainly determined based on industry safety regulations, equipment design specifications, and high-incidence fault ranges from long-term operational statistics. For example, the maximum allowable operating temperature for a certain type of transmission line is 40℃. Exceeding this temperature will accelerate insulation aging. Therefore, a risk threshold for temperature characteristics can be set. ℃.
[0050] Calculations based on parameters: Set the simulation scenario and parameters: feature These represent {temperature, humidity, wind speed, and load rate}, respectively.
[0051] Retrieve data from the pairing list: Meteorological source A ( ): , , , .
[0052] Meteorological source C ( ): , , , .
[0053] Load source B ( ): , .
[0054] Parameter settings: ; ; ; Calculate the fusion risk value for each feature.
[0055] Calculate the initial value of the risk of grid unit integration : ; ; The results indicate that the initial value of the fusion risk of this power grid unit is 0.1796 at the current moment. The higher the value, the higher the risk of power outage. The current value of 0.1796 indicates that the overall risk level is low. The main risk contributions come from humidity and transformer load rate, while the risk contributions from temperature and wind speed are negligible.
[0056] The steps for determining the risk zone attribution are as follows: Based on the initial risk value of grid unit integration, a multi-level risk threshold table is retrieved according to the grid unit type. The thresholds are verified to be ordered and without repetition, and the unified boundary is left closed and right open, forming a multi-level risk threshold sequence for grid units. Based on the multi-level risk threshold sequence of the power grid unit, the initial risk value of the power grid unit is compared step by step from the smallest to the largest. When the threshold hits the lower limit, it is recorded as the interval to which it belongs. When it is equal to the upper limit, it is moved up one level. When it exceeds the limit, it is assigned to the lowest or highest interval respectively, thus generating a risk interval assignment determination.
[0057] Specifically, based on the initial risk value of the power grid unit calculated in the previous step, which is 0.1796, and according to the specific type of the power grid unit currently being evaluated, such as "10kV overhead transmission line" or "box-type transformer", a multi-level risk threshold table matching the type is retrieved from a pre-built multi-level risk threshold library. This threshold table is constructed based on statistical analysis of historical fault data for this type of power grid unit. Specifically, a large number of historical fault cases are collected, and the initial risk value calculation result for each case at the moment before the occurrence is extracted to form a risk value distribution. Then, the percentile method is used to classify the risk level. For example, the 50th percentile (i.e., the median) arranged from smallest to largest in the risk value distribution is used as the dividing point between low and medium risk, the 85th percentile as the dividing point between medium and high risk, and the 98th percentile as the dividing point between high and extremely high risk. Taking "box-type transformer" as an example, the retrieved threshold table might be {"low risk": [0, 0.2)}. The thresholds in the table are first verified after retrieval. This verifies whether all thresholds are arranged in ascending order and whether there are any duplicate values. For example, the sequence 0, 0.2, 0.5, 0.8 is checked to ensure it is monotonically increasing and without repetition. After verification, the interval boundaries are standardized, meaning all intervals are uniformly left-closed and right-open. This means the lower bound of the interval is included, while the upper bound is not. Finally, these verified and standardized threshold boundary points are extracted to form an ordered numerical sequence, namely the multi-level risk threshold sequence of the power grid unit [0, 0.2, 0.5, 0.8].
[0058] Using the multi-level risk threshold sequence of the power grid unit generated in the previous step, whose content is [0, 0.2, 0.5, 0.8], and calling the initial risk value of the power grid unit fusion, which is 0.1796, the risk interval assignment determination process is initiated. This process is implemented through a loop comparison mechanism. Starting from the first element 0 of the threshold sequence, the initial risk value of the power grid unit fusion is compared with the thresholds in the sequence one by one. The comparison rule is to determine whether the initial risk value is greater than or equal to the current threshold and less than the next threshold. First, 0.1796 is compared with the first threshold 0 in the sequence. 0.1796 is greater than or equal to 0, satisfying the lower bound condition. Then, it is compared with the second threshold 0.2 in the sequence. 0.1796 is less than 0.2, satisfying the upper bound condition. Therefore, the initial risk value of the power grid unit fusion, 0.1796, hits the interval [0, 0.2) formed by the thresholds 0 and 0.2. At this time, the determination process immediately terminates and the interval [0, 0.2) is recorded as the interval to which it belongs. If a special case occurs, such as the initial risk value being exactly equal to the upper limit of a certain interval, for example, 0.2, then according to the preset rule of "shifting up one level when equal to the upper limit", the initial risk value will be assigned to the next interval, i.e., [0.2, 0.5). In addition, there is also a handling logic for out-of-bounds cases. If the initial risk value is less than the minimum threshold 0 in the sequence (which will not happen when the risk value is non-negative), it will be directly assigned to the lowest risk interval, i.e., [0, 0.2). If the initial risk value is greater than or equal to the maximum threshold 0.8 in the sequence, it will be assigned to the highest risk interval, i.e., [0.8, 1.0). After completing the above comparison and judgment logic, a clear assignment result is finally output, i.e., the risk interval assignment judgment.
[0059] The steps to obtain the power outage risk warning level are as follows: Based on the risk range attribution determination, a predefined risk level database is retrieved, and the corresponding level identifier is mapped according to the range code to form a power outage risk warning level.
[0060] Specifically, based on the risk range attribution determination generated in the previous step, the result is that the risk value of 0.1796 belongs to the range [0, 0.2). Next, a predefined risk level library is called for query matching. This risk level library is a static data structure that maps risk ranges to specific risk level identifiers. Its construction is based on national or industry early warning standards for power safety production, combined with the company's internal management regulations, assigning a unique range code and corresponding level identifier to each risk range. For example, the contents of this library may be as follows: {"[0, 0.2)":{"Code":"L1", "Identifier":"Blue Warning"}, "[0.2, 0.5)": {"Code":"L2", "Identifier":"Yellow Warning"}, "[0.5, 0.8)": {"Code":"L3", "Identifier":"Orange Warning"}, "[0.8, 1.0]": {"Code":"L4", When searching for the risk level database, the result of the risk range determination, i.e., the range "[0, 0.2)", is used as the query key. The database is searched to find the entry with the key "[0, 0.2)" and the corresponding level identifier "Blue Warning" is extracted from the entry. This "Blue Warning" is the final power outage risk warning level.
[0061] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention in any other way. Any person skilled in the art may make changes or modifications to the above-disclosed technical content to create equivalent embodiments that can be applied to other fields. However, any simple modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the protection scope of the present invention.
Claims
1. A power cut risk assessment method for multi-source data, characterized in that, The method comprises the following steps: Obtain the difference between the generation timestamp of each data source and the current system time, calculate the timeliness score, and then retrieve the deviation between the predicted values of meteorological data and load data in the last period and the actual state records of the power grid to calculate the time-varying credibility set of multi-source data; According to the time-varying credibility set of multi-source data, the credibility values of all data sources are accumulated to obtain a total credibility value, and the credibility value of each data source in the time-varying credibility set of multi-source data is respectively operated with the total credibility value to establish an adaptive data source weight configuration; According to the adaptive data source weight configuration, the temperature, humidity, wind speed and transformer load rate at the current time are extracted to obtain a pairing list, and the pairing list is calculated to obtain a power grid unit fusion risk initial value; According to the power grid unit fusion risk initial value, a step-by-step judgment is performed on the multi-level risk threshold preset for different power grid units such as switch devices to obtain a risk interval attribution determination, and according to the result of the risk interval attribution determination, a corresponding grade identifier is selected from a predefined risk grade library to establish a power outage risk warning grade.
2. The power cut risk assessment method for multi-source data according to claim 1, characterized in that, The acquisition step of the time-varying credibility set of multi-source data is: Parse the generation timestamp of each data source and the current system time, calculate the time difference of the first data source, pair the predicted values of meteorological data and load data in the last period with the actual state record values of the power grid at the same time index to obtain an error sequence, and arrange the time difference and error sequence according to the data source identifier; According to the time difference and error sequence, the time-varying credibility value of the data source is calculated; According to the time-varying credibility value, the time-varying credibility values of each data source are collected in the order of data source identifier, and abnormal values and missing items are removed to generate a time-varying credibility set of multi-source data.
3. The power cut risk assessment method for multi-source data according to claim 1, characterized in that, The acquisition step of the total credibility value is: According to the time-varying credibility set of multi-source data, the time-varying credibility value of each data source is read source by source, a traversal pointer is established in a fixed order of data source identifier, each time-varying credibility value is written into a cumulative register and the index position of the data source identifier is recorded at the same time, the sum of all written values is obtained after the end pointer verification is completed, and the total credibility value is obtained.
4. The method of claim 1, wherein, The acquisition step of the adaptive data source weight configuration is: According to the total credibility value, the time-varying credibility set of multi-source data is traversed again in a fixed order of data source identifier, the corresponding time-varying credibility value is read, a division operation is performed to obtain a normalized value, each normalized value is written into a same sequence result sequence according to the index position, and a normalized weight sequence for power transmission line fault prediction is generated; According to the normalized weight sequence for power transmission line fault prediction, the one-to-one correspondence relationship between the sequence items and the data source identifiers is checked, missing records are removed and index voids are filled, the fixed order is packaged into a key-value pair structure, and the adaptive data source weight configuration is formed.
5. The method of claim 1, wherein, The acquisition step of the pairing list is: According to the adaptive data source weight configuration, the current time temperature value, the current time humidity value, the current time wind speed value and the current time transformer load rate value are extracted according to the data source identification, the time index consistency is checked and the unit system conversion is completed to obtain a pairing list.
6. The method of claim 1, wherein, The power grid unit fusion risk initial value acquisition step is: According to the pairing list, the power grid unit fusion risk initial value is calculated.
7. The method of claim 1, wherein, The risk interval attribution determination acquisition step is: According to the power grid unit fusion risk initial value, a multi-level risk threshold table is retrieved according to the power grid unit type, the threshold is checked to be ordered and non-repetitive, the boundary is unified to be left-closed and right-open, and a power grid unit multi-level risk threshold sequence is formed; According to the power grid unit multi-level risk threshold sequence, the power grid unit fusion risk initial value is compared level by level according to the threshold from small to large, the lower bound is hit when the interval attribution is recorded, the upper bound is equal when the interval is moved up by one level, and the interval is out of range when it is respectively attributed to the lowest or highest interval, and a risk interval attribution determination is generated.
8. The method of claim 1, wherein, The power outage risk early warning level acquisition step is: According to the risk interval attribution determination, a predefined risk level library is searched, the interval code is mapped to the corresponding level identifier, and a power outage risk early warning level is formed.