Multi-temporal-spatial-scale data fusion household number accounting method during power failure of power distribution network
Through multi-spatial-scale data fusion technology, the problems of data islands, rough spatial-temporal particle size and misreporting false alarms in the household count calculation method during power outage in traditional distribution networks are solved, and high-precision power outage event accounting and prediction are achieved, providing more accurate data support for power grid operation and maintenance.
Patent Information
- Application Number
- CN202510270970.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-06-27
Smart Images

Figure CN120218766A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of power systems, and in particular, to a method for calculating the outage household-hours of a distribution network by fusing multi-temporal and multi-spatial scale data. Background Art
[0002] As an important part of the power system, the reliability of the distribution network directly affects the user's power consumption experience and the operation efficiency of the power grid enterprise. The outage household-hours is one of the core indicators for measuring the power supply reliability of the distribution network, which is used to quantify the impact scope and duration of the outage event on users. Accurately calculating the outage household-hours is of great significance for improving the power grid service quality, optimizing the operation and maintenance decision-making, and formulating the power grid planning.
[0003] However, there are many problems in the traditional method for calculating the outage household-hours, resulting in insufficient accuracy of the calculation results and difficulty in meeting the requirements of the efficient operation of modern distribution networks. Specifically, the existing methods mainly face the following technical challenges:
[0004] 1. Data islands and inconsistencies:
[0005] There are various monitoring and management systems in the distribution network (such as SCADA, AMI systems, user repair platforms, etc.), but the data of these systems cannot be effectively fused. Due to the diverse data sources and different formats, traditional methods often have problems of inconsistent information when integrating these data, which further leads to large deviations in the calculation results of the start time, recovery time of the outage event, and the affected scope of users.
[0006] 2. Coarse spatio-temporal granularity:
[0007] Traditional calculation methods usually analyze with a fixed spatio-temporal granularity, without fully considering the dynamic impacts of factors such as the status of distribution network equipment, user power consumption behavior, and meteorological conditions on the outage event. This rough spatio-temporal division method is difficult to accurately reflect the actual spatio-temporal distribution characteristics of the outage event, thus reducing the calculation accuracy.
[0008] 3. Problems of missed reports and false reports:
[0009] In remote areas or instantaneous fault scenarios, traditional methods are difficult to capture the occurrence and recovery processes of outage events in a timely manner, which easily leads to missed reports or false reports. In addition, due to the lack of in-depth mining and intelligent analysis capabilities for multi-source data, traditional methods have poor adaptability to complex outage scenarios, further exacerbating the calculation errors.
[0010] 4. Lack of intelligent analysis means:
[0011] Existing methods mainly rely on manual reports and rule engines, and lack the effective use of big data and artificial intelligence technologies. This makes it difficult to achieve automation and high accuracy in predicting power outage events, assessing the scope of user impact, and filling in data missing areas.
[0012] In summary, the traditional method of calculating the number of households during power outages has obvious shortcomings in data fusion, spatiotemporal modeling, event capture and intelligent analysis, and a new technical solution is urgently needed to solve the above problems. Summary of the invention
[0013] In order to solve the above-mentioned problems of the prior art, the present invention provides a household number accounting method for power outages in distribution networks based on multi-temporal and spatial scale data fusion, aiming to construct a household number accounting method for power outages in distribution networks based on multi-temporal and spatial scale data fusion by introducing multi-source data fusion, spatiotemporal grid modeling, machine learning algorithms and spatial interpolation technology, so as to improve the calculation accuracy, reduce missed reports and false alarms, and provide more accurate data support for reliability assessment and operation and maintenance optimization of distribution networks.
[0014] The technical solution adopted by the present invention is:
[0015] A method for calculating the number of households in a power outage in a distribution network using multi-temporal and spatial scale data fusion, comprising the following steps:
[0016] Step 1 Multi-source data fusion framework: By collecting, cleaning and integrating multi-source heterogeneous data, a high-quality spatiotemporal data set is constructed to provide a basis for subsequent fusion analysis;
[0017] Step 2: Spatiotemporal grid modeling: Based on the spatiotemporal data set, K-means clustering and dynamic time granularity division are used to perform refined spatiotemporal grid modeling of the distribution network and extract key features;
[0018] Step 3: Intelligent accounting algorithm: Through time synchronization correction and mutation point detection, the time range and user impact of power outage events are accurately corrected;
[0019] Step 4 Multidimensional analysis report: Combine machine learning to predict the number of households at the time of potential power outages, and use spatial interpolation to fill in missing data to improve accounting completeness and accuracy;
[0020] Step 5: Household accounting report and decision support during power outages: Generate multi-dimensional visualization reports and push high-risk area data to provide decision support for power grid operation and planning.
[0021] Furthermore, in step 1, the collected data includes real-time and historical data obtained from the SCADA system, AMI system, equipment inventory, historical power outage records, and user repair reporting platform.
[0022] Furthermore, in step 1, the process of cleaning and integrating multi-source heterogeneous data includes:
[0023] Data preprocessing: Denoise the collected data through the Z-score algorithm, align the time based on the timestamp synchronization, and perform normalization processing to ensure the consistency and availability of the data.
[0024] Furthermore, in step 2, first, use the K-means clustering algorithm to perform clustering analysis on the distribution network equipment to determine the regional boundaries such as feeders, substations, and user groups; then, dynamically adjust the time granularity to the minute level, hour level, or daily level according to the frequency and duration of power outage events to improve the time resolution; finally, extract key features from each spatio-temporal unit, and the key features include: equipment failure rate, user density, and historical power outage frequency, so as to describe the impact of power outages.
[0025] Furthermore, the impact of power outages in each spatio-temporal unit is described by the following formula:
[0026]
[0027] In the formula, 影响范围(i,t) represents the power outage impact range of the i-th spatial unit within time t; represents the cumulative sum of all relevant features, where n is the total number of features or regions; ω j represents the weight coefficient of the j-th feature, which is used to measure the importance of this feature to the power outage impact; d ij represents the spatial distance between the i-th spatial unit and the j-th feature; failure rate(j) represents the failure rate of the j-th device or region, which is usually a probability value between 0 and 1; user(j) represents the number of users served by the j-th device or region.
[0028] Furthermore, in step 3, the time correction of power outage events is performed through the time synchronization verification algorithm, which uses the data interpolation method to synchronize the time data from different sources and eliminate the timing error;
[0029] The time correction function of power outage events is:
[0030] T 修正 = T 原始 + ΔT 偏差
[0031] In the formula, T 修正 is the corrected power outage event time; T 原始 is the power outage event time obtained from the original data; ΔT 偏差 is the time deviation, which represents the time error caused by different data sources or acquisition delays.
[0032] Further, in step 3, the influence scope of the power outage event needs to be judged in combination with the user's electricity consumption behavior; there are obvious mutation points in the user's electricity consumption curve when the power outage occurs or is restored. Through the mutation point detection algorithm, these abnormal changes are accurately identified, so as to determine the influence scope of the power outage event;
[0033] For a single user, the mutation point detection can determine the start and end times of the power outage; for multiple users, by comprehensively analyzing the mutation points of all users, the overall influence scope of the power outage event can be inferred.
[0034] Further, the mutation point detection algorithm includes the CUSUM method and cluster analysis;
[0035] The CUSUM method detects mutation points by calculating the cumulative change of the electricity consumption curve, and the calculation formula is:
[0036] S t =max(0,S t-1 +(x t -μ)-k)
[0037] In the formula, S t is the cumulative sum; x t is the electricity consumption at the current moment; t is the index of the data point, only representing the data order and not requiring manual adjustment; μ is the mean value of the electricity consumption; k is the parameter for controlling the sensitivity;
[0038] When S t exceeds the preset threshold, it is determined as a mutation point;
[0039] Cluster analysis groups the characteristics of the user's electricity consumption curve to identify abnormal patterns that are significantly different from the normal electricity consumption pattern. The users in the abnormal category are the users affected by the power outage event.
[0040] Further, in step 4, machine learning uses a random forest model to predict the number of household hours of the power outage event, and its prediction function is:
[0041]
[0042] In the formula: y represents the predicted number of household hours of the power outage event; f(x) represents the prediction function of the random forest model; N represents the number of decision trees in the random forest; α i represents the weight coefficient of the i-th decision tree; T i (X) represents the prediction result of the i-th decision tree for the input feature x;
[0043] For filling the data missing area, the Kriging interpolation method is adopted, and its interpolation formula is:
[0044]
[0045] In the formula: Z(x0) represents the interpolation result of the target point x0; μ(x0) represents the global trend or mean of the target point x0; n represents the number of known data points; λ i represents the weight coefficient of the i-th known point; Z(x i ) represents the actual observed value of the i-th known point; μ(x i ) represents the global trend or mean of the i-th known point; (Z(x i ) - μ(x i )) represents the deviation between the actual observed value of the i-th known point and its trend value.
[0046] Advantages of the present invention:
[0047] The present invention constructs a method for calculating the power outage time - household number based on multi - spatio - temporal scale data fusion by introducing multi - source data fusion, spatio - temporal grid modeling, intelligent accounting algorithms, machine learning prediction, and spatial interpolation techniques, and has the following remarkable beneficial effects:
[0048] 1. Improve the accuracy of power outage time - household number calculation:
[0049] Multi - source data fusion: By integrating multi - source heterogeneous data such as SCADA, AMI systems, equipment ledgers, historical power outage records, and user repair data, the problems of data islands and inconsistencies in traditional methods are solved, ensuring the comprehensiveness and consistency of data.
[0050] Fine - grained spatio - temporal modeling: Using the K - means clustering algorithm to partition space and dynamically adjusting the time granularity according to power outage events significantly improves the resolution of spatio - temporal division and can more accurately reflect the actual distribution characteristics of power outage events.
[0051] Mutation point detection and time correction: By accurately identifying the mutation points of user power consumption curves through the CUSUM method and cluster analysis, and combining with the time synchronization verification algorithm to eliminate timing errors, the calculation accuracy of the start and end times and the affected scope of power outage events is greatly improved.
[0052] 2. Reduce missed reports and false reports:
[0053] Deep mining of multi - source data: Through intelligent analysis of multi - source data, the ability to capture remote areas or instantaneous fault scenarios is enhanced, effectively reducing missed reports and false reports in traditional methods.
[0054] Mutation point detection algorithm: Accurately judge the abnormal changes in user power consumption behavior, avoiding misjudgment caused by human reporting delays or data missing, and further reducing the calculation error.
[0055] 3. Enhance adaptability to complex scenarios:
[0056] Dynamic time granularity division: The time resolution is dynamically adjusted according to the frequency and duration of power outage events, so that the method can flexibly adapt to the power outage event analysis needs in different scenarios.
[0057] Machine learning prediction: The random forest model is used in combination with equipment status, historical fault data, and user distribution characteristics to predict potential power outages and their impact range, improving the adaptability and robustness of the method in complex scenarios.
[0058] 4. Accurately fill in the missing data areas:
[0059] Kriging interpolation method: Kriging interpolation technology is used to fill in the data missing areas, combining the spatial correlation and trend information of known points to generate more accurate estimates, making up for the accounting deviation caused by missing data in traditional methods.
[0060] 5. Provide comprehensive decision support:
[0061] Multi-dimensional visualization report: Generates heat maps of the number of households at the time of power outage by region, device type, and time period, and displays them through the web front end to intuitively present the impact and severity of the power outage.
[0062] High-risk area push: Automatically push high-risk area data to the operation and maintenance management system, providing accurate data support for fault repair, equipment transformation and power grid planning, and improving the efficiency of power grid operation and maintenance and planning.
[0063] 6. Improve the reliability assessment and optimization capabilities of distribution networks:
[0064] High-quality data support: Through multi-source data fusion and cleaning processing, a high-quality spatiotemporal data set is constructed, providing a solid data foundation for distribution network reliability assessment.
[0065] Intelligent analysis methods: The introduction of machine learning and spatial interpolation technology has realized the automatic prediction and accurate calculation of power outage events, significantly improving the intelligent management level of the distribution network.
[0066] 7. Reduce manual intervention and improve work efficiency:
[0067] Automated process: From data collection and preprocessing to intelligent accounting and decision support, the entire process is highly automated, reducing reliance on manual reporting and rule engines, and greatly improving work efficiency.
[0068] Real-time and high efficiency: Through dynamic time granularity division and real-time monitoring data analysis, the accounting and report generation of power outage events can be completed in a short time, meeting the needs of efficient operation of modern distribution networks.
[0069] 8. Promote the intelligent development of the distribution network:
[0070] Technological innovation: The present invention deeply integrates big data, artificial intelligence, and spatio-temporal analysis technologies, promoting the technological innovation of the calculation method for power outage events in the distribution network and providing important technical support for the development of smart grids.
[0071] Wide application prospects: This method is not only applicable to the calculation of power outage hours in the current distribution network but can also be extended to the reliability assessment and optimization fields of other power systems, with wide application value.
[0072] In summary, through multi-temporal and spatial scale data fusion, refined spatio-temporal modeling, intelligent calculation algorithms, and machine learning technologies, the present invention significantly improves the accuracy and efficiency of calculating power outage hours in the distribution network, reduces false positives and false negatives, enhances adaptability to complex scenarios, and provides comprehensive decision-making support for grid operation and planning. This method not only solves the core problems in traditional technologies but also promotes the improvement of the intelligent management level of the distribution network, with important practical application value and broad development prospects. Brief Description of the Drawings
[0073] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments recorded in the present invention. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings.
[0074] Figure 1 It is a schematic flow chart of the method for calculating power outage hours in the distribution network based on multi-temporal and spatial scale data fusion of the present invention. Detailed Embodiments
[0075] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0076] In view of the problems existing in traditional accounting methods, such as data islands, rough spatio-temporal granularity, missing and false reports, and insufficient intelligence. This embodiment provides a method for calculating the power outage household-hours of a distribution network through multi-spatio-temporal scale data fusion. This method constructs a method for calculating the power outage household-hours of a distribution network based on multi-spatio-temporal scale data fusion by introducing multi-source data fusion, spatio-temporal grid modeling, machine learning algorithms, and spatial interpolation technology, so as to improve the calculation accuracy, reduce missing and false reports, and provide more accurate data support for the reliability assessment and operation and maintenance optimization of the distribution network.
[0077] As Figure 1 shown, the method for calculating the power outage household-hours of a distribution network through multi-spatio-temporal scale data fusion includes the following steps:
[0078] Step 1 Multi-source data fusion framework:
[0079] This step aims to construct a high-quality spatio-temporal dataset by collecting and integrating multi-source heterogeneous data, providing a basis for subsequent analysis. It mainly includes three links: data collection, preprocessing, and unified storage:
[0080] Data collection: Obtain real-time and historical data from multiple channels such as the SCADA system, AMI system, equipment ledger, historical power outage records, and user repair platforms. The data includes real-time monitoring data, equipment ledger data, user repair data, historical power outage records, etc. These data are collected from each data source through communication protocols and uploaded to a unified data warehouse.
[0081] Data preprocessing: Denoise, time-align, and normalize the collected data to ensure the consistency and availability of the data.
[0082] Specifically, for the collected data, detect and remove outliers through the Z-score algorithm; judge whether a data point is an outlier by calculating the standard deviation distance between the data point and the mean value;
[0083]
[0084] where z is the standardized value, x is the original data point; μ is the mean value of the dataset; σ is the standard deviation of the dataset.
[0085] The denoising process is as follows: First, calculate the mean value μ and the standard deviation σ of the collected dataset; then, calculate the Z-score value z of each data point x; then, set a threshold: usually, data points with |z| > 3 are set as outliers, that is, points deviating from the mean value by more than 3 times the standard deviation; finally, remove outliers and remove or mark as invalid the data points that satisfy |z| > 3 from the dataset.
[0086] Suppose the current data of a certain device is [10, 12, 11, 50, 9, 13]: then the mean μ = 17.5 and the standard deviation σ = 16.8 are calculated. Calculate the Z-score for each data point: z1 ≈ -0.45, z2 ≈ -0.33, z3 ≈ -0.39, z4 ≈ 1.93, z5 ≈ -0.51, z6 ≈ -0.27; set the threshold |z| > 3, and it is found that all points do not exceed the threshold, so there are no outliers.
[0087] Since there may be deviations in the timestamps of multi-source data, such as the sampling times of the SCADA system and the AMI system being out of sync, it is necessary to align the data in time; in this embodiment, the synchronization process is as follows: First, extract the timestamp information from each data source, such as the timestamp of the SCADA system, the sampling time of the AMI system, etc.; then, determine a unified time reference, such as UTC time, and convert all timestamps to this reference; finally, for data with inconsistent timestamps, linearly interpolate to generate data for the missing time points.
[0088] The linear interpolation formula is:
[0089]
[0090] In the formula, yt is the interpolation result at the target time point; t1, t2 are known time points; yt1, yt2 are the data values corresponding to the known time points.
[0091] Suppose the current data of the SCADA system is [(10:00, 10), (10:05, 12)], and the voltage data of the AMI system is [(10:02, 220), (10:07, 225)]: unify the timestamps to the minute level, for example, 10:00 corresponds to 0 minutes, 10:05 corresponds to 5 minutes; interpolate the SCADA data at 10:02, y 10:02 = 10.8.
[0092] Normalization is to eliminate the differences in the numerical ranges of different data sources and facilitate subsequent analysis;
[0093] The normalization formula is:
[0094]
[0095] In the formula, x' is the normalized value; x is the original data value; x min 、x max are the minimum and maximum values in the dataset.
[0096] For each data type, such as: current, voltage, calculate x min and x max respectively; then calculate the normalized value x' for each data point x; finally, store the normalized data as a new dataset.
[0097] Suppose the current data of a certain device is [10, 12, 11, 9, 13]: then x min = 9, x max = 13; After normalization calculation: x1' = 0.25, x'2 = 0.75, x'3 = 0.5, x'4 = 0, x'5 = 1.
[0098] Build a unified data warehouse: Store the cleaned data as a structured spatio-temporal dataset for subsequent fusion analysis.
[0099] Step 2 Spatio-temporal grid modeling:
[0100] In this step, by spatio-temporally partitioning the distribution network, a refined grid model is established to accurately reflect the spatio-temporal distribution characteristics of power outage events.
[0101] Spatial dimension partitioning: Use the K-means clustering algorithm to partition the distribution network equipment and clarify the spatial boundaries of feeders, substations, and user groups.
[0102] The K-means clustering algorithm is used to divide the data into K clusters, making the data points within the same cluster as similar as possible, while the data points between different clusters are quite different; the calculation formula is:
[0103]
[0104] In the formula, J is the objective function, representing the sum of squared errors within the cluster; K is the number of clusters; C i is all the data points in the i-th cluster; x is a certain data point; μ i is the center of the i-th cluster.
[0105] Specifically, taking the spatial position data of distribution network equipment X = {(2, 10), (2, 5), (8, 4), (5, 8), (7, 5), (6, 4)} and dividing these equipment into K = 2 clusters as an example:
[0106] First, randomly select two initial cluster centers μ1 = (2, 10) and μ2 = (5, 8). Then calculate the distance from each data point to the cluster center and assign it to the nearest cluster:
[0107] The calculation formula is:
[0108]
[0109] In the formula, C i represents the i-th cluster, that is, the set of all data points assigned to this cluster; x represents a certain data point, such as the spatial position or feature vector of a certain device in the distribution network; μ irepresents the center of the i-th cluster, which is the geometric center of all data points in the cluster; ||x - μ i || represents the Euclidean distance from the data point x to the cluster center μ i ; represents for all other clusters.
[0110] Therefore, for the data point (2, 10), its distance to μ1 is 0, and its distance to μ2 is and thus it belongs to C1; for the data point (2, 5), its distance to μ1 is 5, and its distance to μ2 is and thus it belongs to C1. Then, recalculate the cluster centers, μ1 = (2, 7, 5), μ2 = (6.5, 5.25). Finally, repeat the above steps until the cluster centers no longer change or reach the maximum number of iterations; the final result is: C1 = {(2, 10), (2, 5)}, C2 = {(8, 4), (5, 8), (7, 5), (6, 4)}.
[0111] Time dimension division: Dynamically adjust the time granularity according to the frequency and duration of power outage events, such as: minute level, hour level or day level, to improve the time resolution. For example, in the case of high-frequency faults such as bad weather, use the minute-level spatio-temporal granularity for analysis to improve the accuracy of spatio-temporal accounting.
[0112] Feature extraction: Extract key features from each spatio-temporal unit to describe its power outage impact. The key features include: equipment failure rate, user density, and historical power outage frequency.
[0113] Furthermore, the power outage impact of each spatio-temporal unit is described by the following formula:
[0114]
[0115] In the formula, the impact range (i, t) represents the power outage impact range of the i-th spatial unit within time t; represents the cumulative sum of all relevant features, where n is the total number of features or regions; ω j represents the weight coefficient of the j-th feature, used to measure the importance of this feature to the power outage impact; d ij represents the spatial distance between the i-th spatial unit and the j-th feature; the failure rate (j) represents the failure rate of the j-th device or region, usually a probability value between 0 and 1; the user (j) represents the number of users served by the j-th device or region.
[0116] Suppose a distribution network is divided into multiple spatio-temporal units, and one spatial unit i contains three devices within time t, that is, j = 1, 2, 3, and its parameters are as follows:
[0117] Weights: ω1 = 0.5, ω2 = 0.3, ω3 = 0.2; Spatial distance: d i1 = 1, d i2 = 2, d i3 = 3; Failure rate (1) = 0.1, Failure rate (2) = 0.2, Failure rate (3) = 0.15; Number of users: User (1) = 100, User (2) = 200, User (3) = 150;
[0118] Then the power outage impact range of this spatio-temporal unit is calculated as: Impact range (i, t) = 12.5. The final result shows that the power outage impact range of this spatio-temporal unit within time t is 12.5.
[0119] Step 3 Intelligent accounting algorithm:
[0120] This step uses an intelligent algorithm to accurately correct the time and scope of power outage events, eliminate errors, and accurately judge the impact on users:
[0121] Time correction: The time correction of power outage events is carried out through a time synchronization verification algorithm. This algorithm uses data interpolation methods to synchronize time data from different sources, eliminate timing errors, and thus improve the accuracy of power outage event accounting.
[0122] The time correction function of power outage events is:
[0123] T 修正 = T 原始 + ΔT 偏差
[0124] Where, T 修正 is the power outage event time after correction; T 原始 is the power outage event time obtained from the original data; ΔT 偏差 is the time deviation, indicating the time error caused by different data sources or acquisition delays.
[0125] Suppose a power outage event involves the following two data sources: The SCADA system records the power outage time as 10:02, and the AMI system records the power outage time as 10:05. Due to different data sources; therefore, there is a time deviation ΔT 偏差 . Align the timestamps of the two groups of data through interpolation method:
[0126] Suppose the interpolation result shows that the average delay of SCADA and AMI is 3 minutes, then the corrected time is: For SCADA data, T 修正 = 10:02, no delay; For AMI data, T 修正 = 10:05 - 3 = 10:02.
[0127] After calibration, the time of the two sets of data is consistent at 10:02, eliminating the timing error.
[0128] The influence scope of the power outage event needs to be judged in combination with the user's electricity consumption behavior; obvious mutation points appear in the user's electricity consumption curve when the power outage occurs or is restored. Through the mutation point detection algorithm, these abnormal changes are accurately identified, so as to determine the influence scope of the power outage event.
[0129] The mutation point detection algorithm includes the CUSUM method and cluster analysis;
[0130] The CUSUM method detects mutation points by calculating the cumulative change of the electricity consumption curve, and the calculation formula is:
[0131] S t =max(0,S t-1 +(x t -μ)-k)
[0132] In the formula, S t is the cumulative sum; x t is the electricity consumption at the current moment; t is the index of the data point, only representing the data order and no need to be adjusted manually; μ is the average value of the electricity consumption; k is the parameter controlling the sensitivity; when S t exceeds the preset threshold, it is determined as a mutation point.
[0133] Suppose the electricity consumption data of a certain user is: x = [100, 102, 98, 101, 5, 4, 3, 100, 99, 102], the normal electricity consumption average value μ = 100kW, the sensitivity parameter k = 5, and the threshold H = 20.
[0134] Initialize S0 = 0;
[0135] Calculate the cumulative sum S t :
[0136] S1 = max(0, 0+(100 - 100)-5) = 0
[0137] S2 = max(0, 0+(102 - 100)-5) = 0
[0138] S3 = max(0, 0+(98 - 100)-5) = 0
[0139] S4 = max(0, 0+(101 - 100)-5) = 0
[0140] S5 = max(0, 0+(5 - 100)-5) = max(0, -100) = 0
[0141] S6 = max(0, 0+(4 - 100)-5) = max(0, -101) = 0
[0142] S7 = max(0, 0 + (3 - 100) - 5) = max(0, -102) = 0
[0143] S8 = max(0, 0 + (100 - 100) - 5) = 0
[0144] S9 = max(0, 0 + (99 - 100) - 5) = 0
[0145] S 10 = max(0, 0 + (102 - 100) - 5) = 0
[0146] At the 5th moment, the power consumption drops sharply to 5, and the cumulative sum S t begins to decrease significantly, exceeding the threshold H = 20, and is determined as a mutation point. It can be seen that the power outage start time is the 5th moment, and the power outage recovery time is the 8th moment.
[0147] Cluster analysis groups the characteristics of users' power consumption curves to identify abnormal patterns that are significantly different from normal power consumption patterns. Users in the abnormal category are those affected by the power outage event.
[0148] Suppose the power consumption curve characteristics of 5 users in a certain distribution network: mean, variance, are as follows:
[0149] User 1: (100, 10), User 2: (102, 12), User 3: (5, 2), User 4: (4, 1), User 5: (101, 11);
[0150] Using the K-means clustering algorithm, the users are divided into two categories: normal cluster: User 1, User 2, User 5, whose means are close to 100 and variances are small; abnormal cluster: User 3, User 4, whose means are significantly reduced and variances are small. Thus, Users 3 and 4 in the abnormal cluster are identified as users affected by the power outage.
[0151] Furthermore, for a single user, the mutation point detection can determine the start and end times of its power outage; for multiple users, by comprehensively analyzing the mutation points of all users, the overall impact range of the power outage event can be inferred.
[0152] Single-user analysis: For User 3, the mutation point of the power consumption curve appears at the 5th moment, and the power consumption drops sharply from 100 to 5; the power outage start time is at the 5th moment; the power outage recovery time is at the 8th moment.
[0153] Multi-user analysis: By comprehensively analyzing the mutation points of all users, the mutation points of Users 3 and 4 are concentrated at the 5th to 8th moments; it is inferred that the overall impact range of the power outage event is: the time range is from the 5th to 8th moments; the spatial range is the substation area or feeder where Users 3 and 4 are located.
[0154] Step 4 Multidimensional Analysis Report:
[0155] This step uses machine learning and spatial interpolation techniques to predict the time - household number of potential power outage events, fill in missing data, and improve the integrity and accuracy of accounting:
[0156] Machine learning prediction: When using the random forest model in machine learning to predict the time - household number of power outage events, its prediction function is:
[0157]
[0158] In the formula: y represents the predicted time - household number of power outage events; f(x) represents the prediction function of the random forest model; N represents the number of decision trees in the random forest; α i represents the weight coefficient of the i - th decision tree; T i (X) represents the prediction result of the i - th decision tree for the input feature x.
[0159] Suppose the time - household number of power outage events in a certain area is predicted through the random forest model; the input feature x includes: equipment failure rate, user density, historical power outage frequency, and meteorological conditions, and the goal is to predict the time - household number y of power outage events.
[0160] That is, the input feature is: x = [equipment failure rate = 5%, user density = 200, historical power outage frequency = 2, rainfall = 50]. If the number of decision trees N = 3 in the random forest, then:
[0161] The random forest model uses historical data to train the random forest model to generate N = 3 decision trees; each tree makes an independent prediction based on the input feature x. The prediction results of the first tree are: T1(X)=50 time - household numbers, T2(X)=60 time - household numbers, T3(X)=55 time - household numbers. The weight coefficient α i = 1 / 3, and calculate the final predicted value y = 1 / 3(50 + 60 + 55)=55 time - household numbers, that is, the predicted time - household number of power outage events is 55.
[0162] Data filling: For filling in the data - missing areas, the Kriging interpolation method is adopted, and its interpolation formula is:
[0163]
[0164] In the formula: Z(x0) represents the interpolation result of the target point x0; μ(x0) represents the global trend or mean of the target point x0; n represents the number of known data points; λ i represents the weight coefficient of the i - th known point; Z(x i ) represents the actual observed value of the i - th known point; μ(x i ) represents the global trend or mean of the i - th known point; (Z(x i)-μ(x i )) represents the deviation between the actual observation value of the i-th known point and its trend value.
[0165] Assume that there are some areas of a distribution network with missing data, and the number of households in these areas during power outages needs to be filled in by Kriging interpolation. The known data points are: data point number 1, its spatial position is (0,0), and the measured observation value is 50; data point number 2, its spatial position is (1,0), and the measured observation value is 60; data point number 3, its spatial position is (0,1), and the measured observation value is 55; the goal is to predict the number of households Z(x0) during power outages at the target point x0 = (0.5,0.5).
[0166] Calculation process: The global mean of the known points is μ = 55, and the global trend is assumed to be μ(x0) = = 55; the weights are assumed to be λ1 = 0.4, λ2 = 0.3, and λ3 = 0.3;
[0167] Calculated based on the interpolation formula:
[0168] Calculate each item separately: λ1(Z(x1)-55)=-2, λ2(Z(x2)-55)=1.5, λ3(Z(x3)-55)=0; then Z(x0)=54.5, that is, the number of households when the interpolation result of the target point x0=(0.5,0.5) is 54.5.
[0169] Step 5: Household accounting report and decision support during power outage:
[0170] The calculation results of the number of households during power outages are used to generate multi-dimensional reports by power supply area, equipment type, time period, etc. The reports include heat maps of the number of households per hour, and statistical analysis results by power supply area and equipment type. The reports are visualized through the Web front end, using heat maps and dynamic statistical charts to show the impact areas of power outages. For high-risk areas, the system will automatically push them to the operation and maintenance management system to provide real-time fault repair and equipment maintenance decision support.
[0171] The above are only preferred specific implementation modes of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field can make equivalent replacements or changes according to the technical solutions and inventive concepts of the present invention within the technical scope disclosed by the present invention, which should be covered by the protection scope of the present invention.
Claims
1. A method for calculating the number of households during a power outage in a distribution network using multi-temporal and spatial scale data fusion, characterized in that: The multi-time and space scale data fusion distribution network power outage household number accounting method comprises the following steps: Step 1 Multi-source data fusion framework: By collecting, cleaning and integrating multi-source heterogeneous data, a high-quality spatiotemporal data set is constructed to provide a basis for subsequent fusion analysis; Step 2: Spatiotemporal grid modeling: Based on the spatiotemporal data set, K-means clustering and dynamic time granularity division are used to perform refined spatiotemporal grid modeling of the distribution network and extract key features; Step 3: Intelligent accounting algorithm: Through time synchronization correction and mutation point detection, the time range and user impact of power outage events are accurately corrected; Step 4 Multidimensional analysis report: Combine machine learning to predict the number of households at the time of potential power outages, and use spatial interpolation to fill in missing data to improve accounting completeness and accuracy; Step 5: Household accounting report and decision support during power outages: Generate multi-dimensional visualization reports and push high-risk area data to provide decision support for power grid operation and planning.
2. The method for calculating the number of households in a power outage in a distribution network using multi-temporal and spatial scale data fusion according to claim 1 is characterized by: In step 1, the collected data include real-time and historical data from the SCADA system, AMI system, equipment inventory, historical power outage records and user repair platform.
3. The method for calculating the number of households during a power outage in a distribution network using multi-temporal and spatial scale data fusion according to claim 1 is characterized by: In step 1, the process of cleaning and integrating multi-source heterogeneous data includes: Data preprocessing: denoising the collected data through the Z-score algorithm, time alignment based on timestamp synchronization, and normalization to ensure data consistency and availability.
4. The method for calculating the number of households during a power outage in a distribution network using multi-temporal and spatial scale data fusion according to claim 1 is characterized in that: In step 2, first, the K-means clustering algorithm is used to perform cluster analysis on the distribution network equipment to determine the regional boundaries such as feeders, substations, and user groups; then, the time granularity is dynamically adjusted to minutes, hours, or days according to the frequency and duration of power outages to improve the time resolution; finally, key features are extracted from each spatiotemporal unit, including equipment failure rate, user density, and historical power outage frequency, to describe the impact of power outages.
5. The method for calculating the number of households during a power outage in a distribution network using multi-temporal and spatial scale data fusion according to claim 4 is characterized by: The impact of power outages per space-time unit is described by the following formula: In the formula, the impact range (i, t) represents the power outage impact range of the i-th spatial unit within time t; represents the cumulative sum of all relevant features, where n is the total number of features or regions; ω j represents the weight coefficient of the jth feature, which is used to measure the importance of the feature on the power outage; d ij represents the spatial distance between the i-th spatial unit and the j-th feature; failure rate (j) represents the failure rate of the j-th device or area, which is usually a probability value between 0 and 1; users (j) represents the number of users served by the j-th device or area.
6. The method for calculating the number of households during a power outage in a distribution network using multi-temporal and spatial scale data fusion according to claim 1 is characterized by: In step 3, the time correction of the power outage event is performed by a time synchronization verification algorithm, which uses a data interpolation method to synchronize time data from different sources to eliminate timing errors; The time correction function of the power outage event is: T 修正 =T 原始 +ΔT 偏差 Where, T 修正 is the corrected power outage event time; T 原始 is the power outage event time obtained from the original data; ΔT 偏差 It is the time deviation, which indicates the time error caused by different data sources or acquisition delay.
7. The method for calculating the number of households in a power outage in a distribution network using multi-temporal and spatial scale data fusion according to claim 1 is characterized by: In step 3, the impact range of the power outage event needs to be determined in combination with the user's power consumption behavior; the user's power consumption curve has obvious mutation points when the power outage occurs or is restored. Through the mutation point detection algorithm, these abnormal changes can be accurately identified to determine the impact range of the power outage event; For a single user, mutation point detection can determine the start and end time of the power outage; for multiple users, the overall impact range of the power outage can be inferred by comprehensively analyzing the mutation points of all users.
8. The method for calculating the number of households in a power outage in a distribution network using multi-temporal and spatial scale data fusion according to claim 7 is characterized by: The mutation point detection algorithm includes CUSUM method and cluster analysis; The CUSUM method detects mutation points by calculating the cumulative amount of changes in the power consumption curve. The calculation formula is: S t =max(0,S t-1 +(x t -μ)-k) In the formula, S t is the cumulative sum; x t is the power consumption at the current moment; t is the index of the data point; μ is the mean power consumption; k is the parameter for controlling sensitivity; When S t When it exceeds the preset threshold, it is determined as a mutation point; Cluster analysis groups the characteristics of user power consumption curves and identifies abnormal patterns that are significantly different from normal power consumption patterns. Users in the abnormal category are those affected by the power outage.
9. The method for calculating the number of households during a power outage in a distribution network using multi-temporal and spatial scale data fusion according to claim 1 is characterized by: In step 4, machine learning uses the random forest model to predict the number of households during a power outage, and its prediction function is: Where: y represents the number of households at the time of the predicted power outage; f(x) represents the prediction function of the random forest model; N represents the number of decision trees in the random forest; α i Represents the weight coefficient of the i-th decision tree; T i (X) represents the prediction result of the i-th decision tree for the input feature x; To fill in the missing data areas, Kriging interpolation method is used, and its interpolation formula is: Where: Z(x0) represents the interpolation result of the target point x0; μ(x0) represents the global trend or mean of the target point x0; n represents the number of known data points; λ i represents the weight coefficient of the i-th known point; Z(x i ) represents the actual observed value of the i-th known point; μ(x i ) represents the global trend or mean of the i-th known point; (Z(x i )-μ(x i )) represents the deviation between the actual observation value of the i-th known point and its trend value.