Power load scene generation and utility analysis method
By generating high-fidelity power load scenario data through functional principal component analysis and differential privacy protection, the problems of low accuracy in fluctuation pattern recognition and insufficient privacy protection in existing technologies are solved, and high-precision power load scenario data generation and privacy protection are achieved.
Patent Information
- Application Number
- CN202511700268.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-19
- Publication Date
- 2026-02-17
AI Technical Summary
Existing methods for generating power load data ignore the temporal correlation and overall functionality of load data, resulting in low accuracy in identifying fluctuation patterns and a lack of privacy protection, posing a risk of user information leakage.
We employ functional principal component analysis for dimensionality reduction and feature extraction, combined with differential privacy protection and Ziggurat sampling algorithm to generate high-fidelity power load scenario data, and validate the results using functional K-means clustering.
High-fidelity power load scenario data is generated while protecting privacy, improving the accuracy of fluctuation pattern recognition, reducing the risk of user information leakage, and ensuring the statistical characteristics and diversity of the data.
Smart Images

Figure CN121546571A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a method for generating and analyzing power load scenarios. Background Technology
[0002] With the continuous increase in new power systems, the simulation of cascading power system failures and the development of fault mitigation strategies have become a hot topic in order to address the enormous losses caused by the frequent large-scale power outages worldwide in recent years. However, traditional scenario generation methods are insufficient to meet the needs of modern mitigation strategies. Therefore, how to safely and effectively generate high-fidelity power load scenario data to provide technical support for the development of fault mitigation strategies has become an important issue in power load scenario generation.
[0003] In recent years, the construction of intelligent power systems has been continuously advancing, and electricity demand has been increasing. Smart meters, widely deployed in smart grids, can now collect power load data at minute-level or even higher frequencies. The power load data obtained through smart meters is not only massive and highly dense, but also exhibits rich dynamic evolution characteristics across continuous time dimensions. This results in the generation of a set of representative, high-quality power load time series scenario data, allowing system operators' dispatchers to consider the uncertainty of power load when making decisions such as stochastic planning, unit start-up and shutdown, and power trading. This is of vital importance for grid operation planning and safety assurance.
[0004] Most existing methods for generating electricity load data are based on load trajectory data from a single user within the power grid. These methods neglect the inherent temporal correlations and overall functionalities of load changes in the original electricity load data, resulting in low accuracy in identifying fluctuation patterns and distorted characterization of randomness. This, in turn, affects the accuracy and robustness of power system dispatching and planning decisions. Furthermore, large-scale, high-precision electricity load data profoundly reflects highly sensitive information such as users' personal habits, family structures, and economic conditions. Without robust protection measures during data collection, sharing, or mining, user information can easily be leaked and maliciously used for user profiling, targeted fraud, or even physical security threats, posing a serious risk to user safety.
[0005] In recent years, the rise of functional data analysis methods and differential privacy methods has not only effectively captured the dynamic evolution characteristics of power load in the continuous time dimension, but also significantly improved the accuracy of identifying load fluctuation patterns and inherent laws, providing large-scale, high-quality, multimodal synthetic scenario data support for the development and training of large-scale power system load forecasting, scheduling optimization, and risk assessment models. Summary of the Invention
[0006] The purpose of this invention is to provide a method for generating and analyzing power load scenarios, which can effectively generate high-fidelity scenario data while ensuring privacy.
[0007] The technical solution of this invention is: a method for generating and analyzing power load scenarios, comprising: Step 1: Collect raw power load data and preprocess outliers; Step 2: Perform smoothing and fitting processing on the individual user load data; Step 3: Perform functional principal component analysis on the data to achieve dimensionality reduction and feature extraction while ensuring the dynamic characteristics of the data, thereby providing a low-dimensional and operable representation for differential privacy perturbation and subsequent data generation; Step 4: The principal component scores obtained in Step 3 are subjected to a Laplace noise injection mechanism that satisfies differential privacy protection to ensure that the output results meet the quantifiable privacy protection strength. Step 5: Perform probabilistic modeling on the principal component scores after differential privacy protection, and use the Ziggurat sampling algorithm or Gibbs sampling method to generate new principal component score samples to ensure the rationality and diversity of the generated data in terms of statistical characteristics. Step 6: Linearly combine the newly generated principal component scores with the functional principal component basis functions, and reconstruct the functions based on the Karhunen-Loève expansion to obtain new power load data with privacy protection; Step 7: Validate the privacy-preserving data generation results using the simultaneous confidence band theory of functional data and the functional K-means clustering method based on the TSTR evaluation model.
[0008] In the above technical solution, step 1 involves taking measures based on an hourly time interval. The system collects annual electricity load data for each user and performs pre-processing operations on the collected raw electricity load data, including outlier removal.
[0009] In the above technical solution, step 2 involves selecting basis functions and smoothing parameters, performing B-spline basis function fitting on the power load data, and transforming discrete points into smooth function curves for functional data analysis.
[0010] In the above technical solution, the inspection method in step 7 includes: Step 7.1: Check whether the mean function of the newly generated data falls within the simultaneous confidence band at the 95% confidence level constructed based on the original data; Step 7.2: Perform K-means clustering analysis based on functional principal component scores on the original power load data and the generated new dataset, and then evaluate the clustering results.
[0011] The advantages of this invention are: This invention is the first to introduce a differential privacy protection mechanism based on functional principal component analysis. It incorporates the simultaneous confidence band theory of functional data and the functional K-means clustering method into the verification system of privacy-preserving data generation results, and comprehensively verifies the effectiveness of the scene generation algorithm proposed in this invention in generating high-fidelity scene data while ensuring privacy. Attached Figure Description
[0012] The present invention will be further described below with reference to the accompanying drawings and embodiments: Figure 1 This is a flowchart of the method of the present invention.
[0013] Figure 2 This is a schematic diagram of the B-spline smooth trajectory and its mean function for 193 electrical loads according to the present invention.
[0014] Figure 3 This is a schematic diagram of the five functional principal component basis functions of the user power load of the present invention.
[0015] Figure 4 This is a schematic diagram of the probability density of the top 5 principal components of the user's power load according to the present invention.
[0016] Figure 5 This is a graph showing how MAE, MRE, and SSE vary with the privacy budget in this invention.
[0017] Figure 6 This is a schematic diagram comparing the original principal component scores and the noisy principal component scores of the data in this invention.
[0018] Figure 7 This is a schematic diagram of the mean function with a 95% confidence level and simultaneous confidence band test for the present invention.
[0019] Figure 8 This is a schematic diagram of the K-means clustering analysis results of the original data in this invention.
[0020] Figure 9 This is a schematic diagram of the K-means clustering analysis results for the new dataset of electricity load users in this invention. Detailed Implementation
[0021] Example: See Figures 1 to 9 As shown, this invention relates to a method for generating and analyzing power load scenarios, comprising: Step 1: Collect raw power load data and preprocess outliers; Specifically, this includes: taking measures based on one-hour time intervals. The annual electricity load data for each user is shown in the table below. Table 1. Power Load Data Storage Format in Indicates users, totaling One user. Indicates the time of electricity consumption, total That moment. Indicates the first The electricity consumption of a user at time j is measured in kilowatt-hours, and the collected raw power load data undergoes pre-processing operations such as outlier removal.
[0022] Step 2: Perform smoothing and fitting processing on the individual user load data; The specific steps involve selecting basis functions and smoothing parameters, fitting B-spline basis functions to the power load data, and transforming discrete points into smooth function curves for functional data analysis. Specifically, this includes: [The text abruptly ends here, likely due to an incomplete sentence or missing information.] The data observed by each user is represented as ,in For the first The number of times the sample is observed, for the first sample one sample By fitting a function to each discrete observation, we can obtain... .Will Treat as a bounded interval Smooth random function on of 1 independent sample, of which It is also a kind of stochastic process. When fitting a function to observed discrete data, it is assumed that the data... satisfy in, for Implementation of stochastic processes on the surface, Time is the independent variable of Each value can be selected, and its range is: , It is a function of standard deviation. It is a random error with a mean of 0 and a variance of 1.
[0023] A basis function smoothing method is chosen for fitting, which involves fitting the random function by linearly combining known and independent basis functions. Let... Given a set of mutually independent basis functions, a random function can be obtained by linear combination. for in, for exist The coefficient on. At this point, the random function... The expansion was performed on an infinite number of basis functions. After imposing a finite constraint on the number of basis functions, the random function... Transform it into the sum of finite basis functions. The estimated value is Step 3: Perform functional principal component analysis on the data to achieve dimensionality reduction and feature extraction while ensuring the dynamic characteristics of the data, thereby providing a low-dimensional and operable representation for differential privacy perturbation and subsequent data generation; Specifically, this includes prior designation. It is Hilbert space stochastic processes on, where For time intervals, Let be a set of observations. Let B-spline basis function vectors be used. The coefficient vector on the B-spline basis functions is Then the covariance function The estimate is , , yes The Middle 1 element. Mean function The estimated value .
[0024] The replace In ,get ,in, Matrix correspondence The eigenvalues of are and their eigenfunctions are D is the matrix obtained by performing Cholesky decomposition on matrix H, and matrix H is defined as follows: Then, based on the probability model using the Karhunen-Loéve expansion theorem: as well as Using the estimated characteristic function and eigenvalue to estimate The first sample The scores of each principal component are: Ultimately based on Before the election Principal component values.
[0025] Step 4: The principal component scores obtained in Step 3 are subjected to Laplace noise injection mechanism that satisfies differential privacy protection to ensure that the output results meet quantifiable privacy protection strength. users Each element of the dimensional functional principal component score variable was added The noise, of which This represents the maximum range of power load data.
[0026] Step 5: Perform probabilistic modeling on the principal component scores after differential privacy protection, and generate new principal component score samples using the Ziggurat sampling algorithm or the Gibbs sampling method to ensure the rationality and diversity of the generated data in terms of statistical characteristics. Furthermore, this invention classifies the two algorithms into DP-FPCA-MN-SC and DP-FPCA-Gibbs-SC algorithms based on different sampling methods.
[0027] Step 6: Linearly combine the newly generated principal component scores with the functional principal component basis functions, and reconstruct the functions based on the Karhunen-Loève expansion to obtain new power load data with privacy protection.
[0028] Step 7: Validate the privacy-preserving data generation results using the simultaneous confidence band theory of functional data and the functional K-means clustering method based on the TSTR evaluation model.
[0029] Specific testing methods include: Step 7.1: Check whether the mean function of the newly generated data falls within the simultaneous confidence band at the 95% confidence level constructed based on the original data.
[0030] Step 7.2: Perform K-means clustering analysis based on functional principal component scores on the original power load data and the generated new dataset, and then further evaluate the clustering effect using indicators such as SC, DBICHI, and DI.
[0031] In one embodiment of the present invention, data sourced from three feeders in the Midwestern United States in 2017 is used. The electricity consumption data for each user is collected, with a sampling interval of 1 hour for all users. Each user has... There are 100 observation data points. Because some users are vacant, resulting in consistently zero electricity load data, their electricity consumption data was removed, leaving only the following data points. The effective electricity consumption data for each user and the processing status of power load data for the three feeder users are shown in the table below: Table 2 Actual Distribution of Raw Power Load Data After collecting power load data from users along three feeders, a basis function and smoothing parameters are selected, and B-spline basis function fitting is performed on the data. In this example, the number of nodes is set as follows: Based on the principle of minimizing mean square error, this paper employs 80 fourth-order B-spline basis functions. After setting the above parameters, the B-spline smooth trajectories and their mean functions of the 193 power loads selected in this example are as follows: Figure 2 As shown.
[0032] Based on functional principal component analysis (PCA) theory, feature extraction was performed on the fitted curves of user electricity load data. The results of PCA on electricity load data are summarized in the table below: Table 3 Summary of power load data processing results based on data function-based principal component analysis. Therefore, the user power load part of the functional principal component basis functions are as follows: Figure 3 As shown, its probability density function is as follows: Figure 4 As shown, the first five functional principal component basis functions are selected to calculate the first five functional principal component score variables. Then, based on different privacy protection budgets, the MAE, MRE, and SSE of each functional principal component variable are compared to measure the usability of principal component data under different privacy protection budgets. The trends of each variable with the privacy budget are shown below. Figure 5 As shown. Based on the combined results of the three sets of experiments, the privacy budget for the functional principal component variables was set to 0.6. At this point, the three error metrics—MAE, MRE, and SSE—simultaneously entered a stable low-value range, and further increases in the privacy budget yielded limited improvements in data usability. For example... Figure 6 This paper presents a comparison of the distributions of the functional principal component score variables in the original power load data and the functional principal component score variables after differential privacy noise addition.
[0033] After performing differential privacy processing on the selected functional principal component score variables, random sampling based on multivariate normal distribution and Gibbs sampling using Cholesky decomposition and Zigguart inverse transform methods was performed to prepare for generating power load data. Simultaneously, the effectiveness was verified using simultaneous confidence band theory. The results are as follows: Figure 7 As shown.
[0034] Besides requiring theoretical verification for confidence, this invention aims to reveal the group differences and patterns in user electricity load data. It designs a method based on K-means clustering to reveal the inherent characteristics of the original electricity load data. In this example, the K-means clustering algorithm is used to perform functional clustering analysis on the scores of the first five functional principal components. Based on the elbow rule, the optimal cluster size K=3 is determined. The K-means clustering results are as follows: Figure 3 As shown, the analysis results are as follows: Figure 8 As shown in the table below, the number of users in each category is as follows: Table 4 Summary of K-means clustering results for raw data To evaluate the effectiveness of the DP-FPCA-MN-SC and DP-FPCA-Gibbs-SC algorithms in generating new power user load datasets, simulation-real training tests were conducted on the original data and the two new datasets based on functional K-means clustering analysis to identify the overall structure of the new datasets and facilitate comparison with the original data.
[0035] Table 5 Summary of K-means clustering results for each power load dataset Furthermore, to measure the clustering effect of the original data and the newly generated data, this embodiment calculated clustering effect evaluation indicators such as SC, DBI, CHI, and DI, and the results are shown in the table below. As can be seen from the table, the original data showed the best clustering indicators. The analysis results show that an SC of 0.91 indicates that the intra-cluster compactness and inter-cluster separation are ideal; a DBI of 0.61 reflects a good balance between intra-cluster dispersion and inter-cluster distance; a CHI of 434.98 confirms that the inter-cluster variance has an overwhelming advantage over the intra-cluster variance; and a DI of 0.46 indicates that the ratio of the minimum inter-cluster distance to the maximum cluster diameter is at a healthy level. The DP-FPCA-MN-SC generated data still maintains a high level of clustering quality. The SC value is 0.84, a 7% decrease from the original data, proving that multivariate normal sampling effectively maintained the sample distribution structure. The DBI slightly increased to 0.64, suggesting a controllable weakening of inter-cluster separation. The CHI decreased to 414.41, reflecting a slight decrease in the variance ratio advantage. The DI value is 0.4, indicating that the maximum cluster diameter has expanded somewhat. The clustering scores of the DP-FPCA-Gibbs-SC generated data are 0.85, DBI is 0.64, CHI is 421.66 and DI is 0.42, which are all better than those of the DP-FPCA-MN-SC generated data.
[0036] Table 6 Comparison of Clustering Performance Indicators for Various Power Load Data Sets Finally, to verify the effectiveness of the data generated by the DP-FPCA-SC algorithm, simulation training tests were conducted on the original data. Based on the number of functional K-means clusters and cluster center parameters obtained from the new power load datasets generated by the DP-FPCA-MN-SC and DP-FPCA-Gibbs algorithms, the original data itself was re-clustered, and cluster evaluation indicators such as SC, DBI, CHI, and DI were calculated. As shown in the table below, when the original data is re-divided using the clustering parameters determined by the data generated by DP-FPCA-MN-SC, the resulting SC is 0.80, DBI is 0.72, CHI is 399.78, and DI is 0.39. Although the values of each indicator have decreased slightly compared to the direct clustering results of the original data, such as SC decreasing by 0.11, DBI increasing by 0.11, CHI decreasing by 35.20, and DI decreasing by 0.07, the overall indicator values are still at a good level. This indicates that the data generated by the DP-FPCA-MN-SC algorithm can capture the main clustering features of the original data, and its clustering parameters still have good applicability and representativeness for the original data.
[0037] Meanwhile, when the original data was re-partitioned using the clustering parameters determined by the data generated by DP-FPCA-Gibbs-SC, the resulting SC was 0.82, DBI was 0.70, CHI was 402.32, and DI was 0.40. Compared with the results of DP-FPCA-MN-SC, SC was slightly higher by 0.02, DBI was slightly lower by 0.02, CHI was slightly higher by 2.54, and DI was slightly higher by 0.01. This shows that the clustering parameters derived from the data generated by the DP-FPCA-Gibbs-SC algorithm have a slightly better fit with the direct clustering results of the original data when re-partitioning the original data than DP-FPCA-MN-SC, highlighting the potential impact of different generation algorithms on the stability of clustering parameters.
[0038] Table 7 Comparison of Clustering Performance Indicators between Simulation and Real Clustering Data Generated by DP-FPCA-SC Algorithm Of course, the above embodiments are only for illustrating the technical concept and features of the present invention, and their purpose is to enable those skilled in the art to understand the content of the present invention and implement it accordingly. They should not be used to limit the scope of protection of the present invention. All modifications made according to the spirit and essence of the main technical solution of the present invention should be covered within the scope of protection of the present invention.
Claims
1. A method for power load scenario generation and utility analysis, characterized in that, The application comprises the following steps: Step 1: collecting raw data of power load data and pre-processing abnormal values; Step 2: smoothing and fitting individual user load data; Step 3: performing functional principal component analysis on the data to realize dimension reduction and feature extraction under the premise of ensuring the dynamic characteristics of the data, and then providing a low-dimensional and operable representation for subsequent differential privacy disturbance and data generation; Step 4: using the Laplace noise injection mechanism to protect the privacy of the principal component scores obtained in step 3 to ensure that the output results meet the quantifiable privacy protection strength; Step 5: modeling the principal component scores after differential privacy protection, and generating new principal component score samples by using the Ziggurat sampling algorithm or the Gibbs sampling method to ensure the rationality and diversity of the generated data in statistical characteristics; Step 6: linearly combining the newly generated principal component scores with the functional principal component basis functions, and reconstructing the functions based on the Karhunen-Loève expansion to obtain new power load data with privacy protection; Step 7: using the simultaneous confidence band theory of functional data and the functional K-means clustering method based on the TSTR evaluation mode to test the privacy protection data generation results.
2. The method of claim 1, wherein, The step 1 takes one-hour time intervals, and performs data processing pre-operations including outlier rejection on the collected raw power load data.
3. The method of claim 1, wherein, In step 2, the base functions and smoothing parameters are selected, the power load data is fitted by B-spline base functions, the discrete points are converted into smooth function curves, and the functional data analysis is performed.
4. The method of claim 1, wherein, The test method in step 7 comprises: Step 7.1: testing whether the mean function of the newly generated data falls within the 95% confidence level of the simultaneous confidence band constructed based on the original data; Step 7.2: performing K-means clustering analysis based on the functional principal component scores on the original power load data and the generated new data set, respectively, and then evaluating the clustering results.
Citation Information
Patent Citations
Internet of Things time series data prediction method, system and device and storage medium
CN114492561A
Differential privacy linear regression method and system based on principal component analysis and function mechanism
CN114969829A
Function type principal component test method and regression prediction method in function type data
CN115470454A
Multi-target prediction method based on state transition in integrated energy system
CN117291445A