Load sequence scene construction method considering environment and industry elements

By constructing load sequence scenarios for environmental and industry factors, using the ARIMAX model to generate load sequences of different time scales, the problem of existing load prediction methods ignoring external factors is solved, and the accuracy and reliability of prediction are improved.

CN120124005APending Publication Date: 2025-06-10STATE GRID NINGXIA ELECTRIC POWER CO LTD ECO TECH RES INST
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510171603.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-17
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

The existing load prediction methods ignore the impact of environmental and industry factors on load changes, resulting in low accuracy of prediction results.

Method used

The load sequence scenario construction method that takes into account environmental and industry factors is used to generate load sequences of different time scales through data cleaning, time series stationarity inspection, ARIMAX model construction and optimization.

Benefits of technology

It improves the accuracy and reliability of load prediction, can more accurately model historical load data and combine external factors to build load sequences, supporting power system scheduling and demand response.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120124005A_ABST
    Figure CN120124005A_ABST
Patent Text Reader

Abstract

The invention provides a load sequence scene construction method considering environment and industry elements, and relates to the field of power system prediction planning, and the method comprises the steps: carrying out the data cleaning and data conversion processing of historical load data and meteorological data obtained in advance, and obtaining a comprehensive data set; performing time sequence stationarity check on the comprehensive data set, and performing differential processing on non-stationary data in the stationarity check to obtain a stationary comprehensive data set; constructing an ARIMAX model based on the exogenous variables, and training and optimizing the ARIMAX model by using the stationary comprehensive data set; and fitting ARIMAX models of different time scales to obtain load sequences of different time scales. According to the method, factors such as climate change and industry requirements are comprehensively considered, the accuracy and reliability of load prediction are improved, the problem that an existing load prediction method is insufficient in consideration of external influence factors is solved, and the accuracy and reliability of load prediction are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of power system prediction and planning. Specifically, it particularly relates to a method for constructing a load sequence scenario considering environmental and industry factors. Background Art

[0002] With the continuous growth of global energy demand, load forecasting in power systems has become an important part of modern power dispatching, energy management, and optimal dispatching. The accuracy of load forecasting directly affects the stability of the power grid, the effective utilization of energy, and economy. Traditional load forecasting methods mostly rely on historical load data and basic time series analysis. Although they can provide certain reference for short-term load forecasting, they ignore the profound impact of environmental factors and industry factors on load changes. For example, environmental factors such as temperature, humidity, and seasonal changes will significantly affect the fluctuation of power load. According to the research of the National Climate Center of China, during the winter cold wave, the peak power load can be 20%-30% higher than the normal load. In addition, the special needs of industries, such as the operating status of manufacturing industries and industrial production cycles, will also affect power load.

[0003] According to the "Research Report on China's Power Load Forecasting", the characteristics of electricity consumption loads in different industries are significantly different, and the regularity and suddenness of power load changes vary under different time periods and weather conditions. However, existing load forecasting models often lack the consideration of the integration of environmental changes (such as climate change) and industry factors (such as production capacity changes, policy adjustments, etc.), resulting in relatively low accuracy of forecasting results.

[0004] Regarding the problems in the related art, no effective solutions have been proposed yet. Summary of the Invention

[0005] In view of this, the present invention provides a method for constructing a load sequence scenario considering environmental and industry factors to solve the above-mentioned problems.

[0006] To solve the above problems, the specific technical solutions adopted by the present invention are as follows:

[0007] A method for constructing a load sequence scenario considering environmental and industry factors includes:

[0008] S1. Perform data cleaning and data conversion processing on the pre-acquired historical load data and meteorological data to obtain a comprehensive data set;

[0009] S2. Use the PP unit root test method to check the stationarity of the time series of the comprehensive data set, and perform differencing processing on the non-stationary data in the stationarity check to obtain a stationary comprehensive data set;

[0010] S3. Construct an ARIMAX model based on exogenous variables, and train and optimize the ARIMAX model using the stationary comprehensive dataset to obtain the final ARIMAX model;

[0011] S4. Use the final ARIMAX model to fit ARIMAX models of different time scales to obtain load sequences of different time scales.

[0012] Preferably, the data cleaning and data transformation processing of the pre-acquired historical load data and meteorological data to obtain the comprehensive dataset includes:

[0013] S11. Process the missing values in the pre-acquired historical load data and meteorological data;

[0014] S12. Use the box plot method to process the outliers in the historical load data and meteorological data after missing value processing;

[0015] S13. Successively perform timestamp processing, data standardization processing, and unified processing of data time resolution on the historical load data and meteorological data after outlier processing;

[0016] S14. Merge the historical load data and meteorological data after unified processing of data time resolution to obtain the comprehensive dataset.

[0017] Preferably, the use of the PP unit root test method to check the stationarity of the time series of the comprehensive dataset and perform differencing processing on the non-stationary data in the stationarity check to obtain the stationary comprehensive dataset includes:

[0018] S21. For each time series data in the comprehensive dataset, construct a regression model and use the least squares method to estimate the parameter terms in the regression model. The parameter terms include a constant term, a trend term coefficient, and a lag autoregressive coefficient;

[0019] The expression of the regression model is:

[0020] y t = α + βt + γy t-1 + ε t ;

[0021] In the formula, y t represents the time series data at time t, α represents the constant term, β t represents the time trend term, γ represents the lag autoregressive coefficient, γy t-1 represents the autoregressive term, and ε t represents the error term;

[0022] S22. Calculate the test statistic of the unit root according to the regression model, and compare the critical value in the critical value table with the test statistic. If the test statistic is greater than the critical value, it indicates that the time series data in the comprehensive dataset is stationary data; otherwise, it indicates that the time series data in the comprehensive dataset is non-stationary data.

[0023] S23. Based on the difference formula, perform difference processing on the non-stationary data to convert the non-stationary data into stationary data, and obtain a comprehensive dataset with stationarity.

[0024] Preferably, the calculating the test statistic of the unit root according to the regression model includes:

[0025] Calculate the residuals of the time series data in the comprehensive dataset using the regression model;

[0026] Calculate the residual variance based on the residuals of the time series data, and calculate the standard error based on the residual variance;

[0027] Calculate the test statistic of the unit root using the standard error and the lag autoregressive coefficient.

[0028] Preferably, the difference formula is:

[0029] y' t =y t -y t-1 ;

[0030] In the formula, y' t represents the data after the first difference, y t represents the time series data at time t, and y t-1 represents the time series data at time t - 1.

[0031] Preferably, the constructing an ARIMAX model based on exogenous variables and training and optimizing the ARIMAX model using the comprehensive dataset with stationarity to obtain the final ARIMAX model includes:

[0032] S31. Based on the basic time series model, construct an ARIMAX model in combination with exogenous variables;

[0033] S32. Based on the statistical characteristics of the time series data in the comprehensive dataset, determine the order of the ARIMAX model and the lag period of exogenous variables by means of stepwise regression combined with the BIC information criterion;

[0034] S33. Estimate the parameters of the ARIMAX model, and optimize the model parameters using the simulated annealing algorithm;

[0035] S34. Train and evaluate the ARIMAX model with optimized order, exogenous variable lag periods, and parameters using the comprehensive dataset with stationarity, and optimize the ARIMAX model according to the evaluation results to obtain the final ARIMAX model.

[0036] Preferably, the expression of the ARIMAX model is:

[0037]

[0038] In the formula, y t represents the time series data at time t, μ represents the constant term, represents the coefficient of the autoregressive part, θ j represents the coefficient of the moving average part, X t-k represents the value of the exogenous variable matrix at lag period k, β k represents the coefficient of the exogenous variable lag period k, ε t represents the error term, p represents the lag order of the autoregressive part, q represents the order of the moving average part, r represents the lag order of the exogenous variable part, j represents the lag term of the moving average part, i represents the lag term of the autoregressive part, and k represents the lag term of the exogenous variable part.

[0039] Preferably, determining the order of the ARIMAX model and the exogenous variable lag periods by using stepwise regression combined with the BIC information criterion based on the statistical characteristics of the time series data in the comprehensive dataset includes:

[0040] S321. Based on the time series data in the comprehensive dataset, determine the order range of autoregression and moving average by establishing an autocorrelation plot and a partial autocorrelation plot;

[0041] S322. According to the order range of autoregression and moving average, initially set the orders of autoregression and moving average of the ARIMAX model;

[0042] S323. Use the stepwise regression method to gradually introduce the exogenous variable lag periods, and calculate the BIC value introduced in each step according to the BIC information criterion, and select the exogenous variable lag period that minimizes BIC;

[0043] S324. Perform backward elimination on the initially selected model, remove the exogenous variable lag periods one by one, and finally determine the optimal order of the ARIMAX model and the exogenous variable lag periods according to the BIC information criterion.

[0044] Preferably, the parameter estimation of the ARIMAX model includes: moving average parameter, autoregressive parameter estimation, and exogenous variable parameter estimation;

[0045] The exogenous variable parameter estimation uses the ridge regression algorithm. The goal of the ridge regression algorithm is to minimize the loss function with a regularization term, and the expression of the loss function with a regularization term is:

[0046]

[0047] In the formula, ι(β) represents the loss function with a regularization term, x t represents the exogenous variable observation vector at time point t, β represents the regression coefficient corresponding to the exogenous variable, and λ represents the regularization parameter; T represents the number of observation data, j represents the index of the regression coefficient in the regularization term, and m represents the number of features in the ARIMAX model.

[0048] Preferably, the ARIMAX models with different time scales include: the daily time scale ARIMAX model, the weekly time scale ARIMAX model, the monthly time scale ARIMAX model, and the annual time scale ARIMAX model;

[0049] The expression of the daily time scale ARIMAX model is:

[0050]

[0051] The expression of the weekly time scale ARIMAX model is:

[0052]

[0053] The expression of the monthly time scale ARIMAX model is:

[0054]

[0055] The expression of the annual time scale ARIMAX model is:

[0056]

[0057] In the formula, represents the load sequence generated under the daily time scale, represents the load sequence generated under the weekly time scale, represents the load sequence generated under the monthly time scale, represents the load sequence generated under the annual time scale, μ represents the constant term, represents the autoregressive part under the daily time scale, represents the autoregressive part under the weekly time scale, represents the autoregressive part under the monthly time scale, represents the autoregressive part under the annual time scale, Represents the moving average part at the daily time scale, Represents the moving average part at the weekly time scale, Represents the moving average part at the monthly time scale, Represents the moving average part at the annual time scale, Represents the exogenous variables at the daily time scale, Represents the exogenous variables at the weekly time scale, Represents the exogenous variables at the monthly time scale, Represents the exogenous variables at the annual time scale, ε t Represents the error term, Represents the coefficient of the autoregressive part, θ i Represents the coefficient of the moving average part, β k Represents the coefficient of the exogenous variable lag period k. p represents the lag order of the autoregressive part, q represents the order of the moving average part, r represents the lag order of the exogenous variable part, j represents the lag term of the moving average part, i represents the lag term of the autoregressive part, and k represents the lag term of the exogenous variable part.

[0058] The beneficial effects of the present invention are as follows:

[0059] 1. By comprehensively considering factors such as climate change and industry demand, the present invention improves the accuracy and reliability of load forecasting, solves the problem that existing load forecasting methods insufficiently consider external influencing factors, and improves the accuracy and reliability of load forecasting.

[0060] 2. Aiming at constructing load sequence scenarios, the present invention comprehensively considers the impacts of environmental factors and industry elements on load changes, introduces the environmental factors of meteorological changes and the seasonal characteristics of industry demand, establishes a load forecasting framework based on the ARIMAX model, accurately models historical load data, and constructs accurate load sequences in combination with external factors; integrates multi-dimensional constraint conditions, supports the generation and optimization of multiple load scenarios, and can provide accurate load forecasting and decision-making support for power system scheduling, demand response, etc.

[0061] 3. On the basis of the basic time series model, by integrating the impacts of external environmental factors and industry factors on the load sequence, introducing these impacts as exogenous variables into the model, and generating load sequences at different time scales, the present invention improves the accuracy and flexibility of load forecasting. On the basis of the traditional load forecasting model, it considers multi-dimensional exogenous variables and provides more comprehensive and scientific support for load forecasting.

[0062] 4. Through data preprocessing and transformation, the present invention ensures the integrity and accuracy of data by means of interpolation, box plots, etc. By performing PP tests and plotting ACF and PACF diagrams, the stationarity of the model is quantitatively evaluated, the order of the model is determined, and the optimal order of the model is determined by the BIC criterion. Parameter estimation and optimization are carried out by means of maximum likelihood estimation and gradient descent to construct an optimal load sequence model. Residual analysis and RMSE are used to evaluate the accuracy of the model, and the model is optimized according to the evaluation results. Different load sequences on different time scales can be constructed according to different time scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0063] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings. In the drawings:

[0064] Figure 1 is a flowchart of a method for constructing a load sequence scenario considering environmental and industry factors according to an embodiment of the present invention;

[0065] Figure 2 is a schematic diagram of the initial box plot boundaries and outliers in a method for constructing a load sequence scenario considering environmental and industry factors according to an embodiment of the present invention;

[0066] Figure 3 is a box plot after marking outliers in a method for constructing a load sequence scenario considering environmental and industry factors according to an embodiment of the present invention;

[0067] Figure 4 is a box plot after removing outliers in a method for constructing a load sequence scenario considering environmental and industry factors according to an embodiment of the present invention;

[0068] Figure 5 is a schematic diagram of historical load data of a certain place in a method for constructing a load sequence scenario considering environmental and industry factors according to an embodiment of the present invention;

[0069] Figure 6 is a schematic diagram of a daily load sequence in a method for constructing a load sequence scenario considering environmental and industry factors according to an embodiment of the present invention;

[0070] Figure 7 is a schematic diagram of a weekly load sequence in a method for constructing a load sequence scenario considering environmental and industry factors according to an embodiment of the present invention;

[0071] Figure 8It is a schematic diagram of the monthly load series in a method for constructing a load series scenario considering environmental and industry factors according to an embodiment of the present invention. Detailed implementation manners

[0072] In order to enable those skilled in the art of this technology to better understand the technical solutions in this application, the following will clearly and completely describe the technical solutions in the embodiments of this application with reference to the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the protection scope of this application.

[0073] According to an embodiment of the present invention, a method for constructing a load series scenario considering environmental and industry factors is provided.

[0074] Now, the present invention will be further described in combination with the accompanying drawings and specific implementation manners. As Figure 1 shown, the method for constructing a load series scenario considering environmental and industry factors according to an embodiment of the present invention includes:

[0075] S1. Perform data cleaning and data conversion processing on the pre-acquired historical load data and meteorological data to obtain a comprehensive data set;

[0076] As a preferred implementation manner, the performing data cleaning and data conversion processing on the pre-acquired historical load data and meteorological data to obtain a comprehensive data set includes:

[0077] S11. Perform missing value processing on the pre-acquired historical load data and meteorological data;

[0078] S12. Use the box plot method to perform outlier processing on the historical load data and meteorological data after missing value processing;

[0079] S13. Perform timestamp processing, data standardization processing, and unified processing of data time resolution on the historical load data and meteorological data after outlier processing in sequence;

[0080] S14. Merge the historical load data and meteorological data after unified processing of data time resolution to obtain a comprehensive data set.

[0081] It should be added that data cleaning is an important step that must be carried out before time series modeling and prediction. Its main purpose is to ensure the quality and usability of the original data and lay a good foundation for subsequent modeling analysis. The following are the main contents and methods of data cleaning:

[0082] (1) Missing value handling: Time series data may have missing values, and these missing values need to be processed. The processing methods are as follows:

[0083] ① Interpolation method: Use the average value of the previous and subsequent data points or other interpolation methods to fill in the missing values.

[0084] ② Deletion method: Directly delete the records containing missing values (in the case of fewer missing values).

[0085] ③ Filling method: Use zero values or the mean to fill in the missing values.

[0086] (2) Outlier handling: Outliers may be the result of data entry errors or extreme events. Here, the box plot method (quartile method) is used as the outlier handling method. The box plot method is a simple and effective outlier detection method. Its main features are being insensitive to data distribution, simple and intuitive, being able to be used to analyze data sequences with different units and dimensions, and being able to detect outliers in multiple variables or data sets in batches. It is not affected by the data distribution and is easy to understand and apply. The following is its principle:

[0087] The box plot mainly uses the quartiles of the data to identify outliers, and has the following definitions:

[0088] ① The first quartile Q1: 25% of the values in the data are below this value;

[0089] ② The third quartile Q3: 75% of the values in the data are below this value;

[0090] ③ The minimum value: The minimum value in the data set that is not considered an outlier;

[0091] ④ The maximum value: The maximum value in the data set that is not considered an outlier;

[0092] ⑤ The interquartile range IQR: Its value is equal to Q3 - Q1;

[0093] ⑥ Outliers: Values that exceed Q3 + 1.5 * IQR (upper whisker) or are lower than Q1 - 1.5 * IQR (lower whisker).

[0094] As Figures 2 - 4 shown (Boxplot of Original Data refers to the box plot of the original data, Boxplot Afterlabeling Outliers refers to the box plot of the labeled outlier data, Boxplot After Handling Outliers refers to the box plot after removing the outlier data), when using the box plot method to identify outliers, calculate the first quartile, the third quartile, and the interquartile range to determine the values of the upper and lower whiskers. Any value less than the lower whisker or greater than the upper whisker is considered an outlier. Remove these outliers to optimize the accuracy of the data and improve the prediction accuracy.

[0095] Among them, the purpose of data conversion is to convert the original data into a format suitable for modeling and analysis. The detailed steps and methods are as follows:

[0096] (1) Timestamp processing;

[0097] Load data from the data source, convert the timestamp field from a string to a datetime object for time series analysis.

[0098] (2) Data standardization;

[0099] The standardization of data is mainly to convert the data into a format suitable for model processing to improve the stability and prediction performance of the model, and at the same time eliminate the influence of different scales on model training. Here, the stationarity requirements and parameter estimation in the ARIMAX model are mainly considered. The present invention adopts Z-Score standardization to convert the data into a standard normal distribution with a mean of 0 and a standard deviation of 1. The formula is as follows:

[0100]

[0101] Among them, Z represents the value after standardization, X represents the original data, L represents the mean, and σ represents the standard deviation.

[0102] (3) Unification of data time resolution;

[0103] When constructing a load sequence based on the ARIMAX model, unifying the data time resolution is a key step to ensure the accuracy of the model. The following are the detailed steps for unifying the data time resolution:

[0104] ① Determine the resolution;

[0105] Historical load data: Determine the time resolution of the data according to different time granularities respectively.

[0106] Exogenous variable data: Determine the time resolution of exogenous variables (environment, industry) according to different time granularities respectively.

[0107] ② Data resampling;

[0108] Upsampling: When the resolution of the exogenous variable data is higher than the target resolution, use an aggregation method (such as the average value) to convert the high-frequency data into low-frequency data.

[0109] Downsampling: When the resolution of the exogenous variable data is lower than the target resolution, no further processing is required, but it is necessary to ensure that there is no data omission.

[0110] ③ Align the data timestamps;

[0111] Reindex the data using a time index to ensure that the timestamps of all data sources are aligned.

[0112] ④ Data merging;

[0113] Merge the data with unified time resolution into a comprehensive dataset to ensure data consistency and integrity.

[0114] S2. Use the PP unit root test method to check the stationarity of the time series of the comprehensive dataset, and perform differencing on the non-stationary data in the stationarity check to obtain a stationary comprehensive dataset;

[0115] As a preferred implementation, the step of using the PP unit root test method to check the stationarity of the time series of the comprehensive dataset and performing differencing on the non-stationary data in the stationarity check to obtain a stationary comprehensive dataset includes:

[0116] S21. For each time series data in the comprehensive dataset, construct a regression model and use the least squares method to estimate the parameter terms in the regression model. The parameter terms include a constant term, a trend term coefficient, and a lag autoregressive coefficient;

[0117] S22. According to the regression model, calculate the test statistic of the unit root, and compare the critical value in the critical value table with the test statistic. If the test statistic is greater than the critical value, it means that the time series data in the comprehensive dataset is stationary data; otherwise, it means that the time series data in the comprehensive dataset is non-stationary data;

[0118] As a preferred implementation, the step of calculating the test statistic of the unit root according to the regression model includes:

[0119] Calculate the residuals of the time series data in the comprehensive dataset using the regression model;

[0120] Calculate the residual variance based on the time series data residuals, and calculate the standard error based on the residual variance;

[0121] Calculate the test statistic of the unit root using the standard error and the lag autoregressive coefficient.

[0122] S23. Based on the differencing formula, perform differencing on the non-stationary data to convert the non-stationary data into stationary data, and obtain a stationary comprehensive dataset.

[0123] Specifically, data visualization can be used to help understand the basic characteristics of the sequence, and the trend, seasonality, periodicity, and potential anomalies of the time series data can be intuitively seen. For example, Figure 5As shown, a time series plot of the load data is drawn to observe the changes in the data at different time points and whether there are significant trends, seasonal or cyclic fluctuations.

[0124] The Phillips-Perron test (hereinafter referred to as the PP test) is similar to the ADF test and also tests whether a unit root exists in the time series. However, the PP test is more robust to autocorrelation and heteroscedasticity and is more stable than the ADF test when dealing with high-order autocorrelation or heteroscedasticity in the time series. The PP test corrects the problems of heteroscedasticity and autocorrelation in the data by adjusting the standard error in the ADF test and is more suitable for load data with seasonal fluctuations or strong volatility. Therefore, this method uses the PP test to check the stationarity of the time series and perform subsequent processing on non-stationary sequences.

[0125] (1) Unit root;

[0126] In time series analysis, a unit root refers to the situation where the autoregressive coefficient in the time series model is 1; a unit root indicates that the characteristics of the sequence may change significantly over time, that is, the sequence is non-stationary. Specifically, for an autoregressive (AR) model, if the autoregressive coefficient in the model is equal to 1, it means that a unit root exists. This situation will cause the mean and variance of the data to change over time. Here, the PP test is used to check whether the model has a unit root, and differencing is used to ensure the stationarity of the data.

[0127] (2) PP test;

[0128] The assumptions of the PP test are as follows:

[0129] Null hypothesis (H 0 ): The time series has a unit root, which means the sequence is non-stationary.

[0130] Alternative hypothesis (H 1 ): The time series has no unit root, which means the sequence is stationary.

[0131] There are three common forms of the regression model for the PP test. To deal with relatively complex time series in this paper, the form with a constant term, a time trend term, and lag terms is adopted. The expression of the regression model is:

[0132] y t = α + βt + γy t-1 + ε t ;

[0133] In the formula, y t represents the time series data at time t, α represents the constant term, β t represents the time trend term, γ represents the lag autoregressive coefficient, γyt-1 Represents the autoregressive term, that is, it represents the impact of the previous period's value on the current value, used to detect the unit root, ε t Represents the error term;

[0134] Here, the least squares method is used to estimate the parameter terms in the regression model: the constant term α, the trend term coefficient β, and the lag autoregressive coefficient γ, so as to minimize the sum of squared residuals. The following is the process of least squares estimation:

[0135] The sum of squared residuals is defined as:

[0136]

[0137] where, y t Represents the time series data at time t, that is, the actual observed value, (α + γy t-1 + βt) is the predicted value of the regression equation, T represents the length of the time series, and S represents the sum of squared residuals.

[0138] Here, the parameters α, β, and γ are estimated by minimizing the sum of squared residuals;

[0139]

[0140] where, Are the estimated values of the parameters α, β, and γ respectively;

[0141] Specifically expanded, the partial derivative with respect to α:

[0142]

[0143] Set it to zero, and do the same for β and γ in turn, to obtain a set of linear equations about α, β, and γ, and solve them through matrix operations:

[0144]

[0145] where, X represents the matrix containing the independent variables (i.e., y t-1 and t), and y represents the vector containing the dependent variable (i.e., y t ).

[0146] In the PP test, the corrected standard error is a key step to eliminate the effects of autocorrelation and heteroscedasticity. The correction method of the PP test does not rely on the traditional white noise assumption for the residuals, but uses a stationary correction statistic. Calculate the corrected standard error:

[0147]

[0148] where, Represents the standard error, σ 2 Represents the variance of the residuals;

[0149] The calculation formula for the variance of the residuals is:

[0150]

[0151] The residual formula is:

[0152]

[0153] Calculate the test statistic:

[0154]

[0155] After obtaining the statistic, consult the Phillips-Perron critical value table to obtain the corresponding critical value, compare the size with the critical value. If the statistical value is greater than the critical value, reject the original hypothesis and consider the time series to be stationary. Otherwise, consider that the time series has a unit root, that is, it is non-stationary.

[0156] (3) Processing of non-stationary time series;

[0157] If the statistical value of the PP test is less than the critical value, the original hypothesis cannot be rejected, and it is considered that the time series has a unit root and is non-stationary. For non-stationary data, differencing processing is required.

[0158] The differencing formula is:

[0159] y' t = y t - y t-1 ;

[0160] In the formula, y' t represents the data after the first differencing, y t represents the time series data at time t, and y t-1 represents the time series data at time t - 1.

[0161] In addition, differencing processing may not necessarily make the time series stable, and it is necessary to perform the PP test on the time series again until the time series is stationary.

[0162] S3. Construct an ARIMAX model based on exogenous variables, and train and optimize the ARIMAX model using the stationary comprehensive data set to obtain the final ARIMAX model;

[0163] As a preferred implementation manner, the constructing an ARIMAX model based on exogenous variables, and training and optimizing the ARIMAX model using the stationary comprehensive data set to obtain the final ARIMAX model includes:

[0164] S31. Construct an ARIMAX model based on the basic time series model and combined with exogenous variables;

[0165] S32. Based on the statistical characteristics of the time series data in the comprehensive dataset, determine the order of the ARIMAX model and the lag period of exogenous variables by means of stepwise regression combined with the BIC information criterion;

[0166] As a preferred implementation manner, the determining the order of the ARIMAX model and the lag period of exogenous variables by means of stepwise regression combined with the BIC information criterion based on the statistical characteristics of the time series data in the comprehensive dataset includes:

[0167] S321. Based on the time series data in the comprehensive dataset, determine the order range of autoregression and moving average by establishing an autocorrelation plot and a partial autocorrelation plot;

[0168] S322. According to the order range of autoregression and moving average, initially set the order of autoregression and moving average of the ARIMAX model;

[0169] S323. Adopt the stepwise regression method, gradually introduce the lag period of exogenous variables, and calculate the BIC value introduced in each step according to the BIC information criterion, and select the lag period of exogenous variables that minimizes BIC;

[0170] S324. Perform backward elimination on the initially selected model, remove the lag period of exogenous variables one by one, and finally determine the order of the optimal ARIMAX model and the lag period of exogenous variables according to the BIC information criterion.

[0171] S33. Estimate the parameters of the ARIMAX model and optimize the model parameters by using the simulated annealing algorithm;

[0172] As a preferred implementation manner, the estimating the parameters of the ARIMAX model includes: estimating the moving average parameters, autoregressive parameters, and exogenous variable parameters;

[0173] S34. Use the stationary comprehensive dataset to train and evaluate the ARIMAX model with optimized order, lag period of exogenous variables, and parameters, and optimize the ARIMAX model according to the evaluation results to obtain the final ARIMAX model.

[0174] Specifically, ARIMA captures the changing trend in the time series by combining the autoregressive (AR) and moving average (MA) methods for prediction. The ARIMA model mainly consists of three parts: the autoregressive (AR) part, the differencing (I) part, and the moving average (MA) part. The following is a detailed analysis of constructing ARIMA:

[0175] (1) Autoregressive part (AR);

[0176] The autoregressive part uses the past values of the time series itself to predict the current value. Autoregression means that there is a certain linear relationship between the current value and its values at several previous time points. The order (p) of the autoregressive part indicates how many past time points' values are used for prediction. Its expression is:

[0177]

[0178] In the formula, Y t represents the value of the time series at time t, represents the autoregressive coefficient, and ε t represents the white noise error term, indicating an unpredictable random disturbance.

[0179] (2) Differencing part (I);

[0180] The differencing part is used to convert a non-stationary time series into a stationary time series. The mean and variance of a stationary time series do not change over time, so it is suitable for modeling and prediction. The differencing operation is achieved by calculating the differences between adjacent time points.

[0181] Its expression is:

[0182] Y' t = Y t - Y t-1 ;

[0183] where Y' t represents the value of the differenced time series at time t, Y t represents the value of the time series at time t. Y t-1 represents the value of the time series at time t - 1.

[0184] (3) Moving average part (MA);

[0185] The moving average part uses the past prediction errors of the time series to predict the current value. The moving average model means that there is a certain linear relationship between the current value and the prediction errors at several previous time points. The order (q) of the moving average model indicates how many past time points' errors are used for prediction. Its expression is:

[0186] Y t = ε t + θ 1 ε t-1 + θ 2 ε t-2 +... + θ q ε t-q ;

[0187] where Y t represents the value of the time series at time t, εt , ε t-1 ,..., ε t-q represents the prediction error, θ 1 , θ 2 ,..., θ q represents the moving average coefficient.

[0188] (4) The mathematical expression of the basic ARIMA model;

[0189]

[0190] In the formula, Y t represents the value of the time series at time t, ε t , ε t-1 ,..., ε t-q represents the prediction error, θ 1 , θ 2 ,..., θ q represents the moving average coefficient, represents the coefficient of the autoregressive part, Δ d Y t represents the time series after d-order differencing.

[0191] The basic ARIMA model makes the time series stationary through differencing operations, which is beneficial to the fitting and prediction of the model. The autoregressive part and the moving average part respectively use the past values and prediction errors of the time series for modeling, and the final prediction result is obtained after combination.

[0192] In addition, the basic time series model cannot well reflect the influence of exogenous variables such as weather, season, holidays, etc. on the load series. Therefore, this method extends the basic time series model and uses the ARIMAX model to handle the influence of exogenous variables. Its model can be expressed as:

[0193]

[0194] In the formula, y t represents the time series data at time t, μ represents the constant term, represents the coefficient of the autoregressive (AR) part, θ j represents the coefficient of the moving average (MA) part, X t-k represents the value of the exogenous variable matrix at lag k, β k represents the coefficient of the exogenous variable lag k, ε tDenote the error term, p represents the lag order of the autoregressive (AR) part, indicating how many past target values are used to predict the current value, q represents the order of the moving average (MA) part, indicating how many past error terms are used to predict the current error term, r represents the lag order of the exogenous variable part, indicating how many past exogenous variables are used to predict the target time series, j represents the lag term of the moving average (MA) part, i represents the lag term of the autoregressive (AR) part, and k represents the lag term of the exogenous variable part.

[0195] Here, in the basic ARIMA model, environmental factors and industry characteristics are considered as exogenous variables, and the ARIMAX model is constructed for training and optimization to improve the model's ability to explain sequence fluctuations and prediction accuracy, and to realize the construction of the load sequence.

[0196] Specifically, the selection of the model order;

[0197] Based on the time series diagram and the statistical characteristics of the time series, the autocorrelation function (ACF) and partial autocorrelation function (PACF) are plotted to explore the autocorrelation in the load data and exclude the influence of other lag periods, and to determine the approximate range of the orders of the autoregressive and moving average parts of the model:

[0198] When the significance of the PACF diagram disappears after lag P, P is a candidate value for the autoregressive order. When the significance of the ACF diagram disappears after lag q, q is a candidate value for the moving average order.

[0199] On this basis, the present invention uses the method of stepwise regression combined with the BIC information criterion to select the appropriate lag period of the exogenous variable and finally determine the order of the model. The following is the detailed process:

[0200] The formula of the BIC criterion is as follows:

[0201]

[0202] In the formula, n represents the sample size, that is, the number of observed data points in the time series; represents the estimated variance of the model residuals, that is, the variance of the residuals after the model fitting. For the time series model, the formula for the residuals is:

[0203]

[0204] In the formula, y t represents the true value, represents the predicted value;

[0205] p is the number of parameters estimated in the model (the number of free parameters). For the ARIMAX model, this includes all the parameters to be estimated, such as the AR coefficients, MA coefficients, exogenous variable coefficients, etc.

[0206] In the process of gradually introducing exogenous variables, refer to the above method to calculate the estimated variance of the model residuals without adding each exogenous variable using the least squares method, and calculate the BIC value according to the BIC formula. If the addition of the exogenous variable lag period results in a decrease in the BIC value, accept this lag period and continue to select the next lag period. If the addition of the exogenous variable lag period causes an increase in the BIC value, stop the forward selection (that is, gradually introduce variables until there are no variables that can further significantly improve the model performance); based on the existing model, eliminate one exogenous variable lag period one by one. For each model after elimination, recalculate the BIC value. If the BIC value decreases after eliminating a certain lag period, retain this elimination operation. If the BIC value increases after eliminating a certain lag period, restore this lag period and keep the current model.

[0207] Through the combination of forward selection and backward elimination, finally select the model with the minimum BIC value. This is the optimal model order and the number of exogenous variable lag periods obtained through optimization and screening.

[0208] Specifically, the model parameter estimation includes:

[0209] (1) Moving average (MA) parameter estimation;

[0210] Y t =θ 1 ε t-1 +θ 2 ε t-2 +...+θ q ε t-q +δ t ;

[0211] In the formula, Y t represents the load value at time t, θ 1 , θ 2 ,…, θ q represent the autoregressive coefficients, and δ t represents white noise.

[0212] The autocorrelation function is:

[0213]

[0214] In the formula, Y t represents the load value at time t, Y t-k represents the load value at the lagged k time, ρ k represents the autocorrelation coefficient at the lag period k, Cov represents covariance, and Var represents variance.

[0215] Calculate the correlation coefficient through the Yule - Walker equation. The following is the detailed calculation process:

[0216] First, calculate the autocovariance Cov(Y t , Y t-k ), that is, the covariance between Y t and Y t-k .

[0217]

[0218] In the formula, Cov(Y t , Y t-k ) is the autocovariance at lag k, Y i represents a sample of the sequence, represents the sequence mean, and T represents the total number of samples.

[0219] The calculation process of autocovariance needs to be calculated for each lag k, usually starting from k = 1 until the maximum lag required for calculation.

[0220] Similarly, calculate the variance:

[0221]

[0222] In the formula, Var(Y t ) represents the variance, Y i represents a sample of the sequence, represents the sequence mean, and T represents the total number of samples.

[0223] The Yule - Walker equation estimates the autoregressive coefficients by solving these linear equations, and these coefficients describe the autoregressive structure of the time series.

[0224] (2) Autoregressive (AR) parameter estimation;

[0225] The model of the autoregressive part is:

[0226]

[0227] In the formula, Y t represents the load value at time t, represents the autoregressive coefficient, and δ t represents white noise.

[0228] Calculate the partial autocorrelation coefficient:

[0229]

[0230] In the formula, Y t represents the load value at time t, Y t-krepresents the load value at the time lag k, represents the partial autocorrelation coefficient, Cov(y t ,y t-k |{y t-1 ,y t-2 ,...,y t-(k-1)}) represents the conditional covariance, which means that after controlling the influence of the lag period 1, 2, ..., k-1, the current moment y t and the value y of the lag period k t-k The correlation between Var(y t ) and Var(y t-k |y t-1 ,y t-2 ,...,y t-(k-1) ) represents the variance at time t and lag k after controlling for the effects of the intermediate lag period.

[0231] (3) Estimation of exogenous variable parameters (X);

[0232] The exogenous variable parameter estimation adopts the ridge regression algorithm, the goal of the ridge regression algorithm is to minimize the loss function with a regularization term, and the loss function expression with a regularization term is:

[0233]

[0234] In the formula, ι(β) represents the loss function with regularization term, x t represents the observation vector of the exogenous variable at time point t (the values ​​of all exogenous variables corresponding to time point t form a vector), β represents the regression coefficient corresponding to the exogenous variable (β=[β 1 ,β 2 ,...,β m ] T ), λ represents the regularization parameter, which is used to control the size of the regression coefficient (to prevent overfitting); T represents the number of samples in the data set, that is, the number of observations, j represents the index of the regression coefficient in the regularization term, and m represents the number of features in the ARIMAX model.

[0235] In order to find the optimal regression coefficient β, the loss function is minimized, that is, the partial derivative of the loss function with respect to β is taken and set to zero:

[0236]

[0237] In the formula, X represents the T×m exogenous variable matrix (each row is the exogenous variable observation vector at a time point), and y represents the target variable y t A vector (T×1), I represents an m×m identity matrix; λ represents the regularization parameter.

[0238] The coefficient estimates of the exogenous variables can be obtained through calculation.

[0239] In addition, for the parameter optimization of the three parts of autoregressive (AR), moving average (MA), and exogenous variables (X), the present invention uses a simulated annealing algorithm for joint optimization to optimize all the parameters of the AR part, MA part, and exogenous variable part.

[0240] For the joint optimization of the parameters of the AR, MA, and X parts, a target function is defined, which measures the prediction error of the model:

[0241]

[0242] In the formula, J(η) represents the prediction error of the model, y t represents the actual value, represents the prediction calculated by the model (according to the current parameter η).

[0243] In this target function, is the parameter vector to be optimized.

[0244] Among them, the optimization process includes the following steps:

[0245] (1) Initialization;

[0246] Set the initialization parameter η 0 ;

[0247] The temperature decay factor α, 0.95 < α < 0.99;

[0248] The initial temperature T 0 is 10.

[0249] (2) Generate a new solution by perturbation;

[0250] Generate a new parameter η k by adding a perturbation to the current parameter η k+1 : η k+1 = η k + δ;

[0251] where δ - N(0,σ), here δ is random noise and σ is the standard deviation of the noise.

[0252] (3) Calculate the loss function;

[0253] For each new solution η k+1 , calculate the loss function:

[0254]

[0255] where represents the predicted value calculated with the current parameters.

[0256] (4) Probability of accepting a new solution;

[0257] Calculate the difference ΔJ in the objective function between the current solution and the new solution: ΔJ = J(η k+1 ) - J(η k ). If the objective function value of the new solution is smaller (i.e., ΔJ < 0), directly accept the new solution. If the objective function value of the new solution is larger, accept the new solution with a certain probability. The acceptance probability is:

[0258]

[0259] where P(accept) represents the acceptance probability, and T k represents the current temperature. As the iteration progresses, the temperature gradually decreases, and the decreasing temperature makes the algorithm gradually concentrate near the local optimal solution.

[0260] (5) Temperature decay;

[0261] Update the temperature T k , and usually use the following formula for decay:

[0262] T k+1 = αT k ;

[0263] where α is the decay factor.

[0264] In addition, model evaluation includes residual analysis and RMSE (Root Mean Squared Error). The following is the detailed process:

[0265] The residual is the difference between the model prediction value and the true value. If the model fits the data well, then the residuals should conform to white noise, that is, the residuals should be independent and identically distributed random variables with a mean of zero and a constant variance.

[0266] Calculate the residual for each point:

[0267]

[0268] In the formula, N t represents the residual, y t represents the true value, represents the predicted value.

[0269] Check whether the residuals are white noise by looking at the autocorrelation function (ACF) and partial autocorrelation function (PACF) of the residuals. In an ideal situation, if the model fits well, the residuals should have no significant autocorrelation, that is, their ACF and PACF plots should decay rapidly to 0 after a lag of 1.

[0270] Autocorrelation function:

[0271]

[0272] Where ρ k represents the autocorrelation coefficient at lag k, and y k represents each observation in the sequence, and represents the mean of the sequence.

[0273] Partial autocorrelation coefficient:

[0274]

[0275] Where PACF(k) represents the partial autocorrelation coefficient, Cov(ε t , ε t-k |ε t-1 ,......, ε t-(k-1) ) represents the covariance between the t-th and (t - k)-th moments of the time series under the condition of controlling the first k - 1 lags (i.e., the covariance after removing the influence of other lags), Var(ε t ) represents the variance of the time series at time t, Var(ε t-k ) represents the variance of the time series at time t - k, and k represents the lag order, indicating the time interval relative to the time t before the current time.

[0276] If the ACF and PACF of the residuals have no significant correlation, then the residuals of the model can be considered white noise, indicating that the model has well captured the structure of the data.

[0277] Use the Q statistic (Ljung - Box) to quantitatively test whether the residuals are white noise, and its statistic formula is:

[0278]

[0279] Where Q represents the statistic, n represents the sample size, m represents the number of lag periods, represents the autocorrelation coefficient at lag k.

[0280] If the p - value of the Q value is less than the predetermined significance level (e.g., 0.05), then the null hypothesis is rejected, indicating that there is autocorrelation in the residuals and the model needs to be optimized.

[0281] Among them, RMSE calculates the root mean square of the difference between the predicted value and the actual value. The smaller the RMSE, the stronger the prediction ability of the model:

[0282]

[0283] Where RMSE represents the root mean square of the difference between the predicted value and the actual value, n represents the sample size, and y trepresents the true value, represents the predicted value.

[0284] In addition, model optimization will adjust the model parameters according to the evaluation results to achieve better prediction performance and generalization ability. Parameter adjustment: Adjust the orders of the AR, MA, and exogenous variable parts according to the results of residual analysis and AIC / BIC metrics;

[0285] Exogenous variable selection: Select or remove exogenous variables according to the fitting effect of the model to optimize the model;

[0286] Cross-validation: Use the training set and validation set for cross-validation to check the performance of the model on new data. Specifically, it includes:

[0287] (1) Rolling Window Validation: In the time series, fix the size of the training set, roll forward one time step each time, update the training set and validation set, retrain the model, and evaluate the performance.

[0288] (2) Time series split validation: Divide the time series data into multiple training sets and validation sets to ensure that the test set only contains data from future times and avoid data leakage.

[0289] S4. Use the final ARIMAX model to fit ARIMAX models with different time scales to obtain load sequences with different time scales.

[0290] As a preferred implementation, the ARIMAX models with different time scales include: daily time scale ARIMAX model, weekly time scale ARIMAX model, monthly time scale ARIMAX model, and annual time scale ARIMAX model;

[0291] It should be noted that in scenario generation, considering the generation of multi-time scale load sequences with four time scales of day, week, month, and year can more comprehensively understand the change law of the load and effectively predict and optimize the scheduling of the power system from the short-term to the long-term level. The following is the detailed process:

[0292] Based on the constructed ARIMAX model, fit the ARIMAX models of day, week, month, and year respectively to generate load sequences with different time scales;

[0293] The expression of the daily time scale ARIMAX model is:

[0294]

[0295] The expression of the weekly time scale ARIMAX model is:

[0296]

[0297] The expression of the monthly - scale ARIMAX model is as follows:

[0298]

[0299] The expression of the annual - scale ARIMAX model is as follows:

[0300]

[0301] In the formula, represents the load sequence generated on the daily - scale, represents the load sequence generated on the weekly - scale, represents the load sequence generated on the monthly - scale, represents the load sequence generated on the annual - scale, μ represents the constant term, represents the autoregressive part on the daily - scale, represents the autoregressive part on the weekly - scale, represents the autoregressive part on the monthly - scale, represents the autoregressive part on the annual - scale, represents the moving - average part on the daily - scale, represents the moving - average part on the weekly - scale, represents the moving - average part on the monthly - scale, represents the moving - average part on the annual - scale, represents the exogenous variable on the daily - scale, represents the exogenous variable on the weekly - scale, represents the exogenous variable on the monthly - scale, represents the exogenous variable on the annual - scale, ε t represents the error term, represents the coefficient of the autoregressive part, θ i represents the coefficient of the moving - average part, β k represents the coefficient of the exogenous variable lag k, p represents the lag order of the autoregressive (AR) part, which represents how many past target values are used to predict the current value, q represents the order of the moving - average (MA) part, which represents how many past error terms are used to predict the current error term, j represents the lag term of the moving - average (MA) part, r represents the lag order of the exogenous variable part, which represents how many past exogenous variables are used to predict the target time series.

[0302] Daily Scale: Mainly considering the load fluctuations within a day, it is used for short - term load scheduling and prediction;

[0303] Weekly Scale: Focus on the weekly load change pattern, used for daily load scheduling;

[0304] Monthly Scale: Pay attention to seasonal changes, used for medium-term load forecasting and seasonal planning;

[0305] Yearly Scale: Focus on the annual load change trend, used for long-term planning.

[0306] Embodiment

[0307] Taking the historical load data, meteorological data and data of various industries in a certain area for two years as input, a multi-time load sequence model considering environmental and industry factors is constructed respectively, and load sequences of multiple time scales are output, such as Figures 6 - 8 shown (Prediction represents the predicted load sequence curve, Actual represents the actual load sequence curve), which is the load sequence under the daily-weekly-monthly multi-time scale.

[0308] In the above scenario, taking the single-day, weekly and monthly load sequences in February 2024 in a certain place as examples respectively, the comparison curves of the predicted load sequence and the actual load sequence are obtained, and the deviation value of a single point is calculated by the following formula:

[0309]

[0310] In the formula, ε i represents the single-point deviation, y t represents the true value, represents the predicted value.

[0311] Through statistical calculation results, 92.1%, 91%, and 89.7% of the single-point deviations under different time scales are less than 0.2, indicating that in most cases, the prediction error of the model is relatively small, and the difference between the predicted value and the actual value remains within a relatively acceptable range. The model has high accuracy in short-term and medium-term load forecasting and can better fit the actual load change trend. In addition, as the time scale increases, the prediction error slightly expands but still remains within a reasonable tolerance range, showing the robustness and stability of the model under different time scales.

[0312] In summary, by means of the above technical solutions of the present invention, the present invention comprehensively considers factors such as climate change and industry demands, improves the accuracy and reliability of load forecasting, solves the problem that existing load forecasting methods inadequately consider external influencing factors, and improves the accuracy and reliability of load forecasting. The present invention aims to construct load sequence scenarios, comprehensively considers the impacts of environmental factors and industry elements on load changes, introduces environmental factors of meteorological changes and seasonal characteristics of industry demands, establishes a load forecasting framework based on the ARIMAX model, accurately models historical load data, and constructs accurate load sequences in combination with external factors; integrates multi-dimensional constraint conditions, supports the generation and optimization of multiple load scenarios, and can provide accurate load forecasting and decision-making support for power system scheduling, demand response, etc. On the basis of the basic time series model, the present invention integrates the impacts of external environmental factors and industry factors on the load sequence, introduces these impacts as exogenous variables into the model, and generates load sequences of different time scales, improving the accuracy and flexibility of load forecasting. On the basis of the traditional load forecasting model, the present invention considers multi-dimensional exogenous variables, providing more comprehensive and scientific support for load forecasting. The present invention ensures the integrity and accuracy of data through data preprocessing and conversion, using methods such as interpolation and box plots, quantitatively evaluates the stationarity of the model through PP tests and drawing ACF and PACF graphs, determines the order of the model, determines the optimal order of the model through the BIC criterion, performs parameter estimation and optimization using maximum likelihood estimation and gradient descent, constructs the optimal load sequence model, evaluates the accuracy of the model using residual analysis and RMSE, and optimizes the model according to the evaluation results. Different load sequences of different time scales can be constructed according to different time scenarios.

[0313] Those skilled in the art should understand that the embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, optical storage, etc.) containing computer-usable program code.

[0314] The specific embodiments described above have further elaborated on the purpose, technical solutions, and beneficial effects of the present invention. It should be understood that the above description is only for the specific embodiments of the present invention and is not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention should be included in the protection scope of the present invention.

Claims

1. A method for constructing a load sequence scenario taking into account environmental and industry factors, characterized in that: include: S1. Perform data cleaning and data conversion on the historical load data and meteorological data acquired in advance to obtain a comprehensive data set; S2. Use the PP unit root test method to check the stationarity of the time series of the comprehensive data set, and perform differential processing on the non-stationary data in the stationarity check to obtain a stationary comprehensive data set; S3, constructing an ARIMAX model based on exogenous variables, and using a stationary comprehensive data set to train and optimize the ARIMAX model to obtain the final ARIMAX model; S4. Using the final ARIMAX model, fit the ARIMAX models of different time scales to obtain the load series of different time scales.

2. A method for constructing a load sequence scenario taking into account environmental and industry factors according to claim 1, characterized in that: The previously acquired historical load data and meteorological data are cleaned and converted to obtain a comprehensive data set including: S11, processing missing values ​​of the historical load data and meteorological data acquired in advance; S12, using the box plot method to process the outliers of the historical load data and meteorological data after the missing values ​​are processed; S13, performing time stamp processing, data standardization processing and unified processing of data time resolution on the historical load data and meteorological data after abnormal value processing; S14. The historical load data processed uniformly with the data time resolution are merged with the meteorological data to obtain a comprehensive data set.

3. A method for constructing a load sequence scenario taking into account environmental and industry factors according to claim 1, characterized in that: The PP unit root test method is used to perform a time series stationarity check on the comprehensive data set, and the non-stationary data in the stationarity check is subjected to differential processing to obtain a stationary comprehensive data set including: S21. For each time series data in the comprehensive data set, a regression model is constructed, and the parameter items in the regression model are estimated using the least squares method, wherein the parameter items include a constant item, a trend item coefficient, and a lagged autoregressive coefficient; The expression of the regression model is: y t =α+βt+γy t-1 +e t ; In the formula, y t represents the time series data at time t, α represents the constant term, β t represents the time trend term, γ represents the lagged autoregressive coefficient, γy t-1 represents the autoregressive term, ε t represents the error term; S22. According to the regression model, the test statistic of the unit root is calculated, and the critical value in the critical value table is used to compare with the test statistic. If the test statistic is greater than the critical value, it means that the time series data in the comprehensive data set is stationary data; otherwise, it means that the time series data in the comprehensive data set is non-stationary data. S23. Based on the differential formula, the non-stationary data is differentially processed to convert the non-stationary data into stationary data to obtain a comprehensive data set of stationarity.

4. A method for constructing a load sequence scenario taking into account environmental and industry factors according to claim 3, characterized in that: The test statistic for calculating the unit root according to the regression model includes: Use regression models to calculate residuals of time series data in a synthetic dataset; Calculate the residual variance based on the time series data residuals, and calculate the standard error based on the residual variance; The test statistic for the unit root is calculated using the standard error and the lagged autoregressive coefficient.

5. A method for constructing a load sequence scenario taking into account environmental and industry factors according to claim 3, characterized in that: The difference formula is: and' t =and t -and t-1 ; In the formula, y′ t Represents the data after a difference, y t represents the time series data at time t, y t-1 Represents the time series data at time t-1.

6. A method for constructing a load sequence scenario taking into account environmental and industry factors according to claim 1, characterized in that: The ARIMAX model is constructed based on exogenous variables, and the ARIMAX model is trained and optimized using a stationary comprehensive data set to obtain the final ARIMAX model including: S31. Based on the basic time series model, the ARIMAX model is constructed in combination with exogenous variables; S32. Based on the statistical characteristics of the time series data in the comprehensive data set, the order of the ARIMAX model and the number of lags of the exogenous variables are determined by using stepwise regression combined with the BIC information criterion; S33, estimating the parameters of the ARIMAX model and optimizing the model parameters using a simulated annealing algorithm; S34. Use a comprehensive data set of stationarity to train and evaluate the ARIMAX model after the order, lag period of exogenous variables and parameters are optimized, and optimize the ARIMAX model based on the evaluation results to obtain the final ARIMAX model.

7. A method for constructing a load sequence scenario taking into account environmental and industry factors according to claim 6, characterized in that: The expression of the ARIMAX model is: In the formula, y t represents the time series data at time t, μ represents the constant term, represents the coefficient of the autoregressive part, θ j represents the coefficient of the moving average part, X t-k represents the value of the exogenous variable matrix at the lag period k, β k represents the coefficient of the exogenous variable at the lag period k, ε t represents the error term, p represents the lag order of the autoregressive part, q represents the order of the moving average part, r represents the lag order of the exogenous variable part, j represents the lag term of the moving average part, i represents the lag term of the autoregressive part, and k represents the lag term of the exogenous variable part.

8. A method for constructing a load sequence scenario taking into account environmental and industry factors according to claim 6, characterized in that: The method of determining the order of the ARIMAX model and the number of lag periods of the exogenous variables by using stepwise regression combined with the BIC information criterion based on the statistical characteristics of the time series data in the comprehensive data set includes: S321. Based on the time series data in the comprehensive data set, determine the order range of autoregression and moving average by establishing an autocorrelation diagram and a partial autocorrelation diagram; S322. Preliminarily set the order of autoregression and moving average of the ARIMAX model according to the order range of autoregression and moving average; S323, using the stepwise regression method, gradually introduce the lag period of the exogenous variable, and calculate the BIC value introduced at each step according to the BIC information criterion, and select the lag period of the exogenous variable that minimizes the BIC; S324. Perform backward elimination on the initially selected model, remove the lag period of the exogenous variables one by one, and finally determine the optimal ARIMAX model order and the number of lag periods of the exogenous variables based on the BIC information criterion.

9. A method for constructing a load sequence scenario taking into account environmental and industry factors according to claim 6, characterized in that: The estimation of the parameters of the ARIMAX model includes: moving average parameter estimation, autoregressive parameter estimation and exogenous variable parameter estimation; The exogenous variable parameter estimation adopts the ridge regression algorithm, the goal of the ridge regression algorithm is to minimize the loss function with a regularization term, and the loss function expression with a regularization term is: In the formula, ι(β) represents the loss function with regularization term, x t represents the observation vector of the exogenous variable at time point t, β represents the regression coefficient corresponding to the exogenous variable, λ represents the regularization parameter; T represents the number of observations, j represents the index of the regression coefficient in the regularization term, and m represents the number of features in the ARIMAX model.

10. A method for constructing a load sequence scenario taking into account environmental and industry factors according to claim 1, characterized in that: The ARIMAX models of different time scales include: a daily time scale ARIMAX model, a weekly time scale ARIMAX model, a monthly time scale ARIMAX model and a grade time scale ARIMAX model; The expression of the daily time scale ARIMAX model is: The expression of the weekly time scale ARIMAX model is: The expression of the monthly time scale ARIMAX model is: The expression of the grade time scale ARIMAX model is: In the formula, represents the load sequence generated on a daily time scale, represents the load sequence generated on a weekly time scale, represents the load sequence generated on a monthly time scale, represents the load sequence generated at the grade time scale, μ represents the constant term, represents the autoregressive part on a daily time scale, represents the autoregressive part at the weekly time scale, represents the autoregressive part at the monthly time scale, represents the autoregressive part at the grade time scale, represents the moving average part on the daily time scale, represents the moving average part on a weekly time scale, represents the moving average part on a monthly time scale, represents the moving average part in the grade time scale, represents the exogenous variable at the daily time scale, represents the exogenous variable at the weekly time scale, represents the exogenous variable at the monthly time scale, represents the exogenous variable at the grade time scale, ε t represents the error term, represents the coefficient of the autoregressive part, θ i represents the coefficient of the moving average part, β k represents the coefficient of the lag period k of the exogenous variable, p represents the lag order of the autoregressive part, q represents the order of the moving average part, r represents the lag order of the exogenous variable part, j represents the lag term of the moving average part, i represents the lag term of the autoregressive part, and k represents the lag term of the exogenous variable part.

Citation Information

Cited By

  • Carbon emission monitoring method and system fusing decomposition regression, electronic equipment and medium

    CN121352247A