A method for predicting short-term output power of distributed power sources

CN117650511BActive Publication Date: 2026-08-14STATE GRID SHANDONG ELECTRIC POWER CO MARKETING SERVICE CENT (MEASURING CENT) +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-11-20
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

[0007]有鉴于此,本发明实施例提供了一种分布式电源短期出力功率预测方法,以解决如何优化利用历史用电负荷数据对ARIMA模型进行训练的训练结果,以提高分布式电源短期出力功率的预测结果的问题

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117650511B_ABST
    Figure CN117650511B_ABST
Patent Text Reader

Abstract

This invention relates to the field of data processing technology, specifically to a method for predicting the short-term output power of distributed power sources. The method acquires historical electricity load time-series data for a region targeted by any distributed power source within a preset time period; divides the preset time period into at least one cycle using a preset time interval; selects data from any cycle of the historical electricity load time-series data as the cycle time-series data; and obtains the noise anomaly level of the time-series data for each cycle. Based on the noise anomaly level of the time-series data for each cycle and the historical electricity load time-series data, an ARIMA model is trained to obtain a trained ARIMA model for predicting the short-term output power of the distributed power source. The error weights during ARIMA model training are optimized based on the noise anomaly level, improving the training results of the ARIMA model and thus improving the accuracy of the short-term output power prediction results of the distributed power source.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and specifically to a method for predicting the short-term output power of distributed power sources. Background Technology

[0002] Distributed power generation refers to small, discrete energy generation systems, typically deployed directly at the energy end-use location. Distributed power generation usually uses renewable energy or clean energy sources such as fuel cells, and can meet energy needs in any scenario, thereby reducing reliance on traditional centralized power plants, minimizing transmission losses, and promoting the sustainable use of energy.

[0003] Since distributed power generation relies on energy consumption to generate electricity, and the daily electricity demand and the electricity demand at different times of day are different, in order to meet the electricity demand of the area targeted by the power source at all times and to minimize energy waste, it is often necessary to predict the electricity load in the area. Based on the prediction results, energy use can be adjusted in real time, thereby achieving the goal of predicting and optimizing the short-term output power of distributed power sources.

[0004] In existing technologies, historical electricity load data of the area to be predicted is obtained, and the Autoregressive Integrated Moving Average (ARIMA) model is trained using the historical electricity load data to obtain a trained ARIMA model. Then, the trained ARIMA model is used to predict the short-term output power of distributed power sources.

[0005] However, when acquiring historical electricity load data, due to issues such as sensor malfunctions, a large amount of noisy data is present in the acquired historical electricity load data. This noisy data is similar in data fluctuation to the abnormal electricity load data caused by sudden events in the area to be predicted (e.g., prolonged excessive load). However, the latter abnormal electricity load data is of reference value for training the ARIMA model, while the former noisy data interference is not of reference value. Consequently, the ARIMA model trained using historical electricity load data has poor predictive ability.

[0006] Therefore, how to optimize the training results of ARIMA models using historical electricity load data to improve the prediction results of short-term power output of distributed power sources has become an urgent problem to be solved. Summary of the Invention

[0007] In view of this, embodiments of the present invention provide a method for predicting the short-term output power of distributed power sources, in order to solve the problem of how to optimize the training results of ARIMA models using historical electricity load data, so as to improve the prediction results of the short-term output power of distributed power sources.

[0008] This invention provides a method for predicting the short-term output power of distributed power sources, which includes the following steps:

[0009] Obtain historical electricity load time-series data for any distributed power source's target area within a preset time period;

[0010] Based on the local data differences of each data in the historical electricity load time series data, the normal index of electricity load fluctuation in the region targeted by the distributed power source during the preset time period is obtained.

[0011] The preset time period is divided into at least one cycle using a preset time interval. Data within any cycle is taken from the historical electricity load time series data as the cycle time series data. For any cycle time series data, the noise coefficient of each data is obtained based on the local data of each data in the cycle time series data. Based on the noise coefficient of each data in the cycle time series data, the time series fluctuation index of the electricity load in the corresponding cycle is obtained.

[0012] The similarity between the periodic time series data and each other periodic time series data is obtained. Based on the similarity, the time series fluctuation index and the normal fluctuation index of electricity load, the noise anomaly degree of the periodic time series data is obtained.

[0013] The ARIMA model is trained based on the noise anomaly level of each periodic time series data and the historical electricity load time series data to obtain a trained ARIMA model. The trained ARIMA model is then used to predict the short-term output power of the distributed power source.

[0014] Furthermore, the step of obtaining the normal fluctuation index of electricity load in the region targeted by the distributed power source within the preset time period based on the local data differences of each data point in the historical electricity load time series data includes:

[0015] Obtain the first average value of the historical electricity load time-series data;

[0016] For any data in the historical electricity load time series data, take the data as the center, obtain all data contained in a preset size window, obtain the data difference between the maximum data and the minimum data based on all data contained in the preset size window, and count the number of extreme data in the preset size window;

[0017] Calculate the absolute value of the difference between the first data mean and the data, normalize the data difference to obtain the corresponding normalization result, perform negative mapping on the number of extreme data to obtain the corresponding mapping result, and obtain the first difference between the first preset value and the mapping result.

[0018] The product of the absolute value of the difference, the normalization result, and the first difference is obtained as the fluctuation characteristic value of the data.

[0019] Obtain the fluctuation characteristic value of each data point in the historical electricity load time series data, calculate the mean of all fluctuation characteristic values, and use the result after normalizing the mean as the normal fluctuation index of electricity load in the region targeted by the distributed power source during the preset time period.

[0020] Furthermore, the step of obtaining the noise coefficient of each data point based on the local data of each data point in the periodic time series data includes:

[0021] Obtain the extreme value data from the periodic time series data;

[0022] For any data in the periodic time series data, with the data as the center, obtain the extreme value data contained in a preset size window and count the first number of extreme value data. Perform negative mapping on the first number to obtain the corresponding mapping value, and obtain the difference between the target value and the mapping value.

[0023] Based on the extreme value data contained within the preset size window, the time span between each pair of adjacent extreme value data is obtained, and the standard deviation between all time spans is calculated.

[0024] Obtain the product of the difference and the standard deviation, and use the normalized result of the product as the noise coefficient of the data.

[0025] Furthermore, the step of obtaining the time-series fluctuation index of electricity load within the corresponding period based on the noise coefficient of each data point in the periodic time-series data includes:

[0026] Obtain the second data mean of the periodic time series data;

[0027] For any extreme value in the periodic time series data, calculate the difference between the extreme value and the mean of the second data, and obtain the square of the product between the difference and the noise coefficient of the extreme value.

[0028] Based on the squared result of each extreme value in the periodic time series data, the mean of the squared result is calculated, and the result of taking the square root of the mean of the average result is used as the time series fluctuation index of the electricity load in the corresponding period of the periodic time series data.

[0029] Furthermore, obtaining the similarity between the periodic time series data and each other periodic time series data includes:

[0030] For any other periodic time series data, the DTW algorithm is used to perform shortest path matching between data points of the periodic time series data and the other periodic time series data to obtain the corresponding matching results.

[0031] For any set of matching data in the matching results, the distance between the sets of matching data is obtained. The data in the set of matching data that belongs to the periodic time series data is used as the target data. The distance weight of the set of matching data is obtained based on the noise coefficient of the target data.

[0032] Based on the distance weight of each group of matching data in the matching results, the distances of all groups of matching data are summed in a weighted manner to obtain the corresponding weighted sum result. The reciprocal of the weighted sum result is used as the similarity between the periodic time series data and the other periodic time series data.

[0033] Furthermore, the step of obtaining the noise anomaly level of the periodic time series data based on the similarity, the time series fluctuation index, and the normal fluctuation index of electricity load includes:

[0034] Based on the similarity between the periodic time series data and each other periodic time series data, the average similarity is calculated, and the first difference result between the constant 1 and the average similarity is obtained;

[0035] Obtain the second phase difference result between the constant 1 and the normal fluctuation index of the electricity load, and use the product of the first phase difference result, the time series fluctuation index and the second phase difference result as the noise anomaly degree of the periodic time series data.

[0036] Furthermore, the step of training the ARIMA model based on the noise anomaly level of each periodic time series data and the historical electricity load time series data to obtain a trained ARIMA model includes:

[0037] Based on the noise anomaly level of each periodic time series data, the weight coefficient of each data in the historical electricity load time series data is obtained. Based on the historical electricity load time series data and the weight coefficient of each data in the historical electricity load time series data, the ARIMA model is trained to obtain the trained ARIMA model.

[0038] Furthermore, the step of training the ARIMA model based on the historical electricity load time-series data and the weight coefficients of each data point in the historical electricity load time-series data to obtain a trained ARIMA model includes:

[0039] For any data point in the historical electricity load time series data, the data is input into the ARIMA model to obtain the predicted value of the data, the difference between the data and the predicted value is calculated, and the square of the product between the difference and the weight coefficient of the data is obtained.

[0040] The mean of the squares of the products corresponding to each data point in the historical electricity load time series data is calculated, and the square root of the mean of the squares of the products is taken as the root mean square error of the historical electricity load time series data.

[0041] The model parameters of the ARIMA model are corrected in reverse using the gradient descent method until the root mean square error converges, thus obtaining the trained ARIMA model.

[0042] Furthermore, the step of obtaining the weighting coefficient of each data point in the historical electricity load time series data based on the noise anomaly level of each of the said periodic time series data includes:

[0043] For any data in the historical electricity load time series data, the period time series data to which the data belongs is identified as the target period time series data, and the difference between the constant 1 and the noise anomaly degree of the target period time series data is used as the weighting coefficient of the data.

[0044] The embodiments of the present invention have at least the following beneficial effects:

[0045] This invention acquires historical electricity load time-series data for a region targeted by any distributed power source within a preset time period. It then analyzes the electricity load fluctuation status, noise coefficient of each data point within the preset time period, and time-series fluctuation differences between different cycles within the preset time period. This analysis yields the noise anomaly level for each data point in the historical electricity load time-series data, which is used to constrain the noise anomaly. Based on this noise anomaly level, the error weights for each data point in the historical electricity load time-series data are optimized when training the ARIMA model. This reduces the interference of noisy data on the ARIMA model training, improving the training results. The trained ARIMA model then has stronger predictive power and more accurate predictions, ultimately making the predicted short-term power output of distributed power sources more consistent with reality. Attached Figure Description

[0046] To more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0047] Figure 1 The flowchart illustrates the steps of a method for predicting the short-term output power of a distributed power source, as provided in an embodiment of the present invention. Detailed Implementation

[0048] To further illustrate the technical means and effects adopted by the present invention to achieve its intended purpose, the following, in conjunction with the accompanying drawings and preferred embodiments, details the specific implementation, structure, features, and effects of a distributed power source short-term output power prediction method proposed according to the present invention. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. Furthermore, specific features, structures, or characteristics in one or more embodiments can be combined in any suitable form.

[0049] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.

[0050] The following description, in conjunction with the accompanying drawings, details a specific scheme for a distributed power source short-term output power prediction method provided by the present invention.

[0051] Please see Figure 1 The diagram illustrates a flowchart of a short-term output power prediction method for distributed power sources according to an embodiment of the present invention. The method includes the following steps:

[0052] Step S101: Obtain historical electricity load time series data for the area targeted by any distributed power source within a preset time period.

[0053] In this embodiment of the invention, historical electricity load time-series data refers to time-series data composed of electricity consumption at each sampling moment within a historical time period. Since each distributed power source supplies power to a specific area, meaning one distributed power source corresponds to a fixed power supply area, taking one distributed power source as an example, the historical electricity load time-series data for the area served by that distributed power source within a preset time period is obtained. Because external factors such as weather and temperature have a significant impact on electricity consumption, this embodiment of the invention sets one month as the preset time period, meaning that the historical electricity load time-series data for the area served by the distributed power source within one month is obtained.

[0054] Step S102: Based on the local data differences of each data in the historical electricity load time series data, obtain the normal fluctuation index of electricity load in the area targeted by the distributed power source within a preset time period.

[0055] In this embodiment of the invention, considering that distributed power sources typically target a relatively small area, such as a factory, industrial park, residential community, or school, these areas usually exhibit certain electricity load trends or cyclical characteristics. For example, peak electricity consumption periods correspond to commuting times in industrial parks, and residents in residential communities cook, turn on air conditioners, and perform other specific activities at specific times. These behaviors generally exhibit certain temporal cyclical or trend characteristics. Therefore, for the distributed power source to be analyzed, it is necessary to adjust the predictive value of some data in the historical electricity load time-series data when participating in the prediction model construction, based on the historical electricity consumption habits information of the distributed power source. For example, data noise and errors caused by sensor problems often closely resemble the abnormal charge data fluctuations caused by sudden situations (excessive load over a prolonged period) in the area targeted by the distributed power source. However, the latter has reference value when constructing the prediction model, while data interference caused by noise does not have predictive value. Therefore, by analyzing the historical electricity load time-series data, the weights in constructing the prediction model can be optimized, thereby improving the predictive ability of the prediction model.

[0056] Considering that each distributed power source targets a different type of region, and that even within the same type of region, historical electricity consumption characteristics can vary due to the unique characteristics of each region, for example, some factories have large equipment running 24 hours a day and additional electricity consumption from other types of work during the day, while other factories only have large equipment running 24 hours a day, resulting in relatively stable electricity loads. These differences in electricity consumption are reflected differently in historical electricity load time-series data, and the corresponding data fluctuations are also different. Therefore, it is necessary to perform data fluctuation analysis on the historical electricity load time-series data of the region targeted by the distributed power source to understand the electricity fluctuation characteristics of that region. Thus, this embodiment of the invention obtains the normal index of electricity load fluctuation in the region targeted by the distributed power source within a preset time period based on the local data differences of each data point in the historical electricity load time-series data.

[0057] Preferably, based on the local data differences among various data points in the historical electricity load time-series data, the normality index of electricity load fluctuation in the region targeted by the distributed power source within the preset time period is obtained, including:

[0058] Obtain the first average value of the historical electricity load time-series data;

[0059] For any data in the historical electricity load time series data, take the data as the center, obtain all data contained in a preset size window, obtain the data difference between the maximum data and the minimum data based on all data contained in the preset size window, and count the number of extreme data in the preset size window;

[0060] Calculate the absolute value of the difference between the first data mean and the data, normalize the data difference to obtain the corresponding normalization result, perform negative mapping on the number of extreme data to obtain the corresponding mapping result, and obtain the first difference between the first preset value and the mapping result.

[0061] The product of the absolute value of the difference, the normalization result, and the first difference is obtained as the fluctuation characteristic value of the data.

[0062] Obtain the fluctuation characteristic value of each data point in the historical electricity load time series data, calculate the mean of all fluctuation characteristic values, and use the result after normalizing the mean as the normal fluctuation index of electricity load in the region targeted by the distributed power source during the preset time period.

[0063] It should be noted that the method for obtaining extreme value data in the preset size window is as follows: a data change curve is constructed based on all the data contained in the preset size window, and extreme value data is obtained based on the data change curve. The method for obtaining extreme points in the curve is existing technology and will not be elaborated upon here.

[0064] In one embodiment, a one-dimensional window of length L is constructed centered on any data point in the historical electricity load time-series data. The length L of the window can be determined independently; in this embodiment, an empirical value of a 15-minute time-series span is used. It is worth noting that if the left and right ends of some data points cannot be equidistant to construct a one-dimensional window of length L, then data is supplemented at the other end according to the corresponding positions of the data points contained within the window. For example, if the window for the nth data point only contains data 1, 2, and 3 on the left, where data 1 belongs to the (n-1)th data point, data 2 belongs to the (n-2)th data point, and data 3 belongs to the (n-3)th data point, then data is supplemented at the right end of the nth data point, data 1 is supplemented for the (n+1)th data point, data 2 for the (n+2)th data point, and data 3 for the (n+3)th data point, thus forming a one-dimensional window of length L for the nth data point. Based on the local data contained within the window corresponding to each data point, the normal electricity load fluctuation index for the region targeted by the distributed power source within a preset time period is obtained. The calculation expression for the normal electricity load fluctuation index is:

[0065]

[0066] Where c represents the normal fluctuation index of electricity load in the area targeted by any distributed power source within a preset time period, and a n This represents the nth data point in the historical electricity load time-series data. This represents the average of historical electricity load time-series data, also known as the first data mean. N represents the total number of data points included in the historical electricity load time-series data. softmax() represents the normalization exponential function. n V represents the data difference between the maximum and minimum values ​​within the window of the nth data point in historical electricity load time-series data. n This represents the total number of extreme value data points contained in the window containing the nth data point in the historical electricity load time-series data. e represents the natural constant, 1 represents the first preset value, and norm() represents the normalization function. This represents the first difference between the first preset value and the mapping result obtained by negatively mapping the number of extreme value data.

[0067] It should be noted that the average value of historical electricity load time series data is used. The normal baseline for electricity load in the region is defined as the nth data point compared to the average of the data points. Data difference between The larger the value, the more likely the nth data point has fluctuations in electricity consumption. This corresponds to a larger fluctuation characteristic value for the nth data point, and also indicates a greater data difference b between the maximum and minimum values ​​within the local window of the nth data point. n The larger the value, the greater the data fluctuation within the local window of the nth data point, and the larger the corresponding fluctuation characteristic value of the nth data point. Secondly, the total number of extreme value data points V within the local window of the nth data point... n The more frequent the extreme values ​​appear within the local window of the nth data point, the stronger the noise fluctuation characteristics exhibited by all data within the local window of the nth data point. Therefore, the fluctuation characteristic value of the nth data point... The larger the value, the more obvious the local data fluctuation of the nth data. Therefore, the larger the fluctuation characteristic value of each data in the historical electricity load time series data, the larger the normal index of electricity load fluctuation in the area targeted by the distributed power source within the preset time period, indicating that fluctuations are more likely to occur under the time series corresponding to the historical electricity load time series data.

[0068] Step S103: Divide the preset time period into at least one cycle using a preset time interval, take the data within any cycle from the historical electricity load time series data as the cycle time series data, and for any cycle time series data, obtain the noise coefficient of each data based on the local data of each data in the cycle time series data, and obtain the time series fluctuation index of the electricity load in the corresponding cycle based on the noise coefficient of each data in the cycle time series data.

[0069] In this embodiment of the invention, the preset time period is one month, and one day is one cycle. The historical electricity load time series data is divided into multiple cycle time series data. Therefore, for the historical electricity load time series data, the electricity load data from 0:00 to 23:59 on the same day constitutes a cycle time series data.

[0070] Taking a periodic time series data set as an example, based on the local data of each data point in the periodic time series data set, the noise coefficient of each data point is obtained, including:

[0071] Obtain the extreme value data from the periodic time series data;

[0072] For any data in the periodic time series data, with the data as the center, obtain the extreme value data contained in a preset size window and count the first number of extreme value data. Perform negative mapping on the first number to obtain the corresponding mapping value, and obtain the difference between the target value and the mapping value.

[0073] Based on the extreme value data contained within the preset size window, the time span between each pair of adjacent extreme value data is obtained, and the standard deviation between all time spans is calculated.

[0074] Obtain the product of the difference and the standard deviation, and use the normalized result of the product as the noise coefficient of the data.

[0075] In one embodiment, taking the periodic time series data of day w as an example, firstly, based on the criterion that the first derivative is 0, the extreme value data in the periodic time series data of day w is obtained. For the m-th data in the periodic time series data of day w, a one-dimensional window of length L is set with the m-th data as the center of the window, which is the same as the window size in step S102. Then, based on the difference between the extreme value data in the window of the m-th data, the noise coefficient of the m-th data is obtained. The calculation expression of the noise coefficient is:

[0076]

[0077] Where, k wm V represents the noise coefficient of the m-th data point in the periodic time series data of day w, softmax() represents the normalization exponential function, and V wm ε represents the first number of extreme value data points contained in the window containing the m-th data point in the periodic time series data of day w. wm Let e ​​represent the standard deviation of the time span between all adjacent extreme values ​​in the window of the m-th data in the periodic time series data of day w, where e represents the natural constant and 1 represents the target value.

[0078] It should be noted that V is the number of extreme value data points within the window range of the m-th data point in the periodic time series data of day w. wm As the noise level within the corresponding window, compared to fluctuations caused by normal load anomalies, data fluctuations caused by noise interference are relatively more frequent, meaning data fluctuations are more frequent and the number of extreme data points is relatively greater; secondly, ε wm This represents the standard deviation of the time span between all extreme values ​​within the window and their adjacent extreme values. This feature is related to the number of extreme values, V. wm The data fluctuations caused by load anomalies are mutually constrained. Compared to noise, the fluctuations are more disordered, meaning the duration of each fluctuation is different, and the time spans of the corresponding extreme data differ more significantly. Noise, on the other hand, usually manifests as a short and continuous fluctuation. Therefore, within a local range, noise fluctuations are relatively more intense, and the duration of each fluctuation is very similar. Based on these two characteristics, the noise coefficient k of the m-th data point in the periodic time series data of day w is obtained. wm The higher the coefficient, the stronger its weight in the subsequent analysis of the time-series fluctuation of electricity load on day w. In other words, the time-series fluctuation index of electricity load on day w is mainly affected by data with higher noise coefficients.

[0079] After obtaining the noise coefficient of each data point in the periodic time-series data, the time-series fluctuation index of the electricity load in the corresponding period is obtained based on the noise coefficient of each data point in the periodic time-series data. This index is used to characterize the noise anomaly of the periodic time-series data in that period. The method for obtaining the time-series fluctuation index of the electricity load in any period includes:

[0080] Obtain the second data mean of the periodic time series data;

[0081] For any extreme value in the periodic time series data, calculate the difference between the extreme value and the mean of the second data, and obtain the square of the product between the difference and the noise coefficient of the extreme value.

[0082] Based on the squared result of each extreme value in the periodic time series data, the mean of the squared result is calculated, and the result of taking the square root of the mean of the average result is used as the time series fluctuation index of the electricity load in the corresponding period of the periodic time series data.

[0083] In one embodiment, the calculation expression for the time-series fluctuation index of electricity load on day w is:

[0084]

[0085] Among them, F wThis represents the time-series fluctuation index of electricity load on day w, where M represents the total number of extreme values ​​included in the time-series data of day w, and a wi This represents the i-th extreme value in the periodic time series data of day w. Let k represent the mean of the periodic time series data on day w, which is also the second data mean. wi This represents the noise coefficient of the i-th extreme value in the periodic time series data of day w.

[0086] It should be noted that the difference between the i-th extreme value in the periodic time series data on day w and the mean of the periodic time series data on day w is calculated. The load fluctuation change on day w is used as the noise coefficient k of the i-th extreme value data. wi As a weight, the high-noise extreme values ​​have a greater impact on the time-series fluctuation analysis of electricity load. Therefore, the time-series fluctuation index F of electricity load on day w is... w and difference There is a positive correlation, and the time-series fluctuation index F of electricity load on day w is... w The larger the value, the more severe the noise interference is on the electricity load data for day w.

[0087] Step S104: Obtain the similarity between the periodic time series data and each other periodic time series data respectively. Based on the similarity, time series fluctuation index and electricity load fluctuation normal index, obtain the noise anomaly degree of the periodic time series data.

[0088] In this embodiment of the invention, according to step S103, the time series fluctuation index F is known. w This represents the fluctuation characteristics of electricity load on day w. Without considering the similarity of daily electricity load curves over long time periods, the time-series fluctuation index F can be simply defined as... w As a possible source of noise interference in the electricity load data for day w, it is necessary to analyze the dispersion of the trend similarity between the electricity load on day w and the other days, in order to optimize the noise anomaly analysis on day w and make the noise anomaly degree of the subsequent periodic time series data on day w more rigorous.

[0089] First, the similarity between any given period of time series data and each other period of time series data is obtained. The methods for obtaining this similarity include:

[0090] For any other periodic time series data, the DTW algorithm is used to perform shortest path matching between data points of the periodic time series data and the other periodic time series data to obtain the corresponding matching results.

[0091] For any set of matching data in the matching results, the distance between the sets of matching data is obtained. The data in the set of matching data that belongs to the periodic time series data is used as the target data. The distance weight of the set of matching data is obtained based on the noise coefficient of the target data.

[0092] Based on the distance weight of each group of matching data in the matching results, the distances of all groups of matching data are summed in a weighted manner to obtain the corresponding weighted sum result. The reciprocal of the weighted sum result is used as the similarity between the periodic time series data and the other periodic time series data.

[0093] In one embodiment, taking the periodic time series data of day w as an example, the periodic time series data of any day other than day w is used as other periodic time series data. The electricity load data curve constructed from the periodic time series data of each day is obtained. Then, the DTW algorithm is used to perform shortest path matching between the electricity load data curve constructed from the periodic time series data of day w and the electricity load data curve constructed from the periodic time series data of day j. This allows for pairwise matching of data in the periodic time series data of day w and day j. Based on the data distance between any matched data set and the noise coefficient of the data belonging to day w in that data set, the similarity between the periodic time series data of day w and day j is obtained. The formula for calculating the similarity is:

[0094]

[0095] Among them, D wj d represents the similarity between the periodic time series data of day w and the periodic time series data of day j. wmj k represents the data distance between the m-th data point in the periodic time series data of day w and the matching data point in the periodic time series data of day j. wm B represents the noise coefficient of the m-th data point in the periodic time series data of day w. w This represents the total number of data points contained in the periodic time series data for day w.

[0096] It should be noted that a higher noise coefficient for the m-th data point in the periodic time series data of day w indicates more severe noise interference. This suggests that the similarity analysis based on the m-th data point is unreliable. If the anomaly of the m-th data point leads to a large distance between matched data points, the calculated similarity will be lower than the actual similarity. Therefore, it is necessary to weaken the data distance provided by the m-th data point. This can be achieved by reducing the distance weight of the m-th data point, resulting in a higher noise coefficient and a corresponding distance weight (1-k). wmThe smaller the value, the better. The use of the DTW algorithm to pairwise match the periodic time series data of day w with the periodic time series data of day j is an existing technique and will not be elaborated upon here.

[0097] Thus, we can obtain the similarity between the periodic time series data of day w and the periodic time series data of each of the other days.

[0098] After obtaining the similarity between any given period of time series data and each other period of time series data, the noise anomaly level of the given period of time series data is obtained based on the similarity between the given period of time series data and each other period of time series data, the time series fluctuation index of the given period of time series data, and the normal fluctuation index of the electricity load of the given period of time series data. The method for obtaining the noise anomaly level of any given period of time series data includes:

[0099] Based on the similarity between the periodic time series data and each other periodic time series data, the average similarity is calculated, and the first difference result between the constant 1 and the average similarity is obtained;

[0100] Obtain the second phase difference result between the constant 1 and the normal fluctuation index of the electricity load, and use the product of the first phase difference result, the time series fluctuation index and the second phase difference result as the noise anomaly degree of the periodic time series data.

[0101] In one embodiment, taking the periodic time series data of day w as an example, the noise anomaly level of the periodic time series data of day w is obtained based on the similarity between the periodic time series data of day w and the periodic time series data of each other day, the time series fluctuation index of the periodic time series data of day w, and the normal fluctuation index of the electricity load of the periodic time series data of day w. The calculation formula for the noise anomaly level is as follows:

[0102]

[0103] Among them, S w This indicates the degree of noise anomaly in the periodic time series data on day w, where W-1 represents the total number of days within the preset time period excluding day w, and D... wj F represents the similarity between periodic time series data on day w and periodic time series data on day j. w Let c represent the time-series fluctuation index of electricity load on day w, and let c represent the normal fluctuation index of electricity load within a preset time period. (1-c) represents the first phase difference result, and (1-c) represents the second phase difference result.

[0104] It should be noted that, F represents the average similarity of the periodic time series data calculated for day w and the remaining w-1 days. When the similarity between day w and the remaining w-1 days is generally relatively low, it is considered that the electricity load data for day w does indeed show a large difference in the trend of the electricity consumption curve at the same time period due to noise, corresponding to a high degree of noise anomaly; w The noise anomaly of the periodic time series data on day w is characterized, which is the degree of data anomaly quantified based on the fluctuation characteristics of the periodic time series data on day w. This feature and the feature obtained by the similarity of the periodic time series data are mutually adjusted, that is, both features need to be large to jointly represent the noise of the electricity load data on day w; (1-c) characterizes the normal fluctuation index of the electricity load in the area targeted by any distributed power source within a preset time period. As a constraint term for the area targeted by any distributed power source, the larger the value of c, the more likely the corresponding area is to have an abnormal electricity load, the lower the possibility of noise anomaly fluctuation, and the smaller the corresponding degree of noise anomaly. Therefore, the degree of noise anomaly S of the periodic time series data on day w w The larger the value, the more likely the anomalies in the electricity load data on day w over a long time series in that region are due to noise interference.

[0105] Thus, the noise anomaly level S of the periodic time series data on day w within the preset time period can be obtained. w Similarly, according to the methods in steps S103 to S104, the noise level of the periodic time series data for each day within the preset time period is obtained, which is the noise level of the periodic time series data for each period within the preset time period.

[0106] Step S105: Based on the noise anomaly level of the time series data for each cycle and the historical electricity load time series data, train the ARIMA model to obtain the trained ARIMA model, and use the trained ARIMA model to predict the short-term output power of distributed power sources.

[0107] In this embodiment of the invention, the ARIMA model stands for Autoregressive Integrated Moving Average Model. The ARIMA model mainly consists of three parts: an autoregressive model (AR), a differencing process (I), and a moving average model (MA). The basic idea of ​​the ARIMA model is to use historical information from the data itself to predict the future. Therefore, the ARIMA model is trained based on the noise anomaly level of each period's time-series data and historical electricity load time-series data to obtain a trained ARIMA model.

[0108] Preferably, the ARIMA model is trained based on the noise anomaly level of each periodic time series data and the historical electricity load time series data to obtain a trained ARIMA model, including:

[0109] (1) Based on the noise anomaly level of each periodic time series data, obtain the weighting coefficient of each data in the historical electricity load time series data.

[0110] In this embodiment of the invention, since the noise anomaly level of each period time series data obtained in step S104 is obtained by analyzing data within any period of historical electricity load time series data, each data contained in any period time series data in historical electricity load time series data corresponds to the noise anomaly level of the period time series data. Therefore, for any data in the historical electricity load time series data, the period time series data to which the data belongs is identified as the target period time series data, and the difference between the constant 1 and the noise anomaly level of the target period time series data is used as the weighting coefficient of the data.

[0111] It should be noted that the greater the noise anomaly of any data, the smaller the corresponding weight coefficient, and the weaker its impact on ARIMA model training.

[0112] (2) The ARIMA model is trained based on the historical electricity load time series data and the weight coefficient of each data in the historical electricity load time series data to obtain the trained ARIMA model.

[0113] In this embodiment of the invention, for any data in the historical electricity load time series data, the data is input into the ARIMA model to obtain the predicted value of the data, the difference between the data and the predicted value is calculated, and the square of the product between the difference and the weight coefficient of the data is obtained.

[0114] The mean of the squares of the products corresponding to each data point in the historical electricity load time series data is calculated, and the square root of the mean of the squares of the products is taken as the root mean square error of the historical electricity load time series data.

[0115] The model parameters of the ARIMA model are corrected in reverse using the gradient descent method until the root mean square error converges, thus obtaining the trained ARIMA model.

[0116] In one implementation, the expression for calculating the loss function for obtaining the root mean square error when training the ARIMA model is as follows:

[0117]

[0118] Where RMSE represents the root mean square error, N represents the total number of data points included in the historical electricity load time-series data, and d y S represents the difference between the y-th data point in the historical electricity load time series data and its corresponding predicted value. y This represents the noise anomaly level of the y-th data point in the historical electricity load time series data, (1-S y ) represents the weighting coefficient of the y-th data in the historical electricity load time series data.

[0119] It should be noted that when training the ARIMA model, the weight coefficients for model training are obtained based on the degree of noise anomaly of each data point in the historical electricity load time series data. These weight coefficients are then used to optimize the prediction error of each data point in the existing root mean square error loss function. This results in data with higher noise anomalies having a relatively smaller weight in the model error adjustment process, allowing the ARIMA model training to focus on more normal and reliable data, thereby improving the training effect of the ARIMA model. The training process of the ARIMA model is existing technology and will not be described in detail here.

[0120] After the ARIMA model is trained, a well-trained ARIMA model is obtained. At this point, the ARIMA model has strong predictive capabilities, and the prediction results are more consistent with reality when faced with new data. Therefore, the well-trained ARIMA model can be used to predict the short-term output power of distributed power sources. For example, real-time electricity load time-series data for the two hours prior to the current moment can be obtained, and the real-time electricity load time-series data can be input into the well-trained ARIMA model. The output result is the short-term output power of the distributed power source.

[0121] It should be noted that the embodiments of the present invention take the region targeted by a distributed power source as an example to predict the short-term output power of the distributed power source. Therefore, the prediction method for the short-term output power of any distributed power source can refer to the prediction method in the embodiments of the present invention.

[0122] In summary, this embodiment of the invention obtains historical electricity load time-series data for the region targeted by any distributed power source within a preset time period; based on the local data differences of each data point in the historical electricity load time-series data, it obtains the normal fluctuation index of the electricity load in the region targeted by the distributed power source within the preset time period; it divides the preset time period into at least one cycle using a preset time interval, takes data from any cycle in the historical electricity load time-series data as the cycle time-series data, and for any cycle time-series data, it obtains the noise coefficient of each data point based on the local data of each data point in the cycle time-series data, and obtains the time-series fluctuation index of the electricity load in the corresponding cycle based on the noise coefficient of each data point in the cycle time-series data; it obtains the similarity between the cycle time-series data and each other cycle time-series data, and obtains the noise anomaly degree of the cycle time-series data based on the similarity, the time-series fluctuation index, and the normal fluctuation index of the electricity load; it trains the ARIMA model based on the noise anomaly degree of each cycle time-series data and the historical electricity load time-series data, obtains the trained ARIMA model, and uses the trained ARIMA model to predict the short-term output power of the distributed power source. Specifically, the error weights of each data point in the historical electricity load time series data are optimized based on the degree of noise anomaly during ARIMA model training. This reduces the interference of noisy data on ARIMA model training, improves the training results of ARIMA model, and makes the trained ARIMA model more predictive and accurate. Consequently, the prediction results of the short-term output power of distributed power sources using the trained ARIMA model are more consistent with reality.

[0123] It should be noted that the order of the above embodiments of the present invention is merely for descriptive purposes and does not represent the superiority or inferiority of the embodiments. Furthermore, the above description focuses on specific embodiments of this specification. Additionally, the processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired results. In some embodiments, multitasking and parallel processing are possible or may be advantageous.

[0124] The various embodiments in this specification are described in a progressive manner. The same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on describing the differences from other embodiments.

[0125] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for predicting the short-term output power of distributed power sources, characterized in that, The method includes: Obtain historical electricity load time-series data for any distributed power source's target area within a preset time period; Based on the local data differences of each data in the historical electricity load time series data, the normal index of electricity load fluctuation in the region targeted by the distributed power source during the preset time period is obtained. The preset time period is divided into at least one cycle using a preset time interval. Data within any cycle is taken from the historical electricity load time series data as the cycle time series data. For any cycle time series data, the noise coefficient of each data is obtained based on the local data of each data in the cycle time series data. Based on the noise coefficient of each data in the cycle time series data, the time series fluctuation index of the electricity load in the corresponding cycle is obtained. The similarity between the periodic time series data and each other periodic time series data is obtained. Based on the similarity, the time series fluctuation index and the normal fluctuation index of electricity load, the noise anomaly degree of the periodic time series data is obtained. The ARIMA model is trained based on the noise anomaly level of each periodic time series data and the historical electricity load time series data to obtain a trained ARIMA model. The trained ARIMA model is then used to predict the short-term output power of the distributed power source. The step of obtaining the normal fluctuation index of electricity load in the region targeted by the distributed power source within the preset time period based on the local data differences of each data point in the historical electricity load time series data includes: Obtain the first average value of the historical electricity load time-series data; For any data in the historical electricity load time series data, take the data as the center, obtain all data contained in a preset size window, obtain the data difference between the maximum data and the minimum data based on all data contained in the preset size window, and count the number of extreme data in the preset size window; Calculate the absolute value of the difference between the first data mean and the data, normalize the data difference to obtain the corresponding normalization result, perform negative mapping on the number of extreme data to obtain the corresponding mapping result, and obtain the first difference between the first preset value and the mapping result. The product of the absolute value of the difference, the normalization result, and the first difference is obtained as the fluctuation characteristic value of the data. Obtain the fluctuation characteristic value of each data point in the historical electricity load time series data, calculate the mean of all fluctuation characteristic values, and use the result after normalization of the mean as the normal fluctuation index of electricity load in the region targeted by the distributed power source during the preset time period. The step of obtaining the noise coefficient of each data point based on the local data of each data point in the periodic time series data includes: Obtain the extreme value data from the periodic time series data; For any data in the periodic time series data, with the data as the center, obtain the extreme value data contained in a preset size window and count the first number of extreme value data. Perform negative mapping on the first number to obtain the corresponding mapping value, and obtain the difference between the target value and the mapping value. Based on the extreme value data contained within the preset size window, the time span between each pair of adjacent extreme value data is obtained, and the standard deviation between all time spans is calculated. Obtain the product of the difference and the standard deviation, and use the normalized result of the product as the noise coefficient of the data.

2. The method for predicting short-term output power of distributed power sources as described in claim 1, characterized in that, The step of obtaining the time-series fluctuation index of electricity load within a corresponding period based on the noise coefficient of each data point in the periodic time-series data includes: Obtain the second data mean of the periodic time series data; For any extreme value in the periodic time series data, calculate the difference between the extreme value and the mean of the second data, and obtain the square of the product between the difference and the noise coefficient of the extreme value. Based on the squared result of each extreme value in the periodic time series data, the mean of the squared result is calculated, and the result of taking the square root of the mean of the squared result is used as the time series fluctuation index of the electricity load in the corresponding period of the periodic time series data.

3. The method for predicting short-term output power of distributed power sources as described in claim 1, characterized in that, The step of obtaining the similarity between the periodic time series data and each other periodic time series data includes: For any other periodic time series data, the DTW algorithm is used to perform shortest path matching between data points of the periodic time series data and the other periodic time series data to obtain the corresponding matching results. For any set of matching data in the matching results, the distance between the sets of matching data is obtained. The data in the set of matching data that belongs to the periodic time series data is used as the target data. The distance weight of the set of matching data is obtained based on the noise coefficient of the target data. Based on the distance weight of each group of matching data in the matching results, the distances of all groups of matching data are summed in a weighted manner to obtain the corresponding weighted sum result. The reciprocal of the weighted sum result is used as the similarity between the periodic time series data and the other periodic time series data.

4. The method for predicting short-term output power of distributed power sources as described in claim 1, characterized in that, The step of obtaining the noise anomaly level of the periodic time series data based on the similarity, the time series fluctuation index, and the normal fluctuation index of electricity load includes: Based on the similarity between the periodic time series data and each other periodic time series data, the average similarity is calculated, and the first difference result between the constant 1 and the average similarity is obtained; Obtain the second phase difference result between the constant 1 and the normal fluctuation index of the electricity load, and use the product of the first phase difference result, the time series fluctuation index and the second phase difference result as the noise anomaly degree of the periodic time series data.

5. The method for predicting short-term output power of distributed power sources as described in claim 1, characterized in that, The step of training the ARIMA model based on the noise anomaly level of each periodic time series data and the historical electricity load time series data to obtain a trained ARIMA model includes: Based on the noise anomaly level of each periodic time series data, the weight coefficient of each data in the historical electricity load time series data is obtained. Based on the historical electricity load time series data and the weight coefficient of each data in the historical electricity load time series data, the ARIMA model is trained to obtain the trained ARIMA model.

6. The method for predicting short-term output power of distributed power sources as described in claim 5, characterized in that, The step of training the ARIMA model based on the historical electricity load time-series data and the weight coefficient of each data point in the historical electricity load time-series data to obtain a trained ARIMA model includes: For any data point in the historical electricity load time series data, the data is input into the ARIMA model to obtain the predicted value of the data, the difference between the data and the predicted value is calculated, and the square of the product between the difference and the weight coefficient of the data is obtained. The mean of the squares of the products corresponding to each data point in the historical electricity load time series data is calculated, and the square root of the mean of the squares of the products is taken as the root mean square error of the historical electricity load time series data. The model parameters of the ARIMA model are corrected in reverse using the gradient descent method until the root mean square error converges, thus obtaining the trained ARIMA model.

7. The method for predicting short-term output power of distributed power sources as described in claim 5, characterized in that, The step of obtaining the weighting coefficient of each data point in the historical electricity load time series data based on the noise anomaly level of each of the said periodic time series data includes: For any data in the historical electricity load time series data, the period time series data to which the data belongs is identified as the target period time series data, and the difference between the constant 1 and the noise anomaly degree of the target period time series data is used as the weighting coefficient of the data.

Citation Information

Patent Citations

  • Active power distribution network operation situation prediction method based on IEMD-TA-LSTM model

    CN115275991A

  • SOFC (Solid Oxide Fuel Cell) combined heat and power system operation regulation and control method based on short-term load prediction

    CN115965106A