A Distributed Source-Load Resource Migratable Probability Interval Prediction Method Based on Feature Enhancement
By adopting a feature enhancement method in distributed source and load resource prediction, combining gated residual network, two-layer feature enhancement model, timing fusion prediction model and probability interval prediction model, the problem of insufficient limitations and uncertainty information of multi-type, multi-scenario and multi-region prediction tasks in the existing technology is solved, and efficient and accurate power generation and consumption prediction of distributed source and load resource is achieved, and the reliability of the system is improved.
Patent Information
- Application Number
- CN202510437065.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-04-09
AI Technical Summary
The existing distributed source-load resource power generation prediction methods are poor in multiple types, multiple scenarios, and multi-region tasks, and it is difficult to provide sufficient uncertainty information, making it difficult to evaluate the reliability of the prediction results.
The distributed source-load resource migration probability interval prediction method is adopted based on feature enhancement, and efficient and accurate prediction of distributed source-load resource generation and consumption through gated residual network (GRN), two-layer feature enhancement model, timing fusion prediction model and probability interval prediction model are achieved.
This method enhances the model expression ability and feature extraction accuracy, realizes migration prediction of multi-type, multi-scenario, and multi-region distributed source and load resources, significantly improving the system's reliability and scheduling optimization capabilities.
Smart Images

Figure CN119961886B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the prediction technology of distributed energy systems, and particularly to a method for predicting the transferable probability interval of distributed source-load resources based on feature enhancement. Background Art
[0002] Compared with traditional power systems, one of the significant features of the new power system is the transformation from the "source following load" mode to the "source-load interaction" mode. Against this background, in order to ensure the safe and stable operation of the power system, it is necessary to improve the prediction accuracy of power generation and consumption of distributed source-load resources, thereby increasing the utilization rate of load resources to meet the demand for a larger scale of new energy consumption.
[0003] The core issue of predicting power generation and consumption of distributed source-load resources lies in how to use historical data to build a prediction model to accurately predict the changes in distributed resources at future moments or time periods. This ability is crucial for multiple business scenarios such as power grid scheduling, maintenance planning, security and stability analysis, and new energy consumption analysis. Traditional distributed source-load resource prediction technologies mainly rely on physical modeling and mathematical statistics methods. With the development of intelligent measurement devices and the rapid progress of artificial intelligence technology, data-driven artificial intelligence methods have gradually become the mainstream research direction. In particular, load prediction methods based on deep learning, with their advantages in pattern recognition and nonlinear modeling, not only improve the prediction accuracy but also optimize the calculation efficiency. Compared with traditional methods, deep learning-based prediction methods can better handle problems of large-scale data and multi-source data, can achieve real-time prediction and automatic feature extraction, and can also continuously optimize the prediction model.
[0004] In the existing technical research, the data-driven prediction method field can be further divided into two categories: deterministic prediction methods and probability interval prediction methods. Among them, the deterministic point prediction method focuses more on providing specific numerical prediction results without considering their fluctuation range. This type of method is applicable to scenarios where the prediction results with errors can be accepted. In contrast, the probability interval prediction method can not only give numerical prediction results but also provide a range of possibilities, that is, the prediction interval. This method quantifies the uncertainty in the prediction process and is applicable to complex working conditions that need to consider prediction uncertainty. Existing technologies include machine learning models such as long short-term memory neural networks, convolutional neural networks, and Transformers. However, in the prediction of distributed resource power generation and consumption, due to the wide variety of resource types and the influence of multiple uncertainty factors, traditional prediction models are difficult to provide sufficient uncertainty information to comprehensively evaluate the prediction results. In addition, existing models are usually pre-trained for specific resource types and have limitations in generalization ability, and there is still a lack of a general prediction model that can adapt to the transfer requirements of multiple types, multiple scenarios, and multiple regions.
[0005] At present, the method of distributed source-load resource power generation and consumption prediction needs to adjust a large number of hyperparameters for optimization when dealing with prediction tasks of multiple types, multiple scenarios, and multiple regions, and the performance is poor, which is a major technical problem. Currently, the method of distributed resource power generation and consumption prediction usually has good effects after pre-training for specific resource types in specific scenarios. However, as more and more flexible distributed resources participate in the construction of the new power system, it is necessary to design a distributed source-load resource probability interval prediction method that does not require large-scale hyperparameter optimization and has strong transferability for the different operating characteristics of multiple types, multiple scenarios, and multiple regions of distributed resources. In addition, the operation of distributed source-load resources is affected by multiple uncertain factors, which is another important issue in power load prediction. Existing models usually consider these uncertain factors insufficiently, or only predict the power generation and consumption of distributed resources based on steady-state conditions, resulting in their prediction methods being difficult to provide sufficient uncertainty information, making it impossible for users to comprehensively evaluate the reliability of the prediction results. When facing different types of distributed source-load resources, different resources have different influence characteristics. How to design a feature extraction mechanism with both universality and efficiency is another important issue in power generation and consumption prediction. In existing feature extraction methods, most rely on correlation analysis, with limited feature representation ability, unable to effectively isolate the correlation between features, and it is difficult to accurately extract and utilize the influence characteristics of different types of distributed source-load resources.
[0006] It should be noted that the information disclosed in the above background technology section is only used for understanding the background of the present application, and thus may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention
[0007] The main object of the present invention is to overcome the defects existing in the above background technology, and provide a transferable probability interval prediction method for distributed source-load resources based on feature enhancement.
[0008] To achieve the above object, the present invention adopts the following technical solutions:
[0009] A transferable probability interval prediction method for distributed source-load resources based on feature enhancement, comprising the following steps:
[0010] S1. Preprocess the time series data of distributed source-load resource power generation and consumption data and its related characteristic variables;
[0011] S2. Establish a gated residual network (GRN), and by combining the gating mechanism and residual connection, measure the non-linear correlation between the external input and the prediction target, and enhance the expression ability and training stability of the model;
[0012] S3. Based on the output of the gated residual network model, construct a two-layer feature enhancement model. Quantify the correlation between features through high-order partial correlation analysis in statistics, and combine the non-linear feature processing ability of machine learning to dynamically adjust the weights of each feature, realize end-to-end automatic feature extraction, and enhance the data expression ability;
[0013] S4. Based on the output of the two-layer feature enhancement model, construct a time series fusion prediction model. The time series fusion prediction model uses the encoder-decoder of the long short-term memory network (LSTM) to re-encode multi-source data, and introduces a masked self-attention mechanism to capture the long-range dependence relationship between time series features; the time series fusion prediction model also learns the correlation between external features through the gated residual network (GRN) and outputs a prediction result;
[0014] S5. Based on the output of the time series fusion prediction model, construct a probability interval prediction model. Adopt the quantile regression method to simulate the prediction intervals under different confidence levels, quantify the uncertainty of the prediction, and provide more comprehensive uncertainty information for the scheduling of distributed source-load resources.
[0015] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the method for predicting the transferable probability interval of distributed source-load resources based on feature enhancement.
[0016] A computer program product includes a computer program, and when the computer program is executed by a processor, it implements the method for predicting the transferable probability interval of distributed source-load resources based on feature enhancement.
[0017] The present invention has the following beneficial effects:
[0018] The present invention proposes a distributed source and load resource migration probability interval prediction method based on feature enhancement. Aiming at the limitations of the distributed source and load resource prediction model in the prior art in multi-type, multi-scenario, and multi-region tasks, as well as the problem of insufficient consideration of uncertainty factors, the method realizes efficient and accurate prediction of the generation and consumption of distributed source and load resources by fusing the gated residual network (GRN), the two-layer feature enhancement model, the time series fusion prediction model and the probability interval prediction model. The method effectively measures the nonlinear correlation between the external input and the prediction target through GRN to enhance the model expression ability; the two-layer feature enhancement model is combined with the statistical high-order partial correlation analysis and the nonlinear feature processing ability of machine learning to achieve end-to-end automatic feature extraction and improve the accuracy of feature selection; with the help of the LSTM codec and the masked self-attention mechanism in the time series fusion prediction model, the multi-source data is re-encoded and the long-range dependency of the time series features is captured, so as to enhance the adaptability and migration ability of the model to different scenarios; finally, the probability interval prediction model is used to simulate the prediction interval under different confidence levels, quantify the prediction uncertainty, and provide more comprehensive uncertainty information for the scheduling of distributed source and load resources in the new power system, which significantly improves the reliability and scheduling optimization ability of the system.
[0019] Other beneficial effects of the embodiments of the present invention will be further described below. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 The present invention is a flowchart of a method for predicting the probability interval of distributed source and load resource migration based on feature enhancement according to an embodiment of the present invention.
[0021] Figure 2 The present invention is a framework diagram of a distributed source-load resource migration probability interval prediction method based on feature enhancement in an embodiment of the present invention.
[0022] Figure 3 This is a result diagram of a method for predicting the probability interval of distributed source and load resource migration based on feature enhancement in an embodiment of the present invention. DETAILED DESCRIPTION
[0023] The following is a detailed description of the embodiments of the present invention. It should be emphasized that the following description is only exemplary and is not intended to limit the scope and application of the present invention.
[0024] See also Figure 1 The embodiment of the present invention provides a method for predicting the probability interval of distributed source and load resource migration based on feature enhancement, comprising the following steps:
[0025] Step S1, Data preprocessing: Preprocess the time series data of the power generation and consumption data of distributed source-load resources and their related characteristic variables, which may specifically include missing value filling, outlier replacement, and data smoothing to ensure data quality and improve the convergence of the model.
[0026] In a preferred embodiment, the data preprocessing in step S1 includes: filling the missing values in the time series data of the power generation and consumption data of distributed source-load resources and their related characteristic variables, using the data of the previous timestamp to fill the current missing value to maintain the temporal continuity of the data; replacing the outliers outside the reasonable value range with the data of the previous timestamp; smoothing the power generation and consumption data by calculating the average value of the observed values at the current moment and its previous and subsequent timestamps to reduce the impact of data fluctuations on the prediction accuracy.
[0027] Step S2, Establish a gated residual network model: Establish a gated residual network (GRN), by combining the gated mechanism and residual connection, measure the non-linear correlation between the external input and the prediction target, and enhance the expression ability and training stability of the model.
[0028] In a preferred embodiment, the specific process of establishing the gated residual network model (GRN) in step S2 includes: combining the preprocessed data with the context vector, and generating an intermediate layer through an activation function; applying a weight transformation to the intermediate layer to obtain a further processed intermediate layer; inputting this intermediate layer into a gated linear unit (GLU), and combining it with the initial features, and obtaining the output of the GRN through regularization processing; wherein, the gated residual network model measures the non-linear correlation between the external input and the target variable by introducing the gated mechanism and residual connection, and enhances the expression ability of the model; the gated linear unit is used to flexibly suppress the influence of negative values, alleviate the problem of gradient disappearance, and improve the stability of network training.
[0029] Step S3, Establish a two-layer feature enhancement model: Based on the output of the gated residual network model established in step S2, construct a two-layer feature enhancement model, quantify the correlation between features through statistical high-order partial correlation analysis, and combine the non-linear feature processing ability of machine learning to dynamically adjust the weights of each feature, realize end-to-end automatic feature extraction, and enhance the data expression ability;
[0030] In a preferred embodiment, the specific process of establishing the double-layer feature enhancement model in step S3 includes: calculating the high-order partial correlation coefficients between features, quantifying the correlation between features through statistical methods; combining the high-order partial correlation coefficients with the features processed by the gated residual network (GRN), inputting them into the Softmax activation function to obtain the feature weight matrix; in the same time step, inputting each initial feature into the corresponding GRN for non-linear transformation to obtain the feature vector; using the feature weight matrix to weight the feature vector to obtain the feature matrix with correlation weights as the input of the subsequent model. Among them, the double-layer feature enhancement model dynamically adjusts the weight values of each feature by fusing the high-order partial correlation analysis of statistics and the non-linear feature processing ability of machine learning, enhances the model's ability to select key features, and reduces the dependence on large-scale data.
[0031] Step S4, establish a time series fusion prediction model: Based on the output of the double-layer feature enhancement model established in step S3, construct a time series fusion prediction model. The time series fusion prediction model uses the encoder and decoder of the long short-term memory network (LSTM) to re-encode multi-source data and introduces a masked self-attention mechanism to capture the long-range dependence relationship between time series features; the time series fusion prediction model also learns the correlation between external features through the gated residual network (GRN) and outputs the prediction result.
[0032] In a preferred embodiment, the specific process of establishing the time series fusion prediction model in step S4 includes: inputting the feature matrix with correlation weights into the encoder and decoder of the long short-term memory network (LSTM) to re-encode multi-source data to handle the differences in data in different scenarios; learning the re-encoded time series features through the masked self-attention mechanism to capture the long-range dependence relationship between features and prevent the attention mechanism from being interfered by future information during prediction; inputting the extracted time series features into the gated residual network (GRN) module to learn the correlation between each external feature and output the prediction result; among them, the time series fusion prediction model enhances the expression ability of features through the LSTM encoder and decoder and the masked self-attention mechanism, and at the same time uses the multi-head attention mechanism to further capture the complex relationships across time steps to realize the migration prediction of multi-type, multi-scenario, and multi-region distributed source-load resources.
[0033] Step S5, establish a probability interval prediction model: Based on the output of the time series fusion prediction model established in step S4, construct a probability interval prediction model, and use the quantile regression method to simulate the prediction intervals at different confidence levels, quantify the uncertainty of the prediction, and provide more comprehensive uncertainty information for the scheduling of distributed source-load resources.
[0034] In a preferred embodiment, the specific process of establishing the probability interval prediction model in step S5 includes: using the quantile regression method to estimate the conditional quantiles of the target variable at different quantile levels according to the input features; estimating the quantile regression coefficients through a least squares optimization problem to construct the probability distribution function of the target variable; based on the quantile regression results, simulating the prediction intervals at different confidence levels to quantify the prediction uncertainty; wherein, the probability interval prediction model provides the probability distribution of the target variable through quantile regression and represents the prediction uncertainty in the form of prediction intervals, thereby providing more comprehensive decision-making support for the scheduling of distributed source-load resources.
[0035] In a further embodiment, the method further includes the following steps:
[0036] Step S6, establishing an evaluation index system for the prediction results: evaluating the prediction results through deterministic indexes and probability prediction indexes to verify the prediction accuracy, stability and transferability of the model.
[0037] In a preferred embodiment, the specific process of establishing the evaluation index system for the prediction results in step S6 includes: using the symmetric mean absolute percentage error (SMAPE), coefficient of determination ( ), and normalized root mean square error (NRMSE) as deterministic indexes to evaluate the accuracy of the prediction results; introducing the accuracy rate ( ), and passing rate ( ), indexes to further measure the closeness of the prediction results to the true values; using the prediction interval coverage probability (PICP), average interval width (AIW), and true coverage width index (TCWI) as probability prediction indexes to evaluate the coverage ability of the prediction intervals for the true values and the effectiveness of the interval widths. The evaluation index system comprehensively verifies the prediction accuracy, stability and transferability of the model by combining deterministic indexes and probability prediction indexes.
[0038] The proposed method for transferable probability interval prediction of distributed source-load resources based on feature enhancement, wherein the gated residual network (GRN) model can effectively measure the non-linear correlation between external inputs and prediction targets, significantly improving the expression ability of the model; the double-layer feature enhancement model combines statistical and machine learning methods to achieve end-to-end automatic feature extraction while ensuring the independence between features, thereby enhancing the expression ability of the data. In addition, the time series fusion prediction model re-encodes multi-source data through the encoder-decoder of the long short-term memory network (LSTM) and introduces a masked self-attention mechanism to learn deep feature correlations, and then realizes the transfer prediction of distributed source-load resources of multiple types, multiple scenarios and multiple regions. The probability interval prediction model comprehensively considers multiple uncertainty factors, simulates the prediction intervals at different confidence levels, and effectively quantifies the uncertainty information of the prediction.
[0039] Figure 2 Shows the framework of a method for predicting the transferable probability interval of distributed source and load resources based on feature enhancement according to an embodiment of the present invention.
[0040] The double-layer feature enhancement model of the present invention can realize end-to-end automatic feature extraction while isolating the correlation between features, effectively improving the accuracy of feature selection. The re-encoding ability of the LSTM encoder-decoder for multi-source data enables the model to be applicable to distributed source and load resource prediction tasks of multiple types, multiple scenarios, and multiple regions. The masked self-attention mechanism ensures that when predicting the output of the next moment, the model will not be affected by inputs beyond the current moment, thereby enhancing the model's learning ability of the correlation between different time steps. In addition, the probability interval prediction model based on Monte Carlo can simulate the prediction intervals at different confidence levels on the premise of considering multiple uncertainty factors. This provides more comprehensive uncertainty information for the distributed source and load resource scheduling in the new power system, further enhancing the reliability and scheduling optimization ability of the system.
[0041] The following further describes the specific embodiments of the present invention, its algorithm examples, and experimental verification.
[0042] A method for predicting the transferable probability interval of distributed source and load resources based on feature enhancement effectively measures the non-linear correlation between external inputs and prediction targets through a gated residual network model, improving the expression ability of the model; fuses statistical and machine learning methods through a double-layer feature enhancement model, and realizes end-to-end automatic feature extraction on the premise of ensuring the independence of features, enhancing the data expression ability; can re-encode the input data of different scenarios through a time series fusion prediction model, and then realizes the transfer prediction of distributed source and load resources of multiple types, multiple scenarios, and multiple regions, and introduces a masked self-attention mechanism to learn deep feature correlations; can comprehensively consider multiple uncertainty factors through a probability interval prediction model, simulate the prediction intervals at different confidence levels, and effectively quantify the uncertainty information of the prediction. Specifically, the method includes the following steps:
[0043] (1) Preprocess the distributed resource power generation and consumption data and their characteristic variables:
[0044] The characteristic variables include temperature, humidity, wind speed, weekdays, holidays, seasons, etc. However, due to intermittent failures of sensors or statistical errors, there may be missing values and outliers in the original dataset. To ensure data quality, data preprocessing is first performed, including filling in missing values and replacing outliers. The data of the previous timestamp is used to fill in the missing values to maintain the temporal continuity of the data, and a reasonable value range is set for each feature, and the data beyond the acceptable range is replaced with the data of the previous timestamp. In addition, due to the influence of various uncontrollable factors on the internal environment and the time lag of distributed resource loads, the power generation and consumption data may fluctuate greatly, thus affecting the prediction accuracy. Therefore, data smoothing processing is carried out to improve the convergence of the model. The specific data smoothing processing formula is as follows:
[0045]
[0046] where denotes the result of smoothing processing at time denotes the observed value at time denotes the moving window radius. To retain as much original information as possible, the moving window radius is set to 1. The average value is calculated using the data at the current time and the previous and next timestamps (a total of 15 minutes).
[0047] (2) Establish a gated residual network model:
[0048] The exact relationship between the external characteristics of the operation of distributed source-load resources and their power generation and consumption loads is not yet clear, so it is difficult to directly determine which variables are related to them. Therefore, a gated residual network model (Gated Residual Network, GRN) is proposed to measure the correlation between external inputs and target variables. GRN combines the ideas of gated mechanisms and residual connections, and can flexibly capture non-linear correlations and enhance the expressive power of the model.
[0049] The main steps to establish the GRN model are as follows: 1. Combine the preprocessed data with the context vector, add a bias term, and generate an intermediate layer through an activation function ; 2. Apply a weight transformation to the intermediate layer and add a bias term to obtain ; 3. Input into the gated linear unit, combine it with the initial features, and obtain the output of GRN through a regularization layer.
[0050] The input of GRN is the initial feature and an optional context vector :
[0051]
[0052] Among them, is the exponential linear unit activation function, and are the intermediate layers, is the regularization layer, is an index representing weight sharing. When , ELU acts as a feature recognition function, and when , ELU acts as an activation function to produce a constant output, thus exhibiting linear layer behavior.
[0053] When constructing the GRN model, gated linear units (GLU) are also introduced to flexibly suppress the influence of negative values, alleviate the vanishing gradient problem, and improve the stability of network training. The intermediate layer is the input of GLU, and the response formula of its gating mechanism is as follows:
[0054]
[0055] Among them, is the weight matrix, is the bias term, is the sigmoid activation function, is the Hadamard product.
[0056] GLU enables the prediction model to dynamically control the contribution degree of GRN to the initial feature . This layer has an adaptive adjustment ability and can approximately output zero when necessary, thus effectively suppressing the non-linear contribution and enabling GRN to directly skip this layer when no additional transformation is required. In addition, for instances without context vectors, GRN only treats as zero.
[0057] The gated residual network model established in step (2) is used for the establishment of the double-layer feature enhancement model in step (3) and the establishment of the time series fusion prediction model in step (4).
[0058] (3) Establish a double-layer feature enhancement model:
[0059] In the power generation and consumption prediction of distributed source-load resources, the influence of input variables on the prediction results is often unclear, and there may be significant differences in the contribution degrees of different features. To solve this problem, a double-layer feature enhancement model is proposed, and weights are assigned to each variable through feature engineering. In this model, statistics and machine learning are combined, and high-order partial correlation analysis is introduced as prior knowledge to calculate the high-order partial correlation coefficients between variables before each round of machine learning training. It realizes that the model can dynamically adjust the weight values of each feature on the premise of excluding the correlation between features.
[0060] The main steps to establish a double - layer feature enhancement model are as follows: 1. Calculate the high - order partial correlation coefficients of each feature, and input them together with the features processed by GRN into the Softmax activation function to obtain the feature weight matrix in the first stage. ; 2. In the same time step, input each initial feature into its corresponding GRN to obtain the feature vector after non - linear transformation. ; 3. Use the weight matrix to weight the feature vector to obtain the feature matrix with correlation weights .
[0061] The partial correlation coefficient between any two of the three variables is calculated after excluding the influence of the remaining one variable, and is called the first - order partial correlation coefficient. The calculation formula is as follows:
[0062]
[0063] where, is the simple correlation coefficient between variable and , is the simple correlation coefficient between variable and , is the simple correlation coefficient between variable and .
[0064] If there are variables , then for any two variables with order and , the calculation formula for their high - order partial correlation coefficient is:
[0065]
[0066] where, all the terms on the right - hand side of the equal sign are order high - order partial correlation coefficients. is the partial correlation coefficient between variable and variable i after excluding the influence of variable j . is the partial correlation coefficient between variable and variable i after excluding the influence of variable j . is the partial correlation coefficient between variable and variable i after excluding the influence of variable The partial correlation coefficient between is the partial correlation coefficient between variable and j after excluding the influence of variable . The high-order partial correlation coefficient after taking the absolute value indicates no correlation when , weak correlation when , medium correlation when , and strong correlation when .
[0067] Then, the calculated high-order partial correlation coefficient and the features after GRN non-linear transformation are used as the initial inputs, and after activation by the Softmax function, they are normalized:
[0068]
[0069] where is the weight matrix after the first layer of feature selection mechanism.
[0070] At each time step, each original feature is also input into its corresponding GRN for additional non-linear feature capture:
[0071]
[0072] where is the feature vector of the original variable after non-linear capture by GRN processing, and the weights of each are shared among all time steps.
[0073] Finally, according to the weight matrix the feature vector is weighted and combined to obtain the feature matrix with correlation weights:
[0074] .
[0075] The proposed double-layer feature enhancement model combines the advantages of statistics and machine learning. It quantifies the correlation between features at the statistical level through high-order partial correlation analysis, and at the same time uses GRN to flexibly process complex features and effectively resist noise interference. In addition, the model integrates the correlation coefficient into the end-to-end neural network framework, which not only enhances the model's ability to select key features, but also reduces the dependence on large-scale data.
[0076] The double-layer feature enhancement model established in step (3) is used to assign weights to the original features and is used as an input variable in step (4).
[0077] (4)Establish a time series fusion prediction model:
[0078] In the time series prediction task of distributed source-load resources, although there are certain differences in the power generation and consumption data of different scenarios, there are often strong similarities in external influencing factors and pattern characteristics. In view of this phenomenon, a time series fusion prediction model is proposed. The encoder-decoder of the Long Short-Term Memory (LSTM) network is used to automatically process the differences between data. Then, a masked self-attention mechanism is introduced to learn the correlations between features to achieve efficient end-to-end power generation and consumption prediction of distributed source-load resources.
[0079] The main steps to establish the time series fusion prediction model are as follows: 1. Input the feature matrix with correlation weights into the LSTM encoder and decoder respectively to re-encode the multi-source data; 2. Learn the re-encoded time series features through the masked self-attention mechanism to capture the long-range dependencies between features; 3. Input the extracted time series features into the GRN for further learning and output the predicted power load value.
[0080] First, use the LSTM encoder and decoder as the core modules. Input the historical time period into the encoder and input the predicted future time period into the decoder, so as to generate a set of uniform time features as the input of the masked self-attention mechanism. These features are represented by , where is the position index. In addition, to enhance the feature expression ability, gated skip connections are adopted in each layer of the model and are implemented through the following formula:
[0081]
[0082] Among them, represents the feature processed by the gating mechanism, , represents the number of time steps predicted in advance at time t.
[0083] Then, input the re-encoded features into the self-attention mechanism to learn the correlations between features. To improve the "distraction" problem in the traditional self-attention mechanism and more accurately learn the relationships between different time steps, a masked self-attention mechanism is proposed. Generally speaking, the attention mechanism scales the value (V) through the relationship between the key (K) and the query (Q), and its calculation method is as follows:
[0084]
[0085] Among them, is a normalization function. Traditional normalization functions cause the attention mechanism to see the entire input during prediction. To prevent it from being "distracted", a mask matrix M is introduced so that it cannot see the input beyond the next time step when predicting the output of the next time step:
[0086]
[0087] Among them, is the dimension of the key K, which is used for scaling calculations to prevent the vanishing gradient due to too large inner product values. When it means does not pay attention to ; when it means pays attention to .
[0088] To further enhance the model's ability to capture complex relationships across time steps, the self-attention mechanism is extended to the multi-head attention framework. Specifically, the value weight matrix is shared among all attention heads, and additive aggregation is used to combine the attention outputs of each head. The calculation method of the multi-head attention mechanism is as follows:
[0089]
[0090] Among them, is the aggregated attention output, is the output weight matrix, is the shared value weight matrix. It can be seen that the interpretable multi-head attention final output is similar to that of a single-layer attention layer. The key difference lies in the way of calculating the attention weight , which is defined as follows:
[0091]
[0092] Among them, is the number of attention heads, and are the weight matrices of the query and key corresponding to the th attention head. Finally, the calculation formula of the masked self-attention mechanism is obtained:
[0093] .
[0094] After the above feature re-encoding and feature extraction of the self-attention mechanism, the aggregated output is input into the GRN module, and the GRN is used to learn the correlation between external features and output the corresponding prediction results.
[0095] The time-series fusion prediction model established in step (4) is used for the establishment of the probability interval prediction model in step (5).
[0096] (5) Establish a probability interval prediction model:
[0097] Due to the influence of multiple uncertain factors such as weather conditions, user behavior, and holidays on distributed source-load resources, a probability interval prediction model is proposed to provide the probability distribution function of the target variable and simulate the prediction intervals (Prediction Intervals, PI) at different confidence levels, thereby effectively quantifying the uncertainty of the prediction.
[0098] Quantile regression is a method that considers the sample probability distribution. It can provide regression results at specific quantile points, and the uncertainty of the prediction results is usually represented by a bounded closed interval. For a given bounded closed interval real variable, the symbol is used to represent the value range of the variable, and its definition is as follows:
[0099]
[0100] Among them, and represent the lower and upper bounds of the interval variable respectively, represents the set of real numbers.
[0101] Suppose represents the power load demand of distributed resources, represents the input feature. When takes the value of , the quantile regression estimates at the corresponding quantile point . This method can provide the specific quantile estimate of under given conditions and represent the prediction uncertainty through an interval, thus forming the confidence interval of the probability prediction. For example, if the prediction error follows a Gaussian distribution and calculates at the 0.05 quantile and 0.95 quantile of , then the 90% confidence interval of
[0102]
[0103] Among them, represents the conditional quantile of at the quantile when is given. The parameter vector defines the regression coefficients of the specific quantile, represents the quantile level. The parameter vector The estimation problem can be transformed into a least squares optimization problem:
[0104]
[0105] where is the sample number. Embedding this formula as the loss function into the prediction model to achieve regression prediction for a specific quantile and calculate the confidence interval of the load probability prediction of distributed resources.
[0106] (6) Establish an evaluation index system for prediction results:
[0107] (6-1) Deterministic index:
[0108] This method uses the symmetric mean absolute percentage error (SMAPE), coefficient of determination ( ), and normalized root mean square error (NRMSE) to evaluate the accuracy of the prediction results. The definitions and calculations of these evaluation indexes are as follows:
[0109]
[0110] where is the predicted value, is the true value, is the mean of the true values.
[0111] The accuracy rate ( ) and qualification rate ( ) indexes are also introduced as follows:
[0112]
[0113] where is the judgment index, is the maximum output of distributed resources during this time period, is the threshold coefficient.
[0114] (6-2) Probability prediction index:
[0115] This method introduces the prediction interval coverage probability (PICP), average interval width (AIW), and true coverage width index (TCWI) to evaluate the effectiveness of the probability interval. The definitions and calculations of these evaluation indexes are as follows:
[0116]
[0117]
[0118] where is the indicator function, is the true value, is the predicted value, is the number of prediction samples, and represent the upper and lower bounds of the prediction interval, respectively.
[0119] To evaluate the performance, tests were conducted under four different scenarios: photovoltaic power generation, wind power generation, electric vehicle charging stations, and integrated energy systems. The time interval for each dataset is 15 minutes, involving complex external influencing factors such as weather and equipment status. The dataset is divided into an 80% training set and a 20% test set in chronological order. Table 1 shows that when the model after large-scale optimization only for the photovoltaic power generation scenario is migrated and applied to the load forecasting of wind power generation, electric vehicle charging stations, and integrated energy systems, its prediction performance only shows a slight decline, demonstrating good stability and transferability. Under all scenarios, is higher than 85%, is higher than 90%, and PICP(90%) is higher than 90%. These core indicators fully verify the transferability and robustness of the method proposed in the present invention in the power generation and consumption prediction of multi-type, multi-scenario, and multi-region distributed source-load resources.
[0120] Table 1 Evaluation indicators of the method of the present invention migrated to different scenarios
[0121]
[0122] Figure 3 shows the prediction results of the method of the present invention under four different scenarios. 70%PI, 80%PI, and 90%PI represent the prediction intervals PI generated by the prediction model at 70%, 80%, and 90% confidence levels, respectively. It can be seen that without adjusting the hyperparameters, the prediction intervals at different confidence levels can well cover the true values (the proportion of the true values actually falling within the prediction intervals is evaluated by the aforementioned prediction interval coverage probability PICP, and the PICP(90%) data in Table 1 is a quantitative evaluation of the coverage ability of the prediction intervals at the 90% confidence level). This proves that the method proposed in the present invention successfully extracts time and external influence features through its unique components, thereby providing a more reliable, more robust, and highly transferable probabilistic prediction of distributed source-load resources for power generation and consumption.
[0123] The embodiment of the present invention also provides a storage medium for storing a computer program, which when executed, at least executes the method as described above.
[0124] The embodiment of the present invention also provides a control device, including a processor and a storage medium for storing a computer program; wherein, the processor is used to execute the computer program to at least execute the method as described above.
[0125] An embodiment of the present invention further provides a processor, which executes a computer program and at least executes the method described above.
[0126] The storage medium can be implemented by any type of non-volatile storage device, or a combination thereof. Among them, the non-volatile memory can be a read-only memory (ROM, Read Only Memory), a programmable read-only memory (PROM, Programmable Read-Only Memory), an erasable programmable read-only memory (EPROM, Erasable Programmable Read-Only Memory), an electrically erasable programmable read-only memory (EEPROM, Electrically Erasable Programmable Read-Only Memory), a ferromagnetic random access memory (FRAM, Ferromagnetic Random Access Memory), a flash memory (Flash Memory), a magnetic surface memory, an optical disc, or a compact disc read-only memory (CD-ROM, Compact Disc Read-Only Memory); the magnetic surface memory can be a disk memory or a tape memory. The storage medium described in the embodiments of the present invention is intended to include, but is not limited to, these and any other suitable types of memories.
[0127] In several embodiments provided by the present invention, it should be understood that the disclosed system and method can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods, such as: multiple units or components can be combined, or can be integrated into another system, or some features can be ignored, or not executed. In addition, the coupling, direct coupling, or communication connection between the various components shown or discussed with each other can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be electrical, mechanical, or other forms.
[0128] The units described above as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units; some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0129] In addition, in each embodiment of the present invention, each functional unit can be all integrated into one processing unit, or each unit can be separately used as one unit, or two or more units can be integrated into one unit; the above-mentioned integrated unit can be implemented in the form of hardware, or in the form of a hardware plus a software functional unit.
[0130] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps including the above method embodiments; and the foregoing storage medium includes: removable storage devices, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks and other various media that can store program codes.
[0131] Alternatively, if the above-mentioned integrated unit of the present invention is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiments of the present invention, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in various embodiments of the present invention. And the foregoing storage medium includes: removable storage devices, ROM, RAM, magnetic disks, or optical disks and other various media that can store program codes.
[0132] The methods disclosed in several method embodiments provided by the present invention can be arbitrarily combined without conflict to obtain new method embodiments.
[0133] The features disclosed in several product embodiments provided by the present invention can be arbitrarily combined without conflict to obtain new product embodiments.
[0134] The features disclosed in several method or device embodiments provided by the present invention can be arbitrarily combined without conflict to obtain new method embodiments or device embodiments.
[0135] The above content is a further detailed description of the present invention in combination with specific preferred implementation manners. It cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those skilled in the technical field to which the present invention pertains, without departing from the concept of the present invention, several equivalent substitutions or obvious variations can be made, and as long as the performance or use is the same, they should all be regarded as belonging to the protection scope of the present invention.
Claims
1. A distributed source-load resource migration probability interval prediction method based on feature enhancement, characterized in that: The following steps are involved: S1. Preprocessing the time series data of power generation and consumption data of distributed source and load resources and their related characteristic variables; S2. Establish a gated residual network (GRN) that combines the gating mechanism and residual connection to measure the nonlinear correlation between external input and prediction target, thereby enhancing the model's expressiveness and training stability. S3. Based on the output of the gated residual network model, a two-layer feature enhancement model is constructed. The correlation between features is quantified through statistical high-order partial correlation analysis. In combination with the nonlinear feature processing capability of machine learning, the weight of each feature is dynamically adjusted to achieve end-to-end automatic feature extraction and enhance data expression capability. S4. Based on the output of the two-layer feature enhancement model, a time series fusion prediction model is constructed. The time series fusion prediction model uses a long short-term memory network (LSTM) encoder and decoder to re-encode multi-source data, and introduces a masked self-attention mechanism to capture the long-range dependency between time series features; the time series fusion prediction model also learns the correlation between external features through the gated residual network (GRN) and outputs a prediction result; S5. Based on the output of the time series fusion prediction model, a probability interval prediction model is constructed, and the quantile regression method is used to simulate the prediction intervals under different confidence levels, quantify the uncertainty of the prediction, and provide more comprehensive uncertainty information for the scheduling of distributed source and load resources.
2. The method for predicting the probability interval of distributed source and load resource migration based on feature enhancement according to claim 1 is characterized in that: The data preprocessing in step S1 includes: Fill missing values in the time series data of distributed source and load resource power generation and consumption data and related characteristic variables, and use the data of the previous timestamp to fill the current missing values to maintain the temporal continuity of the data; Replace the outliers that are beyond the reasonable value range with the data of the previous timestamp; The power generation and consumption data is smoothed by calculating the average value of the observations at the current time and its previous and subsequent timestamps to reduce the impact of data fluctuations on prediction accuracy.
3. The method for predicting the probability interval of distributed source and load resource migration based on feature enhancement according to claim 1 is characterized in that: The specific process of establishing the gated residual network model (GRN) in step S2 includes: Combine the preprocessed data with the context vector and generate the intermediate layer through the activation function; Apply weight transformation to the intermediate layer to obtain the further processed intermediate layer; The intermediate layer is input into the gated linear unit (GLU) and combined with the initial features to obtain the output of the GRN through regularization. Among them, the gated residual network model measures the nonlinear correlation between external input and target variables by introducing gating mechanism and residual connection, thereby enhancing the expressive power of the model; the gated linear unit is used to flexibly suppress the influence of negative values, alleviate the gradient vanishing problem, and improve the stability of network training.
4. The method for predicting the probability interval of distributed source and load resource migration based on feature enhancement according to claim 1 is characterized in that: The specific process of establishing the double-layer feature enhancement model in step S3 includes: Calculate the high-order partial correlation coefficients between each feature and quantify the correlation between features through statistical methods; Combine the high-order partial correlation coefficient with the features processed by the gated residual network (GRN), input the Softmax activation function, and obtain the feature weight matrix; In the same time step, each initial feature is input into the corresponding GRN for nonlinear transformation to obtain the feature vector; The feature vectors are weighted using the feature weight matrix to obtain a feature matrix with correlation weights as the input of the subsequent model.
5. The method for predicting the probability interval of distributed source and load resource migration based on feature enhancement according to claim 1 is characterized in that: The specific process of establishing the time series fusion prediction model in step S4 includes: The feature matrix with correlation weights is input into the encoder and decoder of the long short-term memory network (LSTM) to re-encode the multi-source data to handle the differences in data from different scenarios; The re-encoded temporal features are learned through a masked self-attention mechanism to capture the long-range dependencies between features and prevent the attention mechanism from being disturbed by future information during prediction. The extracted temporal features are input into the gated residual network (GRN) module to learn the correlation between the external features and output the prediction results.
6. The method for predicting the probability interval of distributed source and load resource migration based on feature enhancement according to claim 1 is characterized in that: The specific process of establishing the probability interval prediction model in step S5 includes: The quantile regression method is used to estimate the conditional quantile of the target variable at different quantile levels based on the input features; The quantile regression coefficients are estimated through the least squares optimization problem and the probability distribution function of the target variable is constructed; Based on the quantile regression results, prediction intervals at different confidence levels are simulated to quantify the uncertainty of the prediction.
7. The method for predicting the probability interval of distributed source and load resource migration based on feature enhancement according to claim 1 is characterized in that: The following steps are also included: S6. Establish a prediction result evaluation index system: Evaluate the prediction results through deterministic indicators and probabilistic prediction indicators to verify the prediction accuracy, stability and transferability of the model.
8. The method for predicting the probability interval of distributed source and load resource migration based on feature enhancement according to claim 7 is characterized in that: The specific process of establishing the prediction result evaluation index system in step S6 includes: The symmetric mean absolute percentage error SMAPE and the coefficient of determination are used. and normalized root mean square error NRMSE as deterministic indicators to evaluate the accuracy of the prediction results; Introducing accuracy and pass rate Indicators, further measure the closeness of the predicted results to the true value; The prediction interval coverage probability PICP, average interval width AIW and true coverage width index TCWI are used as probability prediction indicators to evaluate the coverage ability of the prediction interval for the true value and the effectiveness of the interval width.
9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method for predicting the probability interval of distributed source and load resource migration based on feature enhancement as described in any one of claims 1 to 8 is implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the method for predicting the probability interval of distributed source and load resource migration based on feature enhancement as described in any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Multi-feature multi-modal fusion tight gas reservoir yield deep learning prediction method
CN119129648A
Combined learning method and apparatus using deepening neural network based feature enhancement and modified loss function for speaker recognition robust to noisy environments
US20220208198A1