A two-layer complementary regularized joint prediction method for water, wind and solar energy
By using an adaptive wavelet soft thresholding denoising and a seasonal trend decomposition dual-channel spatiotemporal Transformer model, combined with a loss function of a two-layer complementary regularization term, the problems of noise interference and spatiotemporal coupling feature extraction in the power generation prediction of hydro-wind-solar hybrid systems are solved, achieving high-precision power prediction and improving the scheduling decision-making capability of new power systems.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHINA THREE GORGES UNIV
- Filing Date
- 2026-04-27
- Publication Date
- 2026-07-31
AI Technical Summary
Existing technologies for predicting the power generation of hydro-wind-solar hybrid systems suffer from problems such as large data noise interference, difficulty in extracting complex spatiotemporal coupling features, and lack of physical complementarity constraints in the loss function, which limit the prediction accuracy.
We employ adaptive wavelet soft thresholding denoising, a seasonal trend decomposition dual-channel spatiotemporal Transformer model, and a loss function with dual-layer complementary regularization terms. Noise is removed by the adaptive wavelet soft thresholding denoising algorithm, a seasonal trend decomposition dual-channel spatiotemporal Transformer model is constructed, and a loss function with dual complementary regularization terms at the system layer and topology layer is introduced to achieve high-precision prediction.
It effectively removes high-frequency random noise, accurately captures complex spatiotemporal coupling characteristics, improves the accuracy and stability of power generation prediction for hydro-wind-solar complementary systems, conforms to the real scheduling rules of multi-energy complementary systems, and provides highly reliable decision-making data support for day-ahead power generation planning of new power systems.
Smart Images

Figure CN122495322A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of power system technology, and in particular to power system operation control and new energy power generation prediction technology, as well as deep learning-based spatiotemporal modeling technology. Specifically, it relates to a method and system for joint prediction of power generation of hydro-wind-solar hybrid systems that consider the complementary characteristics of multiple sources. Background Technology
[0002] In existing technologies, constructing a new type of power system based on new energy sources has become crucial for energy transformation. For example, the document titled "Evaluation of New Energy Bundling Capacity of Hydropower-Wind-Solar Complementary System under the Constraint of Hydropower Regulation Capacity" (document number 10.13335 / j.1000-3673.pst.2025.0716) describes a "hydropower-wind-solar complementary system." This system relies on the regulation capacity and transmission channels of hydropower bases to integrate hydropower, wind power, and photovoltaic power into a unified dispatch unit. Through a centralized control center, it coordinates the predicted output and power generation plans of various power sources and internally coordinates the real-time output of each power source to mitigate random fluctuations in wind and solar power and to bundle and transmit hydropower, wind power, and solar power. The essence of this system lies in utilizing the rapid start-up and shutdown, wide regulation range, and natural complementarity between hydropower and wind and solar power on a seasonal scale to improve the stability, reliability, and new energy absorption capacity of multi-energy joint operation. However, wind and solar power are significantly affected by meteorological conditions, exhibiting strong intermittency, volatility, and randomness. Combined with the seasonal variations in hydropower output, the total power generation of a hydro-wind-solar hybrid system displays an extremely complex nonlinear spatiotemporal evolution pattern. Therefore, achieving high-precision joint prediction of power generation for hydro-wind-solar hybrid systems is of great significance for the safe dispatch and optimized operation of the power grid.
[0003] Existing methods for predicting multi-source power generation can be broadly categorized into two types: physics-driven models and data-driven models. Physics models rely on detailed numerical weather prediction (NWP) data and generator physical parameters, resulting in high computational complexity and limitations imposed by the accuracy of parameter acquisition. Data-driven models, especially those based on deep learning such as Long Short-Term Memory (LSTM), Convolutional Neural Networks (CNN), and the standard Transformer, are widely used due to their powerful nonlinear fitting capabilities.
[0004] Despite advancements in power prediction, several key challenges remain for the joint prediction of large-scale hydro-wind-solar hybrid systems: Firstly, multi-source power sequences typically contain significant environmental noise, equipment measurement errors, and high-frequency random disturbances. Secondly, existing data preprocessing methods often employ simple moving averages or globally fixed thresholds for denoising. Globally fixed thresholds struggle to adapt to signal characteristics across different time scales, leading to over-smoothing or under-denoising, which can interfere with subsequent models capturing the true evolutionary trend. Thirdly, hydro-wind-solar hybrid systems exhibit not only long-term seasonal trends and short-term dynamic fluctuations over time, but also complex electrical connections and geographical-climatic relationships between different power plants. Finally, traditional prediction models often treat each energy source or power plant as an independent entity or simply stitch together multi-source data, lacking a deep understanding of the spatiotemporal coupling characteristics. While standard Transformer models excel at handling long sequences, they struggle to simultaneously account for local electrical coupling and global complementarity when dealing with multivariate, multi-scale spatiotemporal dependencies. Most deep learning prediction models optimize by minimizing mean squared error (MSE) or mean absolute error (MAE) during training. MSE focuses solely on numerical approximation, neglecting the inherent physical complementarity between multiple energy sources like water, wind, and solar power. This leads to feature homogenization, hindering the effective learning of coordinated fluctuations and energy substitution relationships among different energy sources. Consequently, prediction accuracy decreases during periods of complex fluctuations, and the models lack physical interpretability.
[0005] In summary, effectively removing noise from multi-source data, accurately capturing complex spatiotemporal coupling characteristics, and guiding the model to learn the inherent energy complementarity mechanism are the technical challenges that urgently need to be solved in the power prediction of hydro-wind-solar complementary systems. Summary of the Invention
[0006] This invention aims to address the problems of high data noise interference, difficulty in extracting complex spatiotemporal coupling features, and limited prediction accuracy due to the lack of physical complementarity constraints in the loss function of existing hydro-wind-solar hybrid system power prediction technologies. It provides a spatiotemporal joint prediction method and system based on dual-layer complementarity regularization. This method achieves high-precision and high-stability prediction of the power generation of multi-source complementary systems by employing adaptive wavelet soft thresholding denoising, dual-channel spatiotemporal modeling through seasonal trend decomposition, and the introduction of dual complementary regularization terms at the system and topology layers in the loss function.
[0007] To solve the above-mentioned technical problems, the present invention proposes the following technical solution: A two-layer complementary regularized joint prediction method for water, wind, and solar energy includes the following steps: Step 1: Obtain multi-source historical power data of the hydro-wind-solar hybrid system, and preprocess the multi-source historical power data to obtain a smoothed signal after noise reduction; Step 2: Construct a seasonal trend decomposition dual-channel spatiotemporal Transformer model, which includes a trend channel and a seasonal channel; Step 3: Input the smoothed signal obtained in Step 1 into the trend channel to extract low-frequency evolution features, and input the original sequence of multi-source historical power data into the seasonal channel to extract periodic patterns and short-term fluctuation features; Step 4: Construct a loss function that includes a two-layer complementarity regularization term; the two-layer complementarity regularization term includes a system-level regularization term and a topology-level regularization term; Step 5: Iteratively train the seasonal trend decomposition dual-channel spatiotemporal Transformer model using the loss function, and use the trained model to generate the joint power prediction results of the water-wind-solar complementary system.
[0008] In step 1, the multi-source historical power data is preprocessed using a wavelet soft thresholding denoising algorithm to obtain a smoothed signal after denoising. The smoothed signal is then divided into a training set, a validation set, and a test set according to a preset ratio, which are used for subsequent model training, hyperparameter selection, and performance evaluation, respectively. The preprocessing includes the following sub-steps: Step 1-1: Use Mallat's multi-scale decomposition algorithm to process the original power output sequence of the hydro-wind-solar hybrid system. Perform discrete wavelet transform to decompose it to the th Layers were used to obtain low-frequency scale coefficients that reflect long-term evolutionary trends. and high-frequency wavelet coefficients containing local perturbations and noise. ; Steps 1-2: Based on Stein's unbiased risk estimation theory, calculate the... Adaptive threshold for layer decomposition scale ; Steps 1-3: Using the soft thresholding function to evaluate the high-frequency wavelet coefficients Nonlinear shrinkage is performed to obtain the denoised high-frequency coefficients. When the coefficient amplitude is less than the threshold The value is set directly to zero when the coefficient amplitude is greater than or equal to the threshold. At that time, its amplitude is reduced by a threshold amount while keeping the sign unchanged; Steps 1-4: Utilizing unsuppressed low-frequency scaling coefficients and the high-frequency coefficients after soft thresholding. Performing inverse wavelet transform generates a denoised time series that retains the true evolutionary trend and eliminates high-frequency random disturbances. .
[0009] In step 1-1, firstly, the multi-source power output signal from water, wind, and solar is considered as a noisy time series, and its true value is expressed as: (1); In the formula, The real value of contributing to the water scenery An ideal output signal with no noise. To follow a standard normal distribution Gaussian white noise, Noise intensity; The non-stationary sequence is decomposed into different time-frequency scales using Discrete Wavelet Transform (DWT), and scale parameters are defined. With translation parameters for: (2); And with mother wavelet Generating subwavelets: (3); Then the signal is at scale ,Location The formula for calculating the wavelet coefficients at point is: (4); Using Mallat's multi-scale decomposition algorithm, a signal can be decomposed into a low-frequency main trend term and a high-frequency detail term: (5); In the formula, For scaling function, For the first The low-frequency scaling coefficients of the layer decomposition reflect the long-term evolution trend of water, wind, and solar power. These are high-frequency wavelet coefficients, which mainly contain local disturbances and noise.
[0010] In steps 1-2, to avoid over-smoothing or under-denoising due to a globally fixed threshold, a scale-adaptive thresholding scheme is adopted based on Stein's unbiased risk estimation theory; for the th Layer decomposition scale, with its adaptive threshold as follows: (6); In the formula, The length of the coefficient at this scale. This is the standard deviation of the detail coefficients at that scale, used to characterize the noise level at that scale; This is a scale correction factor used to adjust the threshold strength at this scale. It is obtained by minimizing the unbiased risk estimate of the denoised reconstructed signal at this scale, i.e.: (7); In the formula, In the correction factor and their corresponding thresholds The unbiased risk estimate corresponding to soft threshold compression under the action; through the above formula, the noise suppression intensity at different scales can be adaptively adjusted, instead of using a single global fixed threshold; In steps 1-3, a soft threshold function is introduced for the first... High-frequency wavelet coefficients of the layer Nonlinear shrinkage is performed to remove noise and retain useful information, resulting in processed coefficients. The calculation formula is as follows: (8); in, The threshold obtained in step 1-2 For sign functions. When When the coefficient is zero, it is considered the dominant noise component and is set to zero directly; when... At that time, the coefficient is according to Amplitude reduction is performed, and the amplitude is adjusted to a threshold value while keeping the sign unchanged. Instead of using the original amplitude for reconstruction, the amplitude is reduced. This operation suppresses high-frequency noise components while ensuring continuity at the threshold by shrinking the high-amplitude coefficients, making the reconstructed signal smoother and more natural.
[0011] After soft thresholding compression is completed in steps 1-4, the processed detail coefficients are used. and unsuppressed scaling factor Perform inverse wavelet transform to obtain the denoised time series. : (9); Reconstructed sequence While accurately depicting the real evolution trend and dynamic characteristics of water, wind, and solar power output, it effectively suppresses high-frequency random disturbances.
[0012] In step 2, a two-channel spatiotemporal Transformer model for seasonal trend decomposition is constructed, which specifically includes the following sub-steps: Step 2-1: Construct the spatial topology of the water-wind-solar hybrid system; Step 2-2: Define the spatial adjacency matrix; Steps 2-3: Build the overall architecture of the dual-channel model; Steps 2-4: Initialize the ST-Former network layer structure and parameters; In step 2-1, a spatial information model for joint prediction of water, wind, and light is established using graph theory methods, and an undirected graph is defined. ;in, This represents the set of all power station nodes in a hydro-wind-solar hybrid system. This represents the total number of nodes. The set of edges represents the physical or electrical connections between nodes. With nodes If there is an actual electrical interconnection, then in If there is a connecting edge, then there is no connecting edge. In step 2-2, based on the spatial topology map Define the adjacency matrix This is used to quantify the spatial topological relationships between nodes; the elements in the matrix represent the connectivity of edges, and if there is a connection between nodes, the corresponding element is non-zero, which is used for subsequent sparsity masking operations in the spatial attention mechanism. In steps 2-3, a seasonal trend decomposition dual-channel spatiotemporal Transformer model (STD-ST-Former) is constructed, configuring two parallel processing channels: Trend channel: used to receive the decomposed trend-stable components, and learn the long-term stationary characteristics of the power sequence through the configured ST-Former network branches; Seasonal channel: used to receive the decomposed seasonal fluctuation components and learn the short-term dynamic characteristics of the power sequence through the configured ST-Former network branch; The output of the seasonal trend decomposition dual-channel spatiotemporal Transformer model is equipped with a fusion module, which is used to linearly superimpose the prediction results of the trend channel and the seasonal channel. Wherein, the adjacency matrix defined in step 2-2 It is integrated into the spatial self-attention unit to guide feature aggregation calculation; In steps 2-4, stacking is performed in the trend channel and the seasonal channel respectively. Layered spatiotemporal attention modules; each layer module contains variable patch temporal attention units and spatial self-attention units; network hyperparameters are set, including input data length. Feature Dimension Hidden layer dimension ; Regarding the first Layered network, defining the patch partitioning parameters: input sequence length and Patch step size The input sequence is divided into The first patch; initializes the learnable parameters of the model, including: the first patch. The first layer The first variable A fake timestamp Temporal attention projection matrix , Spatial self-attention projection matrix , , ; and cross-layer skip connection parameters and output layer perceptron parameters and .
[0013] In step 3, the processed data is input into the seasonal trend decomposition dual-channel spatiotemporal Transformer model to extract features and generate prediction results. This includes the following sub-steps: Step 3-1: Perform time domain decomposition; Step 3-2: Calculate variable patch-time attention; Step 3-3: Calculate self-attention in the sparse space; Steps 3-4: Multi-scale spatiotemporal feature fusion; Steps 3-5: Dual-channel prediction generation and fusion.
[0014] In step 3-1, the multi-source historical power sequence is first processed. Perform time-domain decomposition, breaking it down into trend-stable components. With seasonal fluctuation components ,in, For the length of the input data, The feature dimension is used; the trend-stabilizing component is obtained through a moving average calculation. (10); Next, the trend component is subtracted from the original sequence in the form of a difference to extract the seasonal fluctuation component. The calculation formula is as follows: (11); In the formula, This indicates a moving average operation. This indicates the filling operation. Through this decomposition mechanism, the model constructs trend channels and seasonal channels respectively, enabling the model to focus on dynamic patterns at two different time scales during learning: long-term stable trends and short-term random fluctuations. In step 3-2, the temporal feature extraction stage, a variable patch temporal attention mechanism is used to adapt to input sequences of different lengths; the first... Layer length is The input sequence is divided along the time dimension into There are 1 patch, of which... Indicates the first The patch step size of the layer; let the first layer be... The node at the th The input of each patch is ,in, express The Middle One input, , Define a learnable pseudo-timestamp , Indicates the first The first layer The first variable A pseudo-timestamp is used as a query matrix in the variable patch temporal attention to query all input values in the patch. This represents the hidden dimension. The output of the variable patch-time attention is: (12); In the formula, Indicates the first The first layer The first variable A variable patch-time attention output For variable-patch-time attention operations, express Activation function , Indicates the first The first layer A learnable temporal projection matrix of variable patch-time attention for each variable.
[0015] In step 3-3, a spatial self-attention mechanism is introduced to mine all elements in each layer. The first variable The correlation between patches is used, and the adjacency matrix defined in step 2 is utilized. A sparsity masking operation is performed to preserve the main physical correlations, and the spatial self-attention output is: (13); In the formula, Indicates the first The first layer of topological variables Spatial self-attention output between patches For attention operations in the variable patch space, Indicates the first All layers The first variable A fake timestamp , , Indicates the first The first layer A spatial projection matrix that can be learned through spatial self-attention; Represents the Hadamard product of matrices; In steps 3-4, to simultaneously capture spatiotemporal features at different scales, the models are stacked. A spatiotemporal attention module is used, and multi-scale fusion is achieved through cross-layer skip connections, resulting in the final spatiotemporal feature output. for: (14); in, For the first The output sequence of the layer after spatiotemporal attention processing. Indicates the first Learnable parameters of layer skip connections, ; In steps 3-5, a two-layer perceptual network is used to map spatiotemporal features to power prediction values; Model prediction output for each channel The calculation formula is: (15); Indicates the predicted output. Represents the activation function of the rectified linear unit. To predict future sequence length, and Indicates learnable parameters; Based on the aforementioned dual-channel integration mechanism, the model uses trend-stabilized components respectively. and seasonal fluctuation components Using the input as input, the predicted values of the trend components are calculated through two parallel ST-Former branches. and seasonal component forecast values Finally, the outputs of the two branches are linearly superimposed to obtain the final joint power prediction result of the hydro-wind-solar hybrid system. : (16).
[0016] In step 4, a loss function containing a two-layer complementarity regularization term is constructed. This two-layer complementarity regularization term includes a system-level regularization term and a topology-level regularization term. The system-level regularization term constructs a global complementarity index based on a normalized weighted aggregation sequence, used to characterize multi-source complementary coupling features. The topology-level regularization term targets electrically adjacent nodes, combining cosine similarity and hinge functions to characterize spatial correlation, and is adjusted using gating weights. The construction of the loss function containing the two-layer complementarity regularization term specifically includes the following sub-steps: Step 4-1: Define the general formula for the composite loss function; Step 4-2: Construct a global complementarity regularization term at the system level; Step 4-3: Construct the local complementarity regularization term for the topology layer; In step 4-1, a composite loss function with two-layer complementary regularization is constructed. This function introduces a system-level global complementary regularization term and a topology-level local complementary regularization term into the basic mean square error term. Its expression is: (17); in, This term represents the mean squared error between the model's predicted values and the actual values. and These are the global complementary regularization term at the system layer and the local complementary regularization term at the topology layer, respectively. For the weighting coefficients, this invention employs a grid search strategy based on the validation set to determine the optimal weight combination. Considering the differences in numerical magnitude among different loss terms, this invention presupposes a search space for the weighting coefficients. By comparing the relative root mean square error of the model's total power prediction on the validation set under different parameter combinations, and using error minimization as the criterion, the final result was determined. , ; In step 4-2, in order to guide the model to learn the dynamic coupling law between heterogeneous energy sources, a system-level regularization term is constructed by referring to and expanding the complementarity evaluation index; First, based on the installed capacity of wind power, hydropower, and photovoltaic power. , , Define capacity ratio coefficient , , : (18); Next, a weighted power sequence of the three energy sources and the benchmark complementary pair is constructed to reflect the energy substitution relationship between different energy sources: (19); in, , , They are time points The output of wind power, hydropower, and photovoltaic power; , and These are the weighted power values of wind power, hydropower, and photovoltaic power complements to the benchmark, respectively. Furthermore, for any time step pair in the predicted sequence , Calculate the change in power And define the temporal complementarity coefficient as: (20); in The total duration of the sample. For time-series complementarity coefficients, , , This is the capacity reduction factor, used to correct for the impact of differences in installed capacity on complementarity calculations; Finally, the inverse of the temporal complementarity coefficient is defined as the system-level global complementarity regularization term: (twenty one); When the model predicts a more significant complementary relationship between water, wind, and solar power outputs over time, This reduces the frequency of parameter updates, thus guiding the direction of parameter updates during backpropagation to better align with the energy distribution patterns of multi-energy systems. In step 4-3, in order to prevent spatial feature aliasing and reflect regional complementarity, a topology layer regularization term is constructed based on the correlation of node prediction results. First, compute nodes , The predicted output is , The cosine similarity after removing the mean is: (twenty two); in, Represents a node and The cosine similarity of the predicted output sequence after removing the mean. , This represents the average value of the corresponding sequences. Based on similarity, the local complementarity constraint function is defined as: (twenty three); In the formula, For nodes and Local complementary constraint values between them This is a similarity threshold used to define the maximum permissible correlation between prediction results from different power plants; Combined with adjacency matrix In addition to the historical correlation adjustment coefficient, adaptive gating weights are defined: (twenty four); in, For nodes and Adaptive gating weights between them Derived from the adjacency matrix , representing the spatial connection strength between nodes; These are learnable gating parameters used to dynamically adjust the regularization strength. To prevent the stability constant from having a denominator of zero; The final topology layer complementarity regularization term is: (25); in It represents the set of adjacent nodes across energy types; by introducing a spatially complementary structure, this term enables the model to enhance its ability to identify the output characteristics of different power plants at the feature representation level during the training process.
[0017] Step 5 involves iteratively training the model using the loss function and generating prediction results, specifically including the following sub-steps: Step 5-1: Iterative model training and parameter optimization; Step 5-2: Training terminated and model saved; Step 5-3: Perform offline prediction; Step 5-4: Inverse normalization reconstruction of the prediction results; In step 5-1, the pre-defined training set is input into the seasonal trend decomposition dual-channel spatiotemporal Transformer model; in each iteration, forward propagation is performed, and the current prediction output is generated according to the formula in step 3. ; The composite loss function constructed in step S4 Calculate the total error between the current predicted value and the actual value; The loss function is calculated using the backpropagation algorithm. The gradients of all learnable parameters of the model; these learnable parameters include the temporal attention projection matrix. , Spatial self-attention projection matrix , , ; and cross-layer skip connection parameters and output layer perceptron parameters and ; The gradient descent optimization algorithm is used to update the above parameters based on the calculated gradient in order to minimize the composite loss function. ; During the parameter update process, the system-level global complementary regularization term The topology layer local complementary regularization term guides the model parameters to converge in a direction that conforms to the energy distribution law of multiple energy sources. Suppress spatial feature aliasing, thereby co-optimizing the model's spatiotemporal expressive power during training; In step 5-2, step 5-1 is repeated until a preset convergence condition is met; the convergence condition includes reaching a preset maximum number of iterations, or the loss function value on the validation set no longer decreasing within a certain number of consecutive cycles; After training, the optimal model parameters are saved to obtain the trained power joint prediction model of the hydro-wind-solar hybrid system. In step 5-3, the historical power data of the water-wind-solar hybrid system for the period to be predicted is obtained, and preprocessed according to the wavelet soft thresholding denoising method described in step 1 and the normalization method described in steps 1-3 to generate the model input sequence. The input sequence is fed into the trained model, and through parallel processing and fusion of the trend channel and seasonal channel, a normalized future power prediction sequence is output. ; In step 5-4, the normalized prediction sequence is... Perform an inverse normalization operation to map it back to the original power dimension. The calculation formula is as follows: (26); in, For the final power joint prediction results of the hydro-wind-solar hybrid system, and These are the maximum and minimum values of the historical training data recorded in steps 1-3, respectively. The final output includes the future. The predicted power values for hydropower, wind power, and photovoltaic power at each time step, as well as the predicted total power value of the complementary system.
[0018] Existing deep learning models, when predicting the power of multi-source systems such as water, wind, and solar, mix long-period stable evolution patterns with short-period random fluctuation characteristics, which easily leads to the aliasing of temporal features. At the same time, conventional Transformer models lack constraints on the physical spatial topology when processing multi-site data, resulting in the model's inability to effectively capture local electrical coupling correlations, and a severe decrease in prediction accuracy under complex meteorological fluctuation conditions. To solve this technical problem, this invention also provides a joint prediction system for water, wind, and solar power based on a dual-channel spatiotemporal Transformer with seasonal trend decomposition.
[0019] The system includes: The data acquisition and preprocessing module: Its input terminal is used to input the multi-source historical power sequences of each power station in the hydro-wind-solar hybrid system. The data acquisition and preprocessing module is used to perform noise reduction processing on the historical data and define an adjacency matrix based on the spatial topological relationship of each station. ; Time decomposition module: Its input end and the output end of the data acquisition and preprocessing module are connected. It is used to receive the processed multi-source historical power sequence and perform time domain decomposition, which splits it into a trend stable component that reflects the long-period smooth law and a seasonal fluctuation component that reflects the short-period random fluctuation. Dual-channel prediction model module: Its input end is connected to the output end of the time decomposition module. The dual-channel prediction model module is configured with parallel trend channel and seasonal channel. The dual-channel prediction model module is used to extract the spatiotemporal features of the trend stable component through the trend channel and output the predicted value of the trend component, and at the same time, extract the spatiotemporal features of the seasonal fluctuation component through the seasonal channel and output the predicted value of the seasonal fluctuation component. Result fusion module: Its input end is connected to the output end of the dual-channel prediction model module, and it is used to receive the trend component prediction value and the seasonal component prediction value, and linearly superimpose the two to generate the final power joint prediction result of the water-wind-solar complementary system. The dual-channel prediction model module further integrates: 1) Time attention unit: used to capture the dynamic dependencies of the trend-stabilized component or the seasonal fluctuation component at different time scales through a variable patch mechanism; 2) Spatial Attention Unit: Connected to the matrix output of the data acquisition and preprocessing module, used to utilize the adjacency matrix. As a mask constraint, it aggregates the physical complementary features between adjacent stations. When the system is in operation or in use, the following steps are included: Step 1: Obtain the multi-source historical power sequences of each power station in the hydro-wind-solar hybrid system. And define an adjacency matrix based on the spatial topology of each station. ; Step 2: Perform time-domain decomposition on the multi-source historical power series to obtain trend-stable components and seasonal fluctuation components, which are then input into the trend channel and seasonal channel of the dual-channel spatiotemporal Transformer model, respectively. Step 3: Perform spatiotemporal attention calculations on the input trend stability component and seasonal fluctuation component in the trend channel and seasonal channel, respectively. This includes the following sub-steps: Step 3-1: Extract temporal features using a variable patch temporal attention mechanism and calculate the variable patch temporal attention output; Step 3-2: Introduce a spatial self-attention mechanism and utilize the adjacency matrix defined in Step 1. Perform a sparsity masking operation to obtain a spatial self-attention output; Step 3-3: Stack multiple spatiotemporal attention modules and achieve multi-scale spatiotemporal feature fusion through cross-layer skip connections to obtain the final spatiotemporal feature output; Step 4: Perform dual-channel prediction and fusion generation; including: Step 4-1: Use a perceptual network to map the spatiotemporal feature output to the power prediction value, and obtain the trend component prediction value of the trend channel and the seasonal component prediction value of the seasonal channel respectively. Step 4-2: Linearly superimpose the predicted values of the trend component and the seasonal component to obtain the final power joint prediction result of the hydro-wind-solar hybrid system; this result is used to guide the multi-energy complementary joint dispatch of the power grid and the formulation of day-ahead power generation plans. In step 2, the multi-source historical power sequences are first processed. Perform time-domain decomposition, breaking it down into trend-stable components. With seasonal fluctuation components ,in, For the length of the input data, The feature dimension is used; the trend-stabilizing component is obtained through a moving average calculation. ; Next, the trend component is subtracted from the original sequence in the form of a difference to extract the seasonal fluctuation component. The calculation formula is as follows: ; In the formula, This indicates a moving average operation. This indicates the filling operation. Through this decomposition mechanism, the model constructs trend channels and seasonal channels respectively, enabling the model to focus on dynamic patterns at two different time scales during learning: long-term stable trends and short-term random fluctuations. In step 3-1, the formula for calculating the variable patch temporal attention output is: ; In the formula, Indicates the first The first layer The first variable A variable patch-time attention output For variable-patch-time attention operations, express Activation function , Indicates the first The first layer A variable patch-time attention learnable temporal projection matrix for each variable; In step 3-2, the formula for calculating the spatial self-attention output is: ; In the formula, Indicates the first The first layer of topological variables Spatial self-attention output between patches For attention operations in the variable patch space, Indicates the first All layers The first variable A fake timestamp , , Indicates the first The first layer A spatial projection matrix that can be learned through spatial self-attention; Represents the Hadamard product of matrices; In step 3-3, the final spatiotemporal feature output The calculation formula is: ; in, For the first The output sequence of the layer after spatiotemporal attention processing. Indicates the first The learnable parameters of the layer skip connection, the summation formula above This represents the total number of stacked layers of the spatiotemporal attention modules; In step 4, the calculation formula for prediction and fusion generation is: Model prediction output for each channel (the channel being either a trend channel or a seasonal channel) for: ; in, Represents the activation function of the rectified linear unit. To predict future sequence length, and Indicates learnable parameters; Based on the aforementioned dual-channel integration mechanism, the model uses trend-stabilized components respectively. and seasonal fluctuation components Using the input as input, the predicted values of the trend components are calculated through two parallel ST-Former branches. and seasonal component forecast values Finally, the outputs of the two branches are linearly superimposed to obtain the final joint power prediction result of the hydro-wind-solar hybrid system. The calculation formula is: ; Compared with the prior art, the present invention has the following technical effects: 1) The prediction method of this invention deeply couples adaptive data preprocessing, a dual-channel spatiotemporal feature decoupling mechanism, and a physically heuristic composite loss function, systematically overcoming the performance bottlenecks of existing pure data-driven models when faced with high-frequency noise interference, multi-scale feature aliasing, and prediction outputs that violate the law of energy conservation. The final jointly predicted power sequence not only achieves a significant reduction in mean square error at the numerical level, but also highly conforms to the real scheduling law of multi-energy complementary systems in terms of energy time-series evolution, thus providing highly reliable decision-making data support for the formulation of day-ahead power generation plans, multi-energy collaborative optimization scheduling, and the safe and stable operation of the power grid in new power systems. 2) This invention employs an adaptive wavelet soft-threshold denoising method based on unbiased risk estimation. Compared to traditional fixed-threshold methods, it can automatically adjust the threshold according to the noise level at different decomposition scales. This effectively removes high-frequency random noise while preserving the abrupt changes and true evolution trends of the power signal to the greatest extent, providing high-quality input data for subsequent models. 3) This invention constructs an STD-ST-Former model, employing a trend-seasonal dual-channel decoupled architecture. This design allows the model to focus on learning long-term stationary patterns and short-term random fluctuations respectively, solving the problem of chaotic feature extraction in traditional models when dealing with mixed complex patterns. Simultaneously, by combining variable patch time and a sparse spatial attention mechanism, it can accurately capture the complex dependencies of the water, wind, and light system at different spatiotemporal scales. 4) This invention introduces a two-layer complementarity regularization term. The system-level regularization term incorporates the macroscopic energy complementarity mechanism into the loss function, guiding the model to actively learn the energy substitution relationship between multiple sources. The topology-level regularization term utilizes the graph topology structure to constrain spatial correlation, effectively suppressing prediction bias caused by spatial feature smoothing and improving the model's ability to identify local issues. Attached Figure Description
[0020] The present invention will be further described below with reference to the accompanying drawings and embodiments: Figure 1 This is an overall flowchart of the present invention; Figure 2 This is a diagram of the ST-Former model in this invention; Figure 3 This is a schematic diagram of the hydropower prediction results of the present invention; Figure 4 This is a schematic diagram of the wind power prediction results of the present invention; Figure 5 This is a schematic diagram of the photovoltaic power prediction results of the present invention; Figure 6 This is a schematic diagram of the total power prediction results of the present invention; Figure 7 This is a comparison chart of the prediction accuracy results of the various models in this invention. Detailed Implementation
[0021] like Figure 1 The flowchart shown is a two-layer complementary regularized joint prediction method for water, wind, and solar power provided by this invention. This invention first acquires the original spatiotemporal dataset of multi-source water, wind, and solar power sequences and divides it into training, validation, and test sets. For the training set, high-frequency noise is removed using a wavelet soft thresholding denoising algorithm, and then seasonal-trend decomposition is performed to split the sequence into trend and seasonal components. These two types of components are then fed into the ST-Former model as dual-channel inputs for feature extraction and aggregation output. Based on this, a total loss function containing system-level and topological-level complementary constraints is introduced, and the model parameters are iteratively updated using a backpropagation algorithm. Simultaneously, the training process is monitored using the validation set after the same decomposition process, and the optimal model is selected based on the validation error. Finally, the test set is input into the trained ST-Former model, and the final prediction data is aggregated and analyzed for model evaluation, thereby achieving high-precision joint power prediction of water, wind, and solar power that conforms to physical complementarity characteristics.
[0022] Specifically, the following steps are included: Step 1: Obtain multi-source historical power data of the hydro-wind-solar hybrid system, and preprocess the multi-source historical power data using a wavelet soft threshold denoising algorithm to obtain a denoised smooth signal; the hydro-wind-solar hybrid system is a system in the prior art, such as a "hydro-wind-solar hybrid system" described in "Evaluation of New Energy Bundling Capacity of Hydro-Wind-Solar Hybrid System under Hydropower Regulation Capacity Constraints"; In step 1-1, firstly, the multi-source power output signal from water, wind, and solar is considered as a noisy time series, and its true value is expressed as: (1); In the formula, The real value of contributing to the water scenery An ideal output signal with no noise. To follow a standard normal distribution Gaussian white noise, Noise intensity.
[0023] The non-stationary sequence is decomposed into different time-frequency scales using Discrete Wavelet Transform (DWT), and scale parameters are defined. With translation parameters for: (2); And with mother wavelet Generating subwavelets: (3); Then the signal is at scale ,Location The formula for calculating the wavelet coefficients at point is: (4); Using Mallat's multi-scale decomposition algorithm, a signal can be decomposed into a low-frequency main trend term and a high-frequency detail term: (5); In the formula, For scaling function, For the first The low-frequency scaling coefficients of the layer decomposition reflect the long-term evolution trend of water, wind, and solar power. These are high-frequency wavelet coefficients, which mainly contain local disturbances and noise.
[0024] In steps 1-2, to avoid oversmoothing or under-denoising due to a globally fixed threshold, a scale-adaptive threshold scheme is adopted based on Stein's unbiased risk estimation theory. For the Layer decomposition scale, with its adaptive threshold as follows: (6); In the formula, The length of the coefficient at this scale. This is the standard deviation of the detail coefficients at that scale, used to characterize the noise level at that scale; This is a scale correction factor used to adjust the threshold strength at this scale. It is obtained by minimizing the unbiased risk estimate of the denoised reconstructed signal at this scale, i.e.: (7); In the formula, In the correction factor and their corresponding thresholds The unbiased risk estimate is obtained after soft threshold compression under the influence of the above equation. This allows for adaptive adjustment of noise suppression intensity at different scales, rather than using a single, globally fixed threshold.
[0025] In steps 1-3, a soft threshold function is introduced for the first... High-frequency wavelet coefficients of the layer Nonlinear shrinkage is performed to remove noise and retain useful information, resulting in processed coefficients. The calculation formula is as follows: (8); in, The threshold obtained in step 1-2 For sign functions. When At that time, the coefficient is considered the dominant noise component and is set to zero directly; when At that time, the coefficient is according to Amplitude reduction is performed, and the amplitude is adjusted to a threshold value while keeping the sign unchanged. Reduce, rather than using the original amplitude value in the reconstruction; This operation suppresses high-frequency noise components while ensuring continuity at the threshold by shrinking high-amplitude coefficients, making the reconstructed signal smoother and more natural. After soft thresholding compression is completed in steps 1-4, the processed detail coefficients are used. and unsuppressed scaling factor Perform inverse wavelet transform to obtain the denoised time series. : (9); Reconstructed sequence While accurately depicting the real evolution trend and dynamic characteristics of water, wind, and solar power output, it effectively suppresses high-frequency random disturbances.
[0026] Step 2: Construct a seasonal trend decomposition dual-channel spatiotemporal Transformer model, which includes a trend channel and a seasonal channel; like Figure 2 As shown in the ST-Former model diagram of this invention, the ST-Former model adopts a multi-layer stacked spatiotemporal feature extraction architecture. Its logical flow begins with the multi-source historical power sequence input at the bottom, and the data flows layer by layer from bottom to top. Within each layer, time-dimensional feature mining is first performed, that is, the input sequence is divided into several local patches, and a learnable pseudo-timestamp vector is used as a query matrix. The time feature vector containing the local temporal evolution law is extracted through a variable patch time attention mechanism. Subsequently, the time feature is sent to the sparse spatial self-attention module. Here, the model aggregates the time features of all nodes in the current layer, calculates the spatial dependencies between nodes, and generates hierarchical spatiotemporal features that integrate spatial coupling information. In order to prevent feature degradation in deep networks and to integrate information from different levels of abstraction, the model uses a cross-layer jump connection mechanism to converge the feature sequences output by each layer to the addition node for linear superposition, forming a global spatiotemporal feature representation. Finally, this global feature is sequentially processed through the weights of the first projection layer. , Activation function and weights of the second projection layer The perceptron network completes the mapping from the high-dimensional feature space to the prediction step size dimension, and finally outputs the power prediction sequence for future time steps. .
[0027] In step 2-1, a spatial information model for joint prediction of water, wind, and light is established using graph theory methods, and an undirected graph is defined. ; in, This represents the set of all power station nodes in a hydro-wind-solar hybrid system. This represents the total number of nodes. The set of edges represents the physical or electrical connections between nodes. With nodes If there is an actual electrical interconnection, then in If there is a connecting edge, then there is no connecting edge. In step 2-2, based on the spatial topology map Define the adjacency matrix This is used to quantify the spatial topological relationships between nodes; the elements in the matrix represent the connectivity of edges, and if there is a connection between nodes, the corresponding element is non-zero, which is used for subsequent sparsity masking operations in the spatial attention mechanism. In steps 2-3, a seasonal trend decomposition dual-channel spatiotemporal Transformer model (STD-ST-Former) is constructed, configuring two parallel processing channels: Trend channel: It is equipped with a stationary branch network, whose input is configured to receive the trend stable component obtained by time domain decomposition of the multi-source historical power sequence, and is used to learn the long-term stationary characteristics of the power sequence. Seasonal channel: It is equipped with a fluctuation branch network, whose input is configured to receive the seasonal fluctuation component obtained by time domain decomposition of the multi-source historical power sequence, and to learn the short-term dynamic characteristics of the power sequence. The model output is equipped with a fusion module, which is used to linearly superimpose the prediction results of the two channels; In steps 2-4, stacking is performed in the trend channel and the seasonal channel respectively. Layered spatiotemporal attention module; Each layer module contains variable patch temporal attention units and spatial self-attention units; network hyperparameters are set, including input data length. Feature Dimension Hidden layer dimension ; Regarding the first Layered network, defining the patch partitioning parameters: input sequence length and Patch step size The input sequence is divided into One patch; The learnable parameters for initializing the model include: the first... The first layer The first variable A fake timestamp ; Temporal attention projection matrix , ; Spatial self-attention projection matrix , , ; and cross-layer skip connection parameters and output layer perceptron parameters and .
[0028] Step 3: Input the smoothed signal obtained in Step 1 into the trend channel to extract low-frequency evolution features, and input the original sequence of the multi-source historical power data into the seasonal channel to extract periodic patterns and short-term fluctuation features; In step 3-1, the multi-source historical power sequence is first processed. Perform time-domain decomposition, breaking it down into trend-stable components. With seasonal fluctuation components ,in, For the length of the input data, The feature dimension is defined as follows. The trend-stabilizing component is obtained through a moving average calculation: (10); Next, the trend component is subtracted from the original sequence in the form of a difference to extract the seasonal fluctuation component. The calculation formula is as follows: (11); In the formula, This indicates a moving average operation. This indicates the filling operation. Through this decomposition mechanism, the model constructs trend channels and seasonal channels respectively, enabling the model to focus on dynamic patterns at two different time scales during learning: long-term stable trends and short-term random fluctuations. In step 3-2, the temporal feature extraction stage, a variable patch temporal attention mechanism is used to adapt to input sequences of different lengths; the first... Layer length is The input sequence is divided along the time dimension into There are 1 patch, of which... Indicates the first Layer patch step size; Let the first The node at the th The input of each patch is ,in, express The Middle One input, , Define a learnable pseudo-timestamp , Indicates the first The first layer The first variable A pseudo-timestamp is used as a query matrix in the variable patch temporal attention to query all input values in the patch. Indicates hidden dimensions; The variable patch-time attention output is: (12); In the formula, Indicates the first The first layer The first variable A variable patch-time attention output For variable-patch-time attention operations, , Indicates the first The first layer A learnable temporal projection matrix of variable patch-time attention for each variable.
[0029] In step 3-3, a spatial self-attention mechanism is introduced to mine all elements in each layer. The first variable The correlation between patches is used, and the adjacency matrix defined in step 2 is utilized. A sparsity masking operation is performed to preserve the main physical correlations, and the spatial self-attention output is: (13); In the formula, Indicates the first The first layer of topological variables Spatial self-attention output between patches For attention operations in the variable patch space, Indicates the first All layers The first variable A fake timestamp , , Indicates the first The first layer A spatial projection matrix that can be learned through spatial self-attention.
[0030] In steps 3-4, to simultaneously capture spatiotemporal features at different scales, the models are stacked. A spatiotemporal attention module is used, and multi-scale fusion is achieved through cross-layer skip connections, resulting in the final spatiotemporal feature output. for: (14); in, For the first The output sequence of the layer after spatiotemporal attention processing. Indicates the first Learnable parameters of layer skip connections, ; In steps 3-5, a two-layer perceptual network is used to map spatiotemporal features to power prediction values; Model prediction output for each channel The calculation formula is: (15); Indicates the predicted output. Represents the activation function of the rectified linear unit. To predict the step size, and Indicates learnable parameters; Based on the aforementioned dual-channel integration mechanism, the model uses trend-stabilized components respectively. and seasonal fluctuation components Using the input as input, the predicted values of the trend components are calculated through two parallel ST-Former branches. and seasonal component forecast values Finally, the outputs of the two branches are linearly superimposed to obtain the final joint power prediction result of the hydro-wind-solar hybrid system. : (16); Step 4: Construct a loss function containing a two-layer complementarity regularization term; the two-layer complementarity regularization term includes a system-level regularization term and a topology-level regularization term; wherein, the system-level regularization term constructs a global complementarity index based on a normalized weighted aggregation sequence to characterize multi-source complementary coupling features; the topology-level regularization term targets electrically adjacent nodes, combines cosine similarity and hinge function to characterize spatial correlation, and is adjusted through gating weights; In step 4-1, a composite loss function with two-layer complementary regularization is constructed. This function introduces a system-level global complementary regularization term and a topology-level local complementary regularization term into the basic mean square error term. Its expression is: (17); in, This term represents the mean squared error between the model's predicted values and the actual values. and These are the global complementary regularization term at the system layer and the local complementary regularization term at the topology layer, respectively. For the weighting coefficients, this invention employs a grid search strategy based on the validation set to determine the optimal weight combination. Considering the differences in numerical magnitude among different loss terms, this invention presupposes a search space for the weighting coefficients. By comparing the relative root mean square error of the model's total power prediction on the validation set under different parameter combinations, and using error minimization as the criterion, the final result was determined. , ; In step 4-2, in order to guide the model to learn the dynamic coupling law between heterogeneous energy sources, a system-level regularization term is constructed by referring to and expanding the complementarity evaluation index; First, based on the installed capacity of wind power, hydropower, and photovoltaic power. , , Define capacity ratio coefficient , , : (18); Next, a weighted power sequence of the three energy sources and the benchmark complementary pair is constructed to reflect the energy substitution relationship between different energy sources: (19); in, , , They are time points Wind power, hydropower, and photovoltaic power output, , and These are the weighted power values of wind power, hydropower, and photovoltaic power complements to the benchmark, respectively. Furthermore, for any time step pair in the predicted sequence , Calculate the change in power And define the temporal complementarity coefficient as: (20); in The total duration of the sample. For time-series complementarity coefficients, , , This is the capacity reduction factor, used to correct for the impact of differences in installed capacity on complementarity calculations; Finally, the inverse of the temporal complementarity coefficient is defined as the system-level global complementarity regularization term: (twenty one); When the model predicts a more significant complementary relationship between water, wind, and solar power outputs over time, This reduces the frequency of parameter updates, thus guiding the direction of parameter updates during backpropagation to better align with the energy distribution patterns of multi-energy systems. In step 4-3, in order to prevent spatial feature aliasing and reflect regional complementarity, a topology layer regularization term is constructed based on the correlation of node prediction results. First, compute nodes , The predicted output is , The cosine similarity after removing the mean is: (twenty two); in, Represents a node and The cosine similarity of the predicted output sequence after removing the mean. , This represents the average value of the corresponding sequences. Based on similarity, the local complementarity constraint function is defined as: (twenty three); In the formula, For nodes and Local complementary constraint values between them This is a similarity threshold used to define the maximum permissible correlation between prediction results from different power plants; Combined with adjacency matrix In addition to the historical correlation adjustment coefficient, adaptive gating weights are defined: (twenty four); in, For nodes and Adaptive gating weights between them Derived from the adjacency matrix , representing the spatial connection strength between nodes; These are learnable gating parameters used to dynamically adjust the regularization strength. To prevent the stability constant from having a denominator of zero; The final topology layer complementarity regularization term is: (25); in It represents the set of adjacent nodes across energy types; by introducing a spatially complementary structure, this term enables the model to enhance its ability to identify the output characteristics of different power plants at the feature representation level during the training process.
[0031] Step 5: Iteratively train the seasonal trend decomposition dual-channel spatiotemporal Transformer model using the loss function, and use the trained model to generate the joint power prediction results of the water-wind-solar complementary system.
[0032] In step 5-1, the pre-defined training set is input into the seasonal trend decomposition dual-channel spatiotemporal Transformer model; in each iteration, forward propagation is performed, and the current prediction output is generated according to the formula in step 3. ; Using the composite loss function constructed in step 4 Calculate the total error between the current predicted value and the actual value; The loss function is calculated using the backpropagation algorithm. The gradients of all learnable parameters of the model; these learnable parameters include the temporal attention projection matrix. , Spatial self-attention projection matrix , , ; and cross-layer skip connection parameters and output layer perceptron parameters and ; The gradient descent optimization algorithm is used to update the above parameters based on the calculated gradient in order to minimize the composite loss function. ; During the parameter update process, the system-level global complementary regularization term The topology layer local complementary regularization term guides the model parameters to converge in a direction that conforms to the energy distribution law of multiple energy sources. Suppress spatial feature aliasing, thereby co-optimizing the model's spatiotemporal expressive power during training; In step 5-2, step 5-1 is repeated until a preset convergence condition is met; the convergence condition includes reaching a preset maximum number of iterations, or the loss function value on the validation set no longer decreasing within a certain number of consecutive cycles; After training, the optimal model parameters are saved to obtain the trained power joint prediction model of the hydro-wind-solar hybrid system. In step 5-3, the historical power data of the water-wind-solar hybrid system for the period to be predicted is obtained, and preprocessed according to the wavelet soft thresholding denoising method described in step 1 and the normalization method described in steps 1-3 to generate the model input sequence. The input sequence is fed into the trained model, and through parallel processing and fusion of the trend channel and seasonal channel, a normalized future power prediction sequence is output. ; In step 5-4, the normalized prediction sequence is... Perform an inverse normalization operation to map it back to the original power dimension. The calculation formula is as follows: (26); in, For the final power joint prediction results of the hydro-wind-solar hybrid system, and These are the maximum and minimum values of the historical training data recorded in steps 1-3, respectively; The final output includes the future. The predicted power values for hydropower, wind power, and photovoltaic power at each time step, as well as the predicted total power value of the complementary system.
[0033] Existing deep learning models, when predicting the power of multi-source systems such as water, wind, and solar, mix long-period stable evolution patterns with short-period random fluctuation characteristics, which easily leads to the aliasing of temporal features. At the same time, conventional Transformer models lack constraints on the physical spatial topology when processing multi-site data, resulting in the model's inability to effectively capture local electrical coupling correlations, and a severe decrease in prediction accuracy under complex meteorological fluctuation conditions. To solve this technical problem, this invention also provides a joint prediction system for water, wind, and solar power based on a dual-channel spatiotemporal Transformer with seasonal trend decomposition.
[0034] The system includes: The data acquisition and preprocessing module: Its input terminal is used to input the multi-source historical power sequences of each power station in the hydro-wind-solar hybrid system. The data acquisition and preprocessing module is used to perform noise reduction processing on the historical data and define an adjacency matrix based on the spatial topological relationship of each station. ; Time decomposition module: Its input end and the output end of the data acquisition and preprocessing module are connected. It is used to receive the processed multi-source historical power sequence and perform time domain decomposition, which splits it into a trend stable component that reflects the long-period smooth law and a seasonal fluctuation component that reflects the short-period random fluctuation. Dual-channel prediction model module: Its input end is connected to the output end of the time decomposition module. The dual-channel prediction model module is configured with parallel trend channel and seasonal channel. The dual-channel prediction model module is used to extract the spatiotemporal features of the trend stable component through the trend channel and output the predicted value of the trend component, and at the same time, extract the spatiotemporal features of the seasonal fluctuation component through the seasonal channel and output the predicted value of the seasonal fluctuation component. Result fusion module: Its input end is connected to the output end of the dual-channel prediction model module, and it is used to receive the trend component prediction value and the seasonal component prediction value, and linearly superimpose the two to generate the final power joint prediction result of the water-wind-solar complementary system. The dual-channel prediction model module further integrates: 1) Time attention unit: used to capture the dynamic dependencies of the trend-stabilized component or the seasonal fluctuation component at different time scales through a variable patch mechanism; 2) Spatial Attention Unit: Connected to the matrix output of the data acquisition and preprocessing module, used to utilize the adjacency matrix. As a mask constraint, it aggregates the physical complementary features between adjacent stations. When the system is in operation or in use, the following steps are included: Step 1: Obtain the multi-source historical power sequences of each power station in the hydro-wind-solar hybrid system. And define an adjacency matrix based on the spatial topology of each station. ; Step 2: Perform time-domain decomposition on the multi-source historical power series to obtain trend-stable components and seasonal fluctuation components, which are then input into the trend channel and seasonal channel of the dual-channel spatiotemporal Transformer model, respectively. Step 3: Perform spatiotemporal attention calculations on the input trend stability component and seasonal fluctuation component in the trend channel and seasonal channel, respectively. This includes the following sub-steps: Step 3-1: Extract temporal features using a variable patch temporal attention mechanism and calculate the variable patch temporal attention output; Step 3-2: Introduce a spatial self-attention mechanism and utilize the adjacency matrix defined in Step 1. Perform a sparsity masking operation to obtain a spatial self-attention output; Step 3-3: Stack multiple spatiotemporal attention modules and achieve multi-scale spatiotemporal feature fusion through cross-layer skip connections to obtain the final spatiotemporal feature output; Step 4: Perform dual-channel prediction and fusion generation; including: Step 4-1: Use a perceptual network to map the spatiotemporal feature output to the power prediction value, and obtain the trend component prediction value of the trend channel and the seasonal component prediction value of the seasonal channel respectively. Step 4-2: Linearly superimpose the predicted values of the trend component and the seasonal component to obtain the final power joint prediction result of the hydro-wind-solar hybrid system; this result is used to guide the multi-energy complementary joint dispatch of the power grid and the formulation of day-ahead power generation plans. In step 2, the multi-source historical power sequences are first processed. Perform time-domain decomposition, breaking it down into trend-stable components. With seasonal fluctuation components ,in, For the length of the input data, The feature dimension is used; the trend-stabilizing component is obtained through a moving average calculation. ; Next, the trend component is subtracted from the original sequence in the form of a difference to extract the seasonal fluctuation component. The calculation formula is as follows: ; In the formula, This indicates a moving average operation. This indicates the filling operation. Through this decomposition mechanism, the model constructs trend channels and seasonal channels respectively, enabling the model to focus on dynamic patterns at two different time scales during learning: long-term stable trends and short-term random fluctuations. In step 3-1, the formula for calculating the variable patch temporal attention output is: ; In the formula, Indicates the first The first layer The first variable A variable patch-time attention output For variable-patch-time attention operations, express Activation function , Indicates the first The first layer A variable patch-time attention learnable temporal projection matrix for each variable; In step 3-2, the formula for calculating the spatial self-attention output is: ; In the formula, Indicates the first The first layer of topological variables Spatial self-attention output between patches For attention operations in the variable patch space, Indicates the first All layers The first variable A fake timestamp , , Indicates the first The first layer A spatial projection matrix that can be learned through spatial self-attention; Represents the Hadamard product of matrices; In step 3-3, the final spatiotemporal feature output The calculation formula is: ; in, For the first The output sequence of the layer after spatiotemporal attention processing. Indicates the first The learnable parameters of the layer skip connection, the summation formula above This represents the total number of stacked layers of the spatiotemporal attention modules; In step 4, the calculation formula for prediction and fusion generation is: Model prediction output for each channel (the channel being either a trend channel or a seasonal channel) for: ; in, Represents the activation function of the rectified linear unit. To predict future sequence length, and Indicates learnable parameters; Based on the aforementioned dual-channel integration mechanism, the model uses trend-stabilized components respectively. and seasonal fluctuation components Using the input as input, the predicted values of the trend components are calculated through two parallel ST-Former branches. and seasonal component forecast values Finally, the outputs of the two branches are linearly superimposed to obtain the final joint power prediction result of the hydro-wind-solar hybrid system. The calculation formula is: ; The technical advantages of the proposed water, wind, and solar joint prediction system and method based on a dual-channel spatiotemporal Transformer with seasonal trend decomposition are as follows: 1) Decoupling and Interference Prevention of Heterogeneous Temporal Features: An innovative dual-channel decomposition architecture is proposed, enabling the model to learn the long-period stable characteristics of hydropower output and system base load evolution (trend component) separately from the high-frequency stochastic characteristics of wind power and photovoltaic fluctuations (seasonal component). This effectively avoids feature overlap and mutual interference of multi-source heterogeneous data from hydropower, wind power, and solar power at different time scales, significantly improving the model's joint prediction stability in the face of extreme weather changes. 2) Spatial Complementary Topology Enables High Precision: Adjacency matrices representing the actual physical connections and geographical distances between hydropower stations, wind farms, and photovoltaic power stations (…). This is directly embedded in the mask calculation of spatial self-attention. This guides the model to follow the physical complementary correlations across energy types (such as meteorological linkage compensation between adjacent wind and solar power stations and hydraulic coupling constraints between upstream and downstream hydropower stations) when extracting spatial features, thereby significantly improving the accuracy of overall power output prediction of the hydro-wind-solar complementary system and the physical interpretability of real power grid dispatch.
[0035] Example: This invention selects measured data from a water-wind-solar complementary base in China from January 1, 2018 to December 31, 2018 to construct an original water-wind-solar multi-source sequence spatiotemporal dataset. The data has a time resolution of 15 minutes and contains only power features. The input historical sequence length H=384 and the predicted future sequence length F=384 are set, and the dataset is strictly divided into training set, validation set and test set in a ratio of 7:2:1.
[0036] To comprehensively evaluate the performance advantages of the proposed method in spatiotemporal joint prediction of multiple water, wind, and solar sources, this invention selects STGCN, Informer, GCN-GRU, BiLSTM, and Autoformer as comparison models. All models use the same dataset and input features, and the input length and prediction step size are consistent across models. For the model proposed in this invention, the loss function is constructed jointly using mean squared error and a two-layer complementarity regularization term, while the comparison models use the standard loss function without the complementarity regularization term. The models in the examples were all built in a Python 3.10 environment, and the experimental hardware configuration was a 2.40GHz Intel® Core™ / i7-13620H / 16GB memory. Model parameter settings are shown in Table 1.
[0037] Table 1 Model Parameter Settings
[0038] During the model training and evaluation phases, the features from both channels are aggregated and output. A composite loss function, incorporating global complementary constraints at the system level and local complementary constraints at the topology level, is introduced to calculate the total error. The backpropagation algorithm is used to continuously update the model parameters over 100 iterations to guide the model in learning the physical complementarity mechanism of multi-source power output. Simultaneously, a validation set is used to monitor model performance in real time during training, and the optimal model parameters are selected based on the validation error to avoid overfitting. Finally, the test set is input into the trained model, the output prediction data is aggregated, and inverse normalization is performed to generate the final power prediction result and evaluate its accuracy. The specific implementation schemes for the ablation experiments are divided into three types: 1) STD-ST-Former-W: Directly uses the original power sequence as input, removes the wavelet soft thresholding denoising module, and is used to verify the impact of data preprocessing on model performance.
[0039] 2) STD-ST-Former-S: Independent STD-ST-Former models are established for water, wind and solar energy to perform univariate predictions, without considering multi-source joint modeling and spatiotemporal coordination, in order to verify the role of the joint prediction mechanism.
[0040] 3) STD-ST-Former-L: Removes the two-layer complementary regularization term from the loss function and retains only the basic MSE loss function for training, in order to verify the contribution of the regularization term to improving prediction accuracy.
[0041] The above experimental design was evaluated using the following metrics: (27); (28); (29); In the formula: This represents the number of test samples; for Actual value at any time for Predicted value at any time; The mean of the predicted sequence; This is the mean of the actual sequence. and These represent the standard deviations of the actual and predicted values, respectively. This refers to the rated capacity of the corresponding energy source.
[0042] Table 2 shows the prediction errors of each model: Table 2 Prediction Errors of Each Model
[0043] The statistical results of the error indicators are shown in Table 2. Based on the error data in Table 2, in the hydropower prediction scenario, the method of this invention significantly outperforms the comparative model, with a reduction in rRMSE of 28.2%–56.3%, a reduction in rMAE of 30.1%–59.6%, and an increase in the correlation coefficient of 1.0%. For the more stochastic wind power scenario, the performance advantage of the method of this invention is further demonstrated. Compared with the comparative model, its rRMSE and rMAE are reduced by 22.8%–49.6% and 22.7%–51.7%, respectively, and the correlation coefficient is increased by 1.4%. In scenarios with significant periodic characteristics… In photovoltaic forecasting, the reductions in rRMSE and rMAE were 22.3%–49.4% and 28.6%–57.7%, respectively, with a 1.2% increase in the correlation coefficient, effectively correcting the prediction bias at peak times. In the total power forecasting of multi-source superposition, the method of this invention effectively suppressed the accumulation of errors in a single component through a double-layer complementary regularization term, resulting in a relative reduction in rRMSE and rMAE of 22.0%–50.3% and 28.6%–57.2%, respectively, with a 0.9% increase in the correlation coefficient, further validating its effectiveness in joint forecasting. Overall, the method of this invention achieved significant error reduction in all four scenarios, and the corresponding correlation coefficients achieved the highest values in each scenario, with an improvement of 0.9%–1.4%.
[0044] Table 3 shows the prediction error of the ablation experiment: Table 3. Prediction Errors in Ablation Experiments
[0045] Based on the ablation experiment data in Table 2, compared to variant models lacking specific key components, the complete method employed in this invention achieves comprehensive performance superiority across all energy types and total power predictions. Compared to STD-ST-Former-W, the method of this invention reduces rRMSE by an average of 26.9%–36.1%, rMAE by 27.2%–38.1%, and R² by approximately 0.3%–1.1% across the four prediction scenarios, demonstrating the effectiveness of denoising. Compared to the independent prediction model STD-ST-Former-S, the method of this invention reduces rRMSE by 31.3%–41.6% across the four scenarios, rMAE by 30.3%–43.4%, and correlation coefficients by an average of 0.6%–1.3%, indicating that the joint prediction mechanism helps improve prediction accuracy under multi-source coupling. Compared to STD-ST-Former-L… In comparison, the method of this invention reduces rRMSE and rMAE by 42.7%–46.3% and 36.5%–49.2% respectively in the four scenarios, with an average increase of 1.5% in the correlation coefficient, indicating that complementarity constraints further reduce prediction errors while improving global correlation. In conclusion, only by organically integrating adaptive wavelet denoising, system-level global complementarity constraints, and topology-level local complementarity constraints can the synergistic effect be maximized, thereby achieving the highest accuracy prediction for hydro-wind-solar hybrid systems.
[0046] Depend on Figures 3-6The prediction curves for each scenario reveal significant differences in the dynamic response and trend fitting of different models for multi-source power output. In the hydropower scenario, while Informer and Autoformer can roughly follow the periodic changes, they exhibit obvious oscillations and instability at the peaks, making it difficult to accurately reproduce the high-frequency details of actual power. BiLSTM and GCN-GRU show a tendency for over-smoothing, are not sensitive enough to power abrupt changes, and exhibit significant phase lag during uphill and downhill phases, leading to decreased prediction accuracy in the non-steady-state region. In the wind power scenario, due to the randomness of wind speed, all benchmark models generally perform poorly in this scenario. Informer and BiLSTM show severe amplitude decay, with the overall prediction curve significantly lower than the actual value, failing to effectively capture the power output characteristics during peak periods. Although Autoformer retains... While some trend information is presented, the predicted fluctuations in troughs and transitional phases are chaotic, exhibiting occasional reverse errors. GCN-GRU is weak in capturing small local fluctuations, resulting in a flattened overall curve that loses the fluctuating characteristics of wind power output. In photovoltaic scenarios, Informer and BiLSTM generally predict values lower than the actual values at peak sunshine hours, failing to fully capture output during periods of strongest sunlight. In contrast, STGCN and GCN-GRU show predicted peak values slightly higher than actual values in some periods, and their convergence speed is slower than the actual process during sunset, leading to error accumulation at edge moments and decreased prediction accuracy. The total power output incorporates complex spatiotemporal correlations from multiple sources. Due to the accumulation of errors in the aforementioned single scenarios, the overall trend of benchmark models such as Informer, BiLSTM, and GCN-GRU is significantly lower than the actual load curve, especially in the peak range, where the overall prediction curve exhibits a systematic deviation, making it difficult to cope with synchronous changes in multi-source output. In contrast, the method of this invention can accurately capture the changes in peaks and troughs in various scenarios, achieving more precise prediction output in complex spatiotemporal coupling contexts.
[0047] Depend on Figure 7 The prediction accuracy results show that in hydropower prediction scenarios, the method of this invention exhibits superior performance, achieving an accuracy of 97.4%. In wind power prediction scenarios, the performance of the comparative model shows a significant decline, while the method of this invention, thanks to its two-layer complementary regularization term, maintains an accuracy of 96.9%, achieving an improvement of 2.4% to 6.7%. In photovoltaic prediction scenarios, the method of this invention also performs excellently, with an accuracy as high as 97.9%. Even in the prediction of total power from multiple sources, the model can effectively suppress error accumulation, reaching an accuracy of 97.7%, an improvement of 0.9% to 4.9% compared to the comparative model. Overall, the results indicate that the method of this invention can maintain stable prediction performance in different scenarios, with an accuracy exceeding 96.9% in all scenarios, significantly higher than the comparative model, and exhibiting the smallest fluctuation range.
Claims
1. A method for combined prediction of wind and solar power based on double-layer complementary regularization, characterized in that, Includes the following steps: Step 1: Obtain multi-source historical power data of the hydro-wind-solar hybrid system, and preprocess the multi-source historical power data to obtain a smoothed signal after noise reduction; Step 2: Construct a seasonal trend decomposition dual-channel spatiotemporal Transformer model, which includes a trend channel and a seasonal channel; Step 3: Input the smoothed signal obtained in Step 1 into the trend channel to extract low-frequency evolution features, and input the original sequence of multi-source historical power data into the seasonal channel to extract periodic patterns and short-term fluctuation features; Step 4: Construct a loss function that includes a two-layer complementarity regularization term; the two-layer complementarity regularization term includes a system-level regularization term and a topology-level regularization term; Step 5: Iteratively train the seasonal trend decomposition dual-channel spatiotemporal Transformer model using the loss function, and use the trained model to generate the joint power prediction results of the water-wind-solar complementary system.
2. The method according to claim 1, characterized in that, In step 1, the wavelet soft thresholding denoising algorithm is used to preprocess the multi-source historical power data to obtain the denoised smooth signal. The smooth signal is divided into training set, validation set and test set according to a preset ratio, which are used for subsequent model training, hyperparameter selection and performance evaluation, respectively. The preprocessing includes the following sub-steps: Step 1-1: Use Mallat's multi-scale decomposition algorithm to process the original power output sequence of the hydro-wind-solar hybrid system. Perform discrete wavelet transform to decompose it to the th Layers were used to obtain low-frequency scale coefficients that reflect long-term evolutionary trends. and high-frequency wavelet coefficients containing local perturbations and noise. ; Steps 1-2: Based on Stein's unbiased risk estimation theory, calculate the... Adaptive threshold for layer decomposition scale ; Steps 1-3: Using the soft thresholding function to evaluate the high-frequency wavelet coefficients Nonlinear shrinkage is performed to obtain the denoised high-frequency coefficients. ; When the coefficient magnitude is less than the threshold The value is set directly to zero when the coefficient amplitude is greater than or equal to the threshold. At that time, its amplitude is reduced by a threshold amount while keeping the sign unchanged; Steps 1-4: Utilizing unsuppressed low-frequency scaling coefficients and the high-frequency coefficients after soft thresholding. Performing inverse wavelet transform generates a denoised time series that retains the true evolutionary trend and eliminates high-frequency random disturbances. .
3. The method according to claim 2, characterized in that, In step 1-1, firstly, the multi-source power output signal from water, wind, and solar is considered as a noisy time series, and its true value is expressed as: (1); In the formula, The real value of contributing to the water scenery An ideal output signal with no noise. To follow a standard normal distribution Gaussian white noise, Noise intensity; The non-stationary sequence is decomposed into different time-frequency scales using Discrete Wavelet Transform (DWT), and scale parameters are defined. With translation parameters for: (2); And with mother wavelet Generating subwavelets: (3); Then the signal is at scale ,Location The formula for calculating the wavelet coefficients at a given location is: (4); Using Mallat's multi-scale decomposition algorithm, a signal can be decomposed into a low-frequency main trend term and a high-frequency detail term: (5); In the formula, For scaling function, For the first The low-frequency scaling coefficients of the layer decomposition reflect the long-term evolution trend of water, wind, and solar power. These are high-frequency wavelet coefficients, which mainly contain local disturbances and noise.
4. The method according to claim 3, characterized in that, In steps 1-2, to avoid over-smoothing or under-denoising due to a globally fixed threshold, a scale-adaptive thresholding scheme is adopted based on Stein's unbiased risk estimation theory; for the th Layer decomposition scale, with its adaptive threshold as follows: (6); In the formula, The length of the coefficient at this scale. This is the standard deviation of the detail coefficients at that scale, used to characterize the noise level at that scale; This is a scale correction factor used to adjust the threshold strength at this scale. It is obtained by minimizing the unbiased risk estimate of the denoised reconstructed signal at this scale, i.e.: (7); In the formula, In the correction factor and their corresponding thresholds The unbiased risk estimate corresponding to soft threshold compression under the action; through the above formula, the noise suppression intensity at different scales can be adaptively adjusted, instead of using a single global fixed threshold; In steps 1-3, a soft threshold function is introduced for the first... High-frequency wavelet coefficients of the layer Nonlinear shrinkage processing is performed to remove noise and retain effective information, resulting in processed coefficients. The calculation formula is as follows: (8); in, The threshold obtained in step 1-2 For sign functions; when When the coefficient is zero, it is considered the dominant noise component and is set to zero directly; when At that time, the coefficient is according to Amplitude reduction is performed, and the amplitude is reduced to a threshold value while keeping the sign unchanged. Instead of using the original amplitude for reconstruction, the amplitude is reduced; this operation suppresses high-frequency noise components while ensuring continuity at the threshold by shrinking the high amplitude coefficient, making the reconstructed signal smoother and more natural.
5. The method according to claim 4, characterized in that, After soft thresholding compression is completed in steps 1-4, the processed detail coefficients are used. and unsuppressed scaling factor Perform inverse wavelet transform to obtain the denoised time series. : (9); Reconstructed sequence While accurately depicting the real evolution trend and dynamic characteristics of water, wind, and solar power output, it effectively suppresses high-frequency random disturbances.
6. The method according to any one of claims 1 to 5, characterized in that, In step 2, a seasonal trend decomposition dual-channel spatiotemporal Transformer model, namely the STD-ST-Former network, is constructed, which specifically includes the following sub-steps: Step 2-1: Construct the spatial topology of the hydro-wind-solar hybrid system; Step 2-2: Define the spatial adjacency matrix; Steps 2-3: Build the overall architecture of the dual-channel model; Steps 2-4: Initialize the ST-Former network layer structure and parameters; In step 2-1, a spatial information model for joint prediction of water, wind, and light is established using graph theory methods, and an undirected graph is defined. ;in, This represents the set of all power station nodes in a hydro-wind-solar hybrid system. This represents the total number of nodes. The set of edges represents the physical or electrical connections between nodes. With nodes If there is an actual electrical interconnection, then in If there is a connecting edge, then there is no connecting edge. In step 2-2, based on the spatial topology map Define the adjacency matrix This is used to quantify the spatial topological relationships between nodes; the elements in the matrix represent the connectivity of edges, and if there is a connection between nodes, the corresponding element is non-zero, which is used for subsequent sparsity masking operations in the spatial attention mechanism. In steps 2-3, a seasonal trend decomposition dual-channel spatiotemporal Transformer model is constructed, configuring two parallel processing channels: Trend channel: used to receive the decomposed trend-stable components, and learn the long-term stationary characteristics of the power sequence through the configured ST-Former network branches; Seasonal channel: used to receive the decomposed seasonal fluctuation components and learn the short-term dynamic characteristics of the power sequence through the configured ST-Former network branch; The output of the seasonal trend decomposition dual-channel spatiotemporal Transformer model is equipped with a fusion module, which is used to linearly superimpose the prediction results of the trend channel and the seasonal channel. The adjacency matrix defined in step 2-2 It is integrated into the spatial self-attention unit to guide feature aggregation calculation; In steps 2-4, stacking is performed in the trend channel and the seasonal channel respectively. Layered spatiotemporal attention modules; each layer module contains variable patch temporal attention units and spatial self-attention units; network hyperparameters are set, including input data length. Feature Dimension Hidden layer dimension ; Regarding the first Layered network, defining the patch partitioning parameters: input sequence length and Patch step size The input sequence is divided into The first patch; initializes the learnable parameters of the model, including: the first patch. The first layer The first variable A fake timestamp Temporal attention projection matrix , Spatial self-attention projection matrix , , ; and cross-layer skip connection parameters and output layer perceptron parameters and .
7. The method according to claim 6, characterized in that, In step 3, the processed data is input into the seasonal trend decomposition dual-channel spatiotemporal Transformer model to extract features and generate prediction results. This includes the following sub-steps: Step 3-1: Perform time domain decomposition; Step 3-2: Calculate variable patch-time attention; Step 3-3: Calculate self-attention in the sparse space; Steps 3-4: Multi-scale spatiotemporal feature fusion; Steps 3-5: Dual-channel prediction generation and fusion.
8. The method according to claim 7, characterized in that, In step 3-1, the multi-source historical power sequence is first processed. Perform time-domain decomposition, breaking it down into trend-stable components. With seasonal fluctuation components ,in, For the length of the input data, The feature dimension is used; the trend-stabilizing component is obtained through a moving average calculation. (10); Next, the trend component is subtracted from the original sequence in the form of a difference to extract the seasonal fluctuation component. The calculation formula is as follows: (11); In the formula, This indicates a moving average operation. This indicates the filling operation; through this decomposition mechanism, the model constructs trend channels and seasonal channels respectively, enabling the model to focus on dynamic patterns of two different time scales, namely long-term stable trends and short-term random fluctuations, during learning. In step 3-2, the temporal feature extraction stage, a variable patch temporal attention mechanism is used to adapt to input sequences of different lengths; the first... Layer length is The input sequence is divided along the time dimension into There are 1 patch, of which... Indicates the first Layer patch step size; let the first layer be... The node at the th The input of each patch is ,in, express The Middle One input, , Define a learnable pseudo-timestamp , Indicates the first The first layer The first variable A pseudo-timestamp is used as a query matrix in the variable patch temporal attention to query all input values in the patch. Represents the hidden dimension; the variable patch-time attention output is: (12); In the formula, Indicates the first The first layer The first variable A variable patch-time attention output For variable-patch-time attention operations, express Activation function , Indicates the first The first layer A variable patch-time attention learnable temporal projection matrix for each variable; In step 3-3, a spatial self-attention mechanism is introduced to mine all elements in each layer. The first variable The correlation between patches is used, and the adjacency matrix defined in step 2 is utilized. A sparsity masking operation is performed to preserve the main physical correlations, and the spatial self-attention output is: (13); In the formula, Indicates the first The first layer of topological variables Spatial self-attention output between patches For attention operations in the variable patch space, Indicates the first All layers The first variable A fake timestamp , , Indicates the first The first layer A spatial projection matrix that can be learned through spatial self-attention; Represents the Hadamard product of matrices; In steps 3-4, to simultaneously capture spatiotemporal features at different scales, the models are stacked. A spatiotemporal attention module is used, and multi-scale fusion is achieved through cross-layer skip connections, resulting in the final spatiotemporal feature output. for: (14); in, For the first The output sequence of the layer after spatiotemporal attention processing. Indicates the first Learnable parameters of layer skip connections, ; In steps 3-5, a two-layer perceptual network is used to map spatiotemporal features to power prediction values; Model prediction output for each channel The calculation formula is: (15); Indicates the predicted output. Represents the activation function of the rectified linear unit. To predict future sequence length, and Indicates learnable parameters; Based on the aforementioned dual-channel integration mechanism, the model uses trend-stabilized components respectively. and seasonal fluctuation components Using the input as input, the predicted values of the trend components are calculated through two parallel ST-Former branches. and seasonal component forecast values Finally, the outputs of the two branches are linearly superimposed to obtain the final joint power prediction result of the hydro-wind-solar hybrid system. : (16)。 9. The method according to claim 1, 2, 3, 4, 5, 7, or 8, characterized in that, In step 4, a loss function containing a two-layer complementarity regularization term is constructed. This two-layer complementarity regularization term includes a system-level regularization term and a topology-level regularization term. The system-level regularization term constructs a global complementarity index based on a normalized weighted aggregation sequence, used to characterize multi-source complementary coupling features. The topology-level regularization term targets electrically adjacent nodes, combining cosine similarity and hinge functions to characterize spatial correlation, and is adjusted using gating weights. The construction of the loss function containing the two-layer complementarity regularization term specifically includes the following sub-steps: Step 4-1: Define the general formula for the composite loss function; Step 4-2: Construct a global complementarity regularization term at the system level; Step 4-3: Construct the local complementarity regularization term for the topology layer; In step 4-1, a composite loss function with two-layer complementary regularization is constructed. This function introduces a system-level global complementary regularization term and a topology-level local complementary regularization term into the basic mean square error term. Its expression is: (17); in, This term represents the mean squared error between the model's predicted values and the actual values. and These are the global complementary regularization term at the system layer and the local complementary regularization term at the topology layer, respectively. These are the weighting coefficients; In step 4-2, in order to guide the model to learn the dynamic coupling law between heterogeneous energy sources, a system-level regularization term is constructed by referring to and expanding the complementarity evaluation index; First, based on the installed capacity of wind power, hydropower, and photovoltaic power. , , Define capacity ratio coefficient , , : (18); Next, a weighted power sequence of the three energy sources and the benchmark complementary pair is constructed to reflect the energy substitution relationship between different energy sources: (19); in, , , They are time points The output of wind power, hydropower, and photovoltaic power; , and These are the weighted power values of wind power, hydropower, and photovoltaic power complements to the benchmark, respectively. Furthermore, for any time step pair in the predicted sequence , Calculate the change in power And define the temporal complementarity coefficient as: (20); in The total duration of the sample. For time-series complementarity coefficients, , , This is the capacity reduction factor, used to correct for the impact of differences in installed capacity on complementarity calculations; Finally, the inverse of the temporal complementarity coefficient is defined as the system-level global complementarity regularization term: (21); When the model predicts a more significant complementary relationship between water, wind, and solar power outputs over time, This reduces the frequency of parameter updates, thus guiding the direction of parameter updates during backpropagation to better align with the energy distribution patterns of multi-energy systems. In step 4-3, in order to prevent spatial feature aliasing and reflect regional complementarity, a topology layer regularization term is constructed based on the correlation of node prediction results. First, compute nodes , The predicted output is , The cosine similarity after removing the mean is: (22); in, Represents a node and The cosine similarity of the predicted output sequence after removing the mean. , The average value of the corresponding sequences; the local complementarity constraint function is defined based on similarity as follows: (23); In the formula, For nodes and Local complementary constraint values between them This is a similarity threshold used to define the maximum permissible correlation between prediction results from different power plants; Combined with adjacency matrix In addition to the historical correlation adjustment coefficient, adaptive gating weights are defined: (24); in, For nodes and Adaptive gating weights between them Derived from the adjacency matrix , representing the spatial connection strength between nodes; These are learnable gating parameters used to dynamically adjust the regularization strength. To prevent the stability constant from having a denominator of zero; The final topology layer complementarity regularization term is: (25); in It represents the set of adjacent nodes across energy types; by introducing a spatially complementary structure, this term enables the model to enhance its ability to identify the output characteristics of different power plants at the feature representation level during the training process.
10. The method according to claim 7, characterized in that, Step 5 involves iteratively training the model using the loss function and generating prediction results, specifically including the following sub-steps: Step 5-1: Iterative model training and parameter optimization; Step 5-2: Training terminated and model saved; Step 5-3: Perform offline prediction; Step 5-4: Inverse normalization reconstruction of the prediction results; In step 5-1, the pre-defined training set is input into the seasonal trend decomposition dual-channel spatiotemporal Transformer model; in each iteration, forward propagation is performed, and the current prediction output is generated according to the formula in step 3. ; The composite loss function constructed in step S4 Calculate the total error between the current predicted value and the actual value; The loss function is calculated using the backpropagation algorithm. The gradients of all learnable parameters of the model; these learnable parameters include the temporal attention projection matrix. , Spatial self-attention projection matrix , , ; and cross-layer skip connection parameters and output layer perceptron parameters and ; The gradient descent optimization algorithm is used to update the above parameters based on the calculated gradient in order to minimize the composite loss function. ; During the parameter update process, the system-level global complementary regularization term The topology layer local complementary regularization term guides the model parameters to converge in a direction that conforms to the energy distribution law of multiple energy sources. Suppress spatial feature aliasing, thereby co-optimizing the model's spatiotemporal expressive power during training; In step 5-2, step 5-1 is repeated until a preset convergence condition is met; the convergence condition includes reaching a preset maximum number of iterations, or the loss function value on the validation set no longer decreasing within a certain number of consecutive cycles; After training, the optimal model parameters are saved to obtain the trained power joint prediction model of the water-wind-solar hybrid system. In step 5-3, the multi-source historical power data of the water-wind-solar hybrid system for the period to be predicted is obtained, and preprocessed according to the wavelet soft thresholding denoising method described in step 1 and the normalization method described in steps 1-3 to generate the model input sequence. The input sequence is fed into the trained model, and through parallel processing and fusion of the trend channel and seasonal channel, a normalized future power prediction sequence is output. ; In step 5-4, the normalized prediction sequence is... Perform an inverse normalization operation to map it back to the original power dimension. The calculation formula is as follows: (26); in, For the final power joint prediction results of the hydro-wind-solar hybrid system, and These are the maximum and minimum values of the historical training data recorded in steps 1-3, respectively. The final output includes the future. The predicted power values for hydropower, wind power, and photovoltaic power at each time step, as well as the predicted total power value of the complementary system.