Offshore wind power prediction method based on pre-training time sequence large model

By constructing a two-stage adaptive prediction model for typhoons, combining channel attention enhancement and dynamic time series fusion modules, and adopting a two-stage hybrid optimization training strategy, the problems of high computational cost and difficulty in adaptation of pre-trained time series basic models under extreme weather conditions are solved, achieving efficient and accurate wind power prediction.

CN121787664APending Publication Date: 2026-04-03SICHUAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-30
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Existing pre-trained time series models suffer from high computational overhead and difficulty in black-box adaptation in wind power prediction under extreme weather conditions, and it is difficult to effectively utilize their generalization ability. Traditional methods have insufficient prediction accuracy under extreme weather conditions.

Method used

A two-stage adaptive prediction model for typhoons is constructed, employing a channel attention enhancement module and a dynamic temporal fusion module. By stitching the projection layer and the output projection layer, multi-source features are efficiently coupled with the temporal base model based on lag features. A two-stage hybrid optimization training strategy is adopted, including zero-order optimization and first-order optimization, to extract features and fine-tune the model for extreme weather.

Benefits of technology

It improves the accuracy and computational efficiency of wind power prediction, achieves zero-sample cross-scenario transfer, and significantly reduces errors, especially under extreme weather conditions, adapting to wind power prediction for wind farms with different terrains.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121787664A_ABST
    Figure CN121787664A_ABST
Patent Text Reader

Abstract

The invention discloses an offshore wind power prediction method based on a pre-training time sequence large model, and relates to the technical field of new energy power generation. Constructing a TTPAPM model, and constructing a feature extraction framework composed of a channel attention enhancement module and a dynamic time sequence fusion module by taking the pre-training time sequence basic model as a core; the multi-source features and the time sequence basic model based on the lag features are efficiently coupled through the splicing projection layer and the output projection layer. The training strategy adopts two-stage hybrid optimization: in the first stage, zero-order optimization is performed on an input side on conventional weather data, and first-order optimization is performed on an output side; and in the second stage, only the DTFM module and the output module are finely adjusted in terms of typhoon data. The TTPAPM model achieves consistency improvement on two indexes of mean absolute error and root-mean-square error. Meanwhile, excellent generalization performance is shown in a zero sample scene, and memory occupation in the model training process is remarkably reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of new energy power generation technology, and in particular to a method for predicting offshore wind power based on a pre-trained time series large model. Background Technology

[0002] In the global energy structure transformation, wind power has become a core support. Accurate wind power forecasting is crucial for grid dispatch, system stability, and market transactions. Especially under extreme weather conditions such as typhoons, strong fluctuations and non-stationarity of wind speed can lead to drastic power surges, threatening grid security. Therefore, accurate forecasting under extreme weather conditions is particularly important.

[0003] Traditional forecasting methods mainly include statistical models, machine learning, and deep learning. Statistical models (such as ARIMA) struggle to handle nonlinearity and abrupt changes; machine learning methods (such as SVM and random forest) are insufficient for modeling long-term sequence dependencies; and while deep learning methods (such as LSTM) can uncover complex temporal patterns, their performance degrades significantly under extreme weather conditions, and they are computationally expensive. To address the problems of limited sample size and poor data quality under extreme weather conditions, existing research has explored various approaches. For example, data completion, constructing derived features, or employing oversampling strategies can improve the insufficient sample size; or generative models can be used for data augmentation to expand the training set. However, these traditional augmentation methods often fail to preserve the original physical characteristics and temporal dependencies of the data, and the quality of samples generated from biased numerical weather prediction data is limited, thus restricting the improvement of forecast accuracy.

[0004] Transfer learning reduces reliance on target domain data through knowledge transfer, but it is prone to "negative transfer" due to inter-domain differences, and usually still requires a small number of target samples for fine-tuning, making it difficult to achieve true zero-shot cross-scenario applications. In recent years, the emergence of pre-trained time series foundation models has provided a new approach to solving the above problems. These models, after pre-training on large-scale time series data, have shown strong cross-domain generalization and zero-shot inference potential. However, directly applying them to extreme weather wind power prediction still faces significant challenges: on the one hand, the model parameters are huge, resulting in high computational costs, which is not conducive to real-time deployment; on the other hand, power sequences under extreme weather conditions are highly non-stationary and noisy, requiring targeted optimization of the model. However, existing foundation models are mostly black-box or closed-source structures, and their weights and training details are unavailable, hindering effective domain adaptation and performance improvement.

[0005] Therefore, there is an urgent need for a new wind power prediction method that can efficiently utilize the generalization ability of pre-trained time series basic models, while overcoming their shortcomings such as heavy computational burden and difficulty in black box adaptation, and is specifically optimized for the non-stationary characteristics of extreme weather. Summary of the Invention

[0006] To address the aforementioned issues, this invention proposes a method for predicting offshore wind power based on a pre-trained time-series large model. The method uses the pre-trained time-series basic model as its core and constructs a feature extraction framework consisting of a channel attention enhancement module and a dynamic time-series fusion module. By stitching together a projection layer and an output projection layer, multi-source features are efficiently coupled with the time-series basic model based on lag features. A two-stage hybrid optimization training strategy is employed to resolve the aforementioned problems.

[0007] This application discloses a method for predicting offshore wind power based on a pre-trained time-series large model, including the following steps: S1. Acquire offshore wind power data and numerical weather prediction data; S2. Construct a two-stage adaptive prediction model for typhoons. The two-stage adaptive prediction model for typhoons includes a channel SE module, a dynamic temporal fusion module, a stitching projection layer, a pre-trained temporal base model, and an output mapping layer connected in sequence. The channel temporal fusion module includes a parallel temporal convolution module and a dynamic module decomposition module. S3. A two-stage hybrid optimization training strategy is adopted to train the two-stage adaptive prediction model for typhoons. In the first stage, the input module is trained on regular weather data, and in the second stage, the dynamic time series fusion module and the output mapping layer are fine-tuned on typhoon data. S4. Input the data obtained in S1 into the trained two-stage adaptive prediction model of typhoon obtained in S3 to obtain the predicted wind power value.

[0008] Preferably, the channel SE module employs a multi-scale dynamic channel attention enhancement module, and the implementation process is as follows: The multi-scale dynamic channel attention enhancement module transforms the input feature map into a spatiotemporal representation through multi-scale convolutional branches. The multi-scale convolutional branches include parallel temporal branches and simulated periodic branches. The temporal branches capture pure temporal dependencies, while the simulated periodic branches reshape the sequence into a two-dimensional image pattern. The time-series branch output and the simulation periodic branch output are dynamically compressed, and the weighted global description vector is calculated by combining the typhoon sign as a conditional input. Channel recalibration is performed, and the two-layer fully connected network is extended into a residual cascade structure to ensure stable gradient propagation. The calculation process of the multi-scale dynamic channel attention enhancement module is as follows: (1) (2) (3) in, For time step, For the number of channels, Indicates the first One channel, For the first Global characteristics of each channel For learnable physical constraint parameters, and Here are the weight matrices for the two fully connected layers. It is the sigmoid activation function. express Activation function Indicates a time index. Indicates the first The time step, the first The original input feature values ​​in each feature channel For the first The typhoon status indicator variables corresponding to each time step This is a channel-level global feature vector. Here is the channel attention weight vector. For the channel attention weight vector corresponding to the first The weighting coefficients of each channel, The output feature value is modulated by the multi-scale dynamic channel attention mechanism.

[0009] Preferably, the implementation process of the temporal convolution module is as follows: The weighted features output by the channel SE module Divide into blocks, cut into A fixed-length overlapping block ; Each overlapping block Projected into a higher-dimensional space via a linear compression-mapping structure Generate local timing input ; For local timing input A bidirectional, multi-scale causal temporal convolutional structure is performed, followed by weighted modeling of the bidirectional convolutional output, feature reshaping, and nonlinear enhancement. Additionally, a typhoon identifier variable is introduced, and learnable coefficients are used to perform the modeling. The attention output is modulated by disaster intensity, and finally the convolution enhancement result is added to the original local input through residual connection to obtain the output of the temporal convolution module; The calculation process of the temporal convolution module is as follows: (4) (5) (6) in, For weighted features The first one obtained by dividing in the time dimension A locally overlapping time block, For local timing input Intermediate feature representation after applying bidirectional multi-scale dilated causal convolution. The final output features of the temporal convolution module are... The weights of the compression layer, For the bias of the compression layer, The weights of the mapping layer, For the bias of the mapping layer, For the set of dilation rates of dilated causal convolutions, This represents a three-layer dilated causal convolutional group that operates in parallel forward and backward directions. This is a multi-head self-attention mechanism. This is a one-dimensional convolution operation. The learnable typhoon enhancement coefficient, This is an indicator variable for typhoon conditions.

[0010] Preferably, the implementation process of the dynamic module decomposition module is as follows: Weighted features of the SE module output of the channel Application layer normalization is obtained by stabilizing the feature distribution. For the normalized feature map, a sliding window expansion operation is used to extract a feature map of length [length missing] over the time dimension. subsequence of: (7) in, To output the shape, The number of sliding windows. This indicates the operation of expanding the sliding window. Weighted features of the SE module output of the channel The feature matrix obtained after applying layer normalization, This represents the dimension index in which the sliding window expansion operation is performed. Indicates the length of the sliding window. This indicates the step size of the sliding window in the time dimension; The dimensions are rearranged using the permute operation to construct a multidimensional Hankel matrix: (8) in, For Hankel matrix, Indicates the permute operation; right Perform low-rank singular value decomposition and take the first-order singular value. Each singular value defines a delay coordinate. Construct the shifted data and then use the last coordinate. As an external force channel, the rest of the front Each is a system state. Using least squares estimation: (9) in, This represents the hidden state vector of the system at the next time step. Represents the system state transition matrix. Indicates the first The system's hidden state vector at each time step Represents the vector of external force coupling coefficients. Indicates the first The external force channel variables corresponding to each time step Indicates a discrete-time index; For the hidden state vector of the system With the corresponding external force channel variable Concatenate the vectors to generate dynamic modular decomposition embedding vectors. Before stitching and projection, the channels are aligned with those of the temporal convolution module, and a linear mapping is applied: (10) in, For dynamic module decomposition embedding, it represents the dynamic feature representation after linear mapping. Represents dynamic embedding vectors The transpose form, The linear mapping weight matrix represents the embedding of dynamic module decomposition into a unified feature space. This represents the bias term in the linear mapping process.

[0011] Preferably, the splicing projection layer is used to fuse the output of the temporal convolution module and the output of the dynamic module decomposition module and map them to the input dimension of the pre-trained temporal base model. The calculation process is as follows: (11) (12) in, Indicates will With The joint feature representation formed by concatenating features along the feature dimension. This represents the feature concatenation operation function. Indicates batch size, This represents the final fused feature representation output by the stitching projection layer. This indicates a random deactivation operation. Presentation layer normalization operation, This represents the linear mapping weight matrix of the spliced ​​projection layer. This indicates the bias term corresponding to the splicing projection layer.

[0012] Preferably, the pre-trained temporal base model adopts the Lag-Llama model, which is based on the Transformer architecture with only a decoder, uses lag features and time covariates as inputs, and obtains a general temporal representation through pre-training.

[0013] Preferably, the output mapping layer maps the hidden representation of the pre-trained temporal base model to the predicted wind power value, and the calculation process is as follows: (13) in, Indicates the first The sample, the first The model outputs predicted values ​​at each prediction time step. The model represents the first time. The sample, the first The high-dimensional hidden state vector corresponding to each time step This represents the linear mapping weight matrix of the output mapping layer. This represents the bias term of the output mapping layer.

[0014] Preferably, the two-stage hybrid optimization training strategy includes: In the first stage, a regular weather dataset is used as input. The zero-order optimization method is used to update the parameters of the channel SE module and the dynamic time series fusion module, and the first-order optimization method is used to update the output mapping layer. The second stage uses the typhoon dataset as input, employs a zero-order optimization method to update the parameters of the dynamic time series fusion module, and a first-order optimization method to update the output mapping layer.

[0015] Preferably, the zero-order optimization method is calculated using the following formula: (14) in, To estimate the gradient, For the model's parameter vector, For the number of queries, The random perturbation step size, For random perturbation direction, and loss function Loss values ​​in different perturbation directions.

[0016] right Update: (15) in, Update the step size for the parameters. Indicates the first The parameter vector of the model at the next iteration. Indicates the first In the next iteration, the corresponding parameters The loss function value.

[0017] Preferably, the first-order optimization method includes the following steps: Using the standard backpropagation mechanism, the gradients of the output mapping layer weights and biases are calculated respectively: (16) in, To output the gradient of the weight parameters of the mapping layer, To output the gradient of the bias parameters of the mapping layer, For the first The bias parameters of the output mapping layer are output before the next iteration; In the In this iteration, based on the gradient of the current loss function: (17) Calculate the first and second moments of the gradient with respect to the weight parameters and bias parameters, respectively. The first moment is used to characterize the exponentially weighted average of the gradient, and its calculation formula is as follows: (18) in, This is the first-order moment decay coefficient, used to control the degree to which historical gradient information is preserved; After correcting for the deviation of the first moment, we get: (19) The second moment is used to characterize the exponentially weighted average of the squared gradient, and its calculation formula is as follows: (20) in, The second-order moment attenuation coefficient, This represents the element-wise square of the gradient; After correcting for the deviation of the second moment, we get: (twenty one) Construct a standardized update direction vector: (twenty two) in, It is a numerically stable term; Introducing hierarchical trust ratio in weight updates: (twenty three) in, For the first Output the weight parameters of the mapping layer before the next iteration. This is a standardized update direction vector constructed based on the first and second moments; Based on this, complete the parameter update: (twenty four) (25) in, For the first The weight parameters of the mapping layer are output after the next iteration. To indicate the first The bias parameters of the output mapping layer are output after the next iteration. This represents the global learning rate.

[0018] The beneficial effects of this invention are: (1) Higher prediction accuracy: The TTPAPM proposed in this invention outperforms traditional time series models (such as Transformer, Informer, etc.) under both normal and extreme weather conditions, and the error is significantly reduced, especially in high-fluctuation scenarios such as typhoons.

[0019] (2) Strong zero-sample generalization ability: The TTPAPM proposed in this invention can be directly applied to wind farms with different terrains without retraining, and adapts to cross-scene migration.

[0020] (3) Better computational efficiency: The present invention adopts a two-stage training and modular fine-tuning strategy to reduce the need for complete fine-tuning of pre-trained large models, thereby reducing computational overhead and memory usage. Attached Figure Description

[0021] Figure 1 This is a framework diagram of a two-stage adaptive prediction model for typhoons according to an embodiment of the present invention. Figure 2 This is a schematic diagram of the channel SE module structure according to an embodiment of the present invention; Figure 3 This is a schematic diagram of the temporal convolution module structure according to an embodiment of the present invention; Figure 4 This is a schematic diagram of the dynamic module decomposition module structure according to an embodiment of the present invention; Figure 5 This is a schematic diagram of the splicing projection layer structure according to an embodiment of the present invention; Figure 6 This is a schematic diagram of the Lag-Llama model structure according to an embodiment of the present invention; Figure 7 This is a schematic diagram of the training process of the two-stage adaptive prediction model for typhoons according to an embodiment of the present invention. Figure 8 This is a schematic diagram of wind power prediction results for offshore wind farms under normal weather conditions, according to an embodiment of the present invention. Figure 9 This is a schematic diagram of wind power prediction results for offshore wind farms under typhoon weather, according to an embodiment of the present invention. Figure 10 This is a schematic diagram illustrating the correlation analysis between wind speed and wind power in offshore wind farms and hilly wind farms according to an embodiment of the present invention. Figure 11 This is a schematic diagram of zero-sample prediction results for routine weather in a hilly wind farm according to an embodiment of the present invention. Figure 12 This is a schematic diagram of zero-sample prediction results for a hilly wind farm during typhoon weather, according to an embodiment of the present invention. Figure 13 This is a schematic diagram of the results of a comparative experiment on conventional weather conditions at an offshore wind farm, as described in an embodiment of the present invention. Figure 14 This is a schematic diagram of the comparative experimental results of offshore wind farms in typhoon weather according to an embodiment of the present invention. Detailed Implementation

[0022] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided with reference to the accompanying drawings and embodiments.

[0023] This application discloses a method for predicting offshore wind power based on a pre-trained time-series large model, including the following steps: S1. Acquire offshore wind power data and numerical weather prediction data.

[0024] Based on the official dataset of the 5th "Goldwind Cup" Energy Innovation Challenge finals, and combined with a supplementary typhoon-specific dataset, validation work was carried out to obtain offshore wind power operation data and numerical weather prediction data. Specifically, the raw data was first parsed and sorted by time column, and complete observation records for two years were selected to ensure sufficient timeliness and representativeness. Subsequently, feature columns (including key meteorological variables such as wind speed, wind direction, temperature, and humidity) and target power columns were automatically identified. Missing values ​​were filled using time-based linear interpolation combined with forward and backward imputation strategies to ensure the continuity of the time series without introducing significant bias. In the typhoon-specific data, the typhoon_flag field was further annotated according to the provided typhoon impact period dictionary to distinguish between regular weather and typhoon weather samples. Normalization employs a dual-scaler strategy: regular weather samples are standardized using a global MinMaxScaler, while typhoon weather samples are normalized using a separate MinMaxScaler. This preserves the extreme numerical distribution characteristics during typhoon seasons while avoiding excessive smoothing of the normalization parameters for typhoon samples by regular samples. The target power is uniformly normalized using a single MinMaxScaler to ensure consistent physical meaning after denormalization. After normalization, hourly, daily, and monthly features are extracted as temporal covariates, and a supervised sequence incorporating historical power is constructed using a sliding window approach. The input sequence simultaneously includes historical meteorological features, historical power, and temporal features, and the prediction target is the normalized power value at multiple time steps.

[0025] S2. Construct a two-phase adaptive prediction model for typhoons (TTPAPM). The model framework is as follows: Figure 1 As shown, the two-stage adaptive prediction model for typhoons comprises a channel SE module (Squeeze-and-Excitation, SE), a dynamic temporal fusion module (DTFM), a stitching projection layer, a pre-trained temporal base model, and an output mapping layer, connected sequentially. In this embodiment, the pre-trained temporal base model adopts the Lag-based Large Language Model. The channel SE module and the dynamic temporal fusion module form a feature extraction framework to separate and characterize periodic and abrupt components. The stitching projection layer and the output mapping layer efficiently couple multi-source features with the pre-trained temporal base model based on lag features, achieving end-to-end mapping from features to power.

[0026] Wind power forecasting relies on various meteorological variables (such as wind speed, wind direction, and air pressure), but each variable contributes differently to power output. To optimize feature representation, the channel SE module structure in this embodiment is as follows: Figure 2 As shown, a Multi-Scale Dynamic Channel Attention Module (MSDCAM) is employed. Compared to the original channel SE module, MSDCAM is no longer limited to simple channel recalibration after global average pooling. Instead, it introduces a two-dimensional spatiotemporal cascaded attention mechanism and physically constrained adaptive weight learning, which significantly improves the sensitivity to non-stationary abrupt changes in meteorological variables such as wind speed and air pressure under extreme typhoon weather conditions, while reducing computational overhead.

[0027] Specifically, MSDCAM first processes the input feature map (with dimensions of 10 ... ,in For time step, The sequence (number of channels) is transformed into a spatiotemporal representation through multi-scale convolutional branches (1D and 2D in parallel): the 1D branch is a temporal branch used to capture pure temporal dependencies. The 2D branch simulates periodicity, reshaping the sequence into a two-dimensional image form. To simulate multi-period attribute decomposition, a dynamic squeezing operation is introduced, using the typhoon flag as a conditional input to compute a weighted global description vector. This operation dynamically amplifies the channel response during typhoon seasons, avoiding excessive smoothing dominated by regular weather conditions. Next, channel recalibration is performed, and the two-layer fully connected network (FC) is expanded into a residual cascade structure (Excitation with Residual). This residual connection ensures stable gradient propagation, improving the model's robustness to extreme perturbations.

[0028] The calculation process of the multi-scale dynamic channel attention enhancement module is as follows: (1) (2) (3) in, Indicates the first One channel, For the first Global characteristics of each channel These are learnable physical constraint parameters used to adjust the intensity of the impact of extreme weather conditions on the channel feature aggregation process. and Here are the weight matrices for the two fully connected layers. It is the sigmoid activation function. express Activation function. This represents the time index, corresponding to the first element in the input feature sequence. Each time step Indicates the first The time step, the first The original input feature values ​​on each feature channel. For the first The typhoon status indicator variable corresponding to each time step is used to characterize whether the current moment is in the typhoon impact stage. It serves as the condition input for the dynamic compression and channel weighting process, so as to distinguish the characteristic response mechanism under normal and extreme operating conditions. This represents the channel-level global feature vector. This represents the channel attention weight vector, which reflects the importance of each feature channel under the current operating conditions and their relative contribution. For the channel attention weight vector corresponding to the first The weight coefficients of each channel are used to adaptively amplify or suppress the features of that channel, thereby achieving dynamic recalibration at the channel level. This represents the output feature value modulated by the multi-scale dynamic channel attention mechanism.

[0029] MSDCAM not only enhances the spatiotemporal fusion of feature representations, but also provides adaptive weight support for zero-sample cross-terrain migration. The channel temporal fusion module comprises a parallel temporal convolutional module (TCN) and a dynamic module decomposition module (HAVOK). DTFM is the core component of the TTPAPM model, responsible for extracting and fusing multi-source features from the weighted features output by the channel SE module to improve wind power prediction accuracy. This module captures multi-scale temporal features through a parallel architecture: on the one hand, it utilizes the temporal convolutional module to extract local dependencies and generate local temporal embeddings. On the other hand, the dynamic modular decomposition (HAVOK) module is used to extract the periodicity and abrupt change patterns of key variables such as wind speed, generating dynamic modular decomposition embeddings. By fusing feature representations, this module enhances the ability to model complex dynamic systems under typhoon weather, effectively improving prediction accuracy.

[0030] The temporal convolution module improves upon the advantages of the block-residual structure by significantly expanding the temporal receptive field through bidirectional multi-scale dilated convolution and enhancing sensitivity and interpretability to power surge events through intra-block multi-head self-attention and typhoon conditional modulation, thereby improving prediction accuracy. The module's structure is as follows: Figure 3 As shown.

[0031] The temporal convolution module first processes the weighted features (i.e., output feature values) Divide into N overlapping blocks of fixed length. Each overlapping block Projection from linear mapping layer to high-dimensional space Generate local timing input For local timing input A bidirectional, multi-scale causal temporal convolutional structure operation is performed. Subsequently, a multi-head self-attention mechanism is introduced within the block to weight and model the bidirectional convolutional output. Convolution and ReLU activation functions are used to achieve feature reshaping and nonlinear enhancement. Additionally, a typhoon identifier variable is introduced. And through learnable coefficients The attention output is modulated with disaster intensity, and finally, the convolutional enhancement result is added to the original local input through a residual connection to obtain the output of the temporal convolutional module. This is achieved by introducing nonlinearity using an activation function, and the residual connection modulates the local temporal input... Added to the convolution output .

[0032] The calculation process of the temporal convolution module is shown below: (4) (5) (6) in, For weighted features The first one obtained by dividing along the time dimension A locally overlapping time block, For local timing input Intermediate feature representation after applying bidirectional multi-scale dilated causal convolution. The final output features of the temporal convolution module are... The weights of the compression layer, For the bias of the compression layer, The weights of the mapping layer, Here, represents the bias of the mapping layer, and represents the set of dilation rates for dilated causal convolutions. In this embodiment, we take . , This represents a three-layer dilated causal convolutional group that operates in parallel forward and backward directions. It employs a multi-head self-attention mechanism, calculating only within each overlapping block to avoid leaking future information across blocks. This is a one-dimensional convolution operation. The typhoon enhancement coefficient is a learnable parameter that automatically amplifies the weight of high-frequency disturbance heads during typhoon seasons. This represents the typhoon condition indicator variable, used to identify whether the current time block is in the typhoon impact stage.

[0033] Dynamic module decomposition module structure as follows Figure 4As shown, based on dynamic theory, the computational complexity is reduced by constructing an improved time-delay embedded Hankel matrix. The evolution of the nonlinear system is then mapped to an approximately linear observation space, and the dominant linear dynamics and external driving terms are decomposed. Compared to deep models that rely on gradient optimization, the Dynamic Modular Decomposition (DMD) module can directly extract the main dynamic modes and transient perturbation features from the time series without gradient learning. The DMD module first normalizes and reconstructs the weighted features output by the channel SE module using a sliding window to form a Hankel matrix. Then, it uses Singular Value Decomposition (SVD) to extract the dominant mode coefficients and solves the linear state equations using least squares regression, thereby obtaining the evolution of the system's hidden states and external force terms. The hidden state reflects stable periodic or slowly varying dynamics, while the external force terms characterize sudden perturbations and non-stationary features.

[0034] Through this "linear + external force" decomposition, the dynamic module decomposition module can accurately capture the transient change patterns of variables such as wind speed and air pressure under typhoon scenarios. The low-dimensional dynamic feature vector output by the dynamic module decomposition module is fused with the local temporal embedding generated by the temporal convolution module in the stitching projection layer, and then linearly mapped and input into the frozen Lag-Llama model backbone network to complete the wind power prediction task under extreme weather conditions. Its implementation includes the following steps: Weighted features of the SE module output of the channel Application layer normalization is obtained by stabilizing the feature distribution. For the normalized feature map, a sliding window unfolding operation is used to extract a feature map of length [length missing] over the time dimension. subsequence of: (7) in, To output the shape, The number of sliding windows. This represents the unfold operation. Weighted features of the SE module output of the channel The feature matrix obtained after applying layer normalization, This represents the dimension index on which the unfold operation is performed; in this embodiment, it is taken as... That is, expanding the sliding window over the time dimension. This represents the length of the sliding window, i.e., the number of time steps contained in each subsequence. , This represents the step size of the sliding window in the time dimension. The example uses... 。

[0035] The dimensions are rearranged using the permute operation to construct a multidimensional Hankel matrix: (8) in, For Hankel matrix, This indicates the permute operation.

[0036] right Perform low-rank singular value decomposition and take the first-order singular value. Each singular value defines a delay coordinate. Construct shifted data and then use the last coordinate. As an external force channel, the rest of the front Each is a system state. Using least squares estimation: (9) in, It represents the hidden state vector of the system in the next time step, reflecting the evolution result of the system under the combined action of internal dynamic mechanisms and external disturbances; This represents the system state transition matrix, used to characterize the hidden states. The linear evolution law under the condition of no external force reflects the inherent dynamic structure of wind power-related meteorological variables under normal operating conditions. Indicates the first The system hidden state vector at each time step is obtained by decomposing the delay coordinates. The system consists of several dominant modes, which are used to characterize the low-dimensional dynamic state of key meteorological variables such as wind speed and air pressure at the current moment. This represents the external force coupling coefficient vector, used to characterize the external force path. The influence intensity and direction of the system's hidden state evolution are determined, enabling explicit modeling of the effects of extreme weather disturbances. Indicates the first The external force channel variables corresponding to each time step are composed of the last dominant mode in the delayed coordinate decomposition, and are used to characterize the sudden disturbances or non-stationary effects introduced by extreme weather factors such as typhoons. This represents the discrete-time index, used to describe the gradual evolution of the system in the delay coordinate space, corresponding to the [number]th [time] after the sliding window is expanded. A time state.

[0037] For the hidden state vector of the system With the corresponding external force channel variable Concatenate the vectors to generate dynamic modular decomposition embedding vectors. Before stitching and projection, the channels are aligned with those of the temporal convolution module, and a linear mapping is applied: (10) in, The dynamic module decomposition embedding represents the dynamic feature representation after linear mapping. It is used to fuse with the local temporal embedding generated by the temporal convolution module in the splicing projection layer and is fed as input features into the frozen Lag-Llama backbone network. Represents dynamic embedding vectors The transpose of is used to match the linear mapping weight matrix in dimension, which facilitates subsequent linear projection operations. This represents a linear mapping weight matrix that embeds dynamic module decomposition into a unified feature space. It is used to align dynamic features and deep temporal features in the channel dimension, ensuring that different feature branches can be effectively fused. This represents the bias term in the linear mapping process, used to improve the flexibility and numerical stability of feature mapping and enhance the model's adaptability under different extreme weather intensities. and It has the same output dimension as the temporal convolution module.

[0038] Dynamic module decomposition module from weighted features By extracting dominant dynamic features and capturing the periodic fluctuations and abrupt changes in meteorological variables such as wind speed and wind direction, the two-stage adaptive prediction model for typhoons can accurately capture the complex nonlinear relationship between wind power and meteorological variables.

[0039] In the two-stage adaptive prediction model architecture for typhoon wind power forecasting, the temporal convolution module generates the input embedding. Dynamic module decomposition module generates dynamic module decomposition embedding The outputs of the two are generated by concatenation. Then, the input dimension of the Lag-Llama model is required to be... Therefore, a concatenated projection layer is needed as the output fusion layer to process the concatenated features into a suitable input format for the Lag-Llama model. This ensures the matching of different features with the input space of the pre-trained Lag-Llama model, avoiding performance degradation caused by insufficient information alignment. Simultaneously, it achieves information fusion, regularization, and numerical stability. The structure of the concatenated projection layer is as follows: Figure 5 As shown, the calculation process is as follows: (11) (12) in, Indicates the embedding of temporal convolutions With dynamic module decomposition and embedding The joint feature representation formed by concatenating features along the feature dimension. This represents a feature concatenation operation function, used to concatenate feature vectors from different sources along a specified dimension, thereby achieving explicit fusion of multi-source information. This indicates the batch size, which is the number of samples processed in parallel and is used to improve the computational efficiency of model training and inference. This represents the final fused feature representation output by the stitching projection layer. This represents a random deactivation operation, used to suppress feature overfitting during the training phase and improve the model's generalization ability under different typhoon intensities and paths. The representation layer normalization operation is used to standardize the numerical distribution of different feature channels, alleviate the problem of inconsistent feature scales, and improve the training stability after feature fusion. This represents the linear mapping weight matrix of the stitching projection layer, used to stitch high-dimensional features. from The feature dimensions required for projection back to the Lag-Llama model This achieves feature space alignment. This represents the bias term corresponding to the splicing projection layer, used to enhance the flexibility of feature mapping and improve numerical stability.

[0040] Linear layers are used for mosaic projection layers. , from dimension Projected to To ensure compatibility with pre-trained Lag-Llama models and retain their pre-training knowledge, ReLU enhances nonlinear expressiveness and adapts to the complex patterns of wind power data. Dropout prevents the model from overfitting to a single feature (such as over-reliance on DMD features), allowing the model to... The data is input into a pre-trained Lag-Llama model for wind power prediction.

[0041] The pre-trained temporal base model uses Lag-Llama, and its structure is as follows: Figure 6As shown, Lag-Llama is based on a decoder-only Transformer architecture, with its core employing a tokenization scheme of "lag features + temporal covariates". The model concatenates the lag features corresponding to historical observations with temporal features to form the input token, which is then mapped to the attention latent space through a shared linear projection layer. Pre-normalization and relative position encoding are introduced into the network to enhance its ability to model temporal dependencies. The main body of the model consists of a Transformer decoding layer with a causal mask, ensuring that predictions are made using only historical information. At the output, a distributed prediction head maps the decoder's hidden state to directly predict the probability distribution parameters of the target variable at the next time step, and training is performed based on log-likelihood, thus achieving probabilistic prediction and uncertainty characterization. Lag-Llama was pre-trained on a large-scale cross-domain time-series dataset covering 27 fields, including energy and transportation, learning a generalized time-series representation with good generalization ability. It exhibits excellent zero-shot and fine-tuning performance in downstream tasks, providing a fundamental model support for time-series prediction under complex conditions.

[0042] The output mapping layer is the final module in the two-stage adaptive prediction model architecture for typhoons, responsible for mapping the high-dimensional hidden output representation of the Lag-Llama model. Mapping to target prediction dimension This layer facilitates the transformation from the abstract semantic space to the concrete physical quantity space, playing a crucial bridging role in the prediction process. The mathematical expression is: (13) in, Indicates the first The sample, the first The model outputs predicted values ​​at each prediction time step, corresponding to the prediction results of the target physical quantity (such as wind power). This indicates that the Lag-Llama model is in the... The sample, the first The high-dimensional hidden state vector corresponding to each time step integrates historical lag features, temporal covariates, and information provided by preceding modules (temporal convolution and dynamic decomposition), reflecting the model's comprehensive understanding of temporal structure and the impact of extreme weather in the abstract semantic space. The linear mapping weight matrix of the output mapping layer is used to project the high-dimensional hidden representation from the semantic feature space to the target physical quantity space, thereby achieving scale and semantic alignment between the internal representation of the model and the actual predicted quantity. This represents the bias term of the output mapping layer, used to compensate for systematic offsets that may be introduced during the linear mapping process, thereby improving the numerical stability and overall fitting ability of the prediction results.

[0043] S3. A two-stage hybrid optimization training strategy is used to train the two-stage adaptive prediction model for typhoons.

[0044] The dynamic characteristics inherent in typhoon scenarios are often sudden and highly nonlinear. These characteristics account for a small proportion of conventional data and are difficult to fully capture with general representations, leading to limited prediction performance when directly applying the Lag-Llama model. Therefore, this embodiment proposes a two-stage hybrid optimization training paradigm. Under black-box conditions, it only relies on the inference output of the Lag-Llama model to optimize the parameters of the input and output mapping layers without accessing their internal structure and gradients. The first stage uses a conventional weather training dataset obtained from the official platform of the 5th "Goldwind Cup" Energy Innovation Challenge to enhance the model's generalization ability: zero-order optimization (ZOO) is applied to all input modules of the Lag-Llama model, while first-order optimization is applied to the output mapping layer to ensure stable convergence. The second stage, based on the typhoon dataset, continues to perform lightweight fine-tuning of the DTFM module and the output mapping layer to achieve targeted calibration for the severe non-stationarity under typhoon scenarios, thereby improving the accuracy of wind power prediction under extreme weather conditions. The specific optimization process is detailed in [link to optimization process]. Figure 7 .

[0045] The first phase involves training on a standard weather dataset to enhance the generalization ability of the two-stage adaptive prediction model for typhoons. To achieve efficient parameter learning, all internal parameters of the pre-trained Lag-Llama model's backbone network are frozen during training, and only the circumstantial adaptation layers are trained. A hybrid optimization strategy is employed: zero-order optimization is used to update parameters for all input modules of the Lag-Llama model (channel SE module and dynamic temporal fusion module), while standard first-order optimization is used for the output mapping layer to ensure stable convergence during training.

[0046] Zero-order optimization uses the stochastic gradient estimator (RGE) method, which approximates the gradient by performing finite difference lookups in random directions. This allows for effective parameter updates even in cases where differentiability is lacking or the gradient is difficult to compute. The specific implementation steps are as follows: (14) in, To estimate the gradient, For the number of queries, The random perturbation step size, For random perturbation direction, and loss function Loss values ​​in different perturbation directions, and loss function Loss values ​​in different perturbation directions, Indicates the first random perturbation direction Above, along the positive direction (positive perturbation) on the parameters The applied amplitude is The value of the loss function after a small perturbation. This term reflects how the loss function changes as the parameter "moves forward" in this random direction. Indicates the same random direction Above, along the negative direction (reverse perturbation) on the parameters The applied amplitude is The value of the loss function corresponding to a small perturbation. This term reflects how the loss function changes when the parameter is "shifted backward" in this direction. An index representing the direction of random perturbation, used to distinguish different perturbation directions involved in stochastic gradient estimation. This represents the model's parameter vector, which contains all trainable parameters that need to be updated during the zero-order optimization process.

[0047] Based on the obtained gradient, Update: (15) in, Update the step size for the parameters. Indicates the first The parameter vector of the model at each iteration contains all the adjustable parameters that need to be optimized. Indicates the first In the next iteration, the corresponding parameters The loss function value is used to measure the current model's predictive performance. This represents the iteration index, used to identify the number of iterations in a zero-order optimization algorithm.

[0048] To further improve the convergence stability and training efficiency of the output mapping layer, this embodiment employs the LAMB (Layer-wise Adaptive Moments optimizer for Batch training) first-order optimization algorithm. This method obtains the loss gradient based on the standard backpropagation mechanism and introduces adaptive moment estimation and layer confidence ratio adjustment to achieve stable updates of large-scale parameters and multi-scale features during the training process of the output mapping layer. In each iteration, a loss function is constructed based on the model's predicted output and the actual observed values. ,in .

[0049] Using the standard backpropagation mechanism, the gradients of the output mapping layer weights and biases are calculated respectively: (16) in, To output the gradient of the weight parameters of the mapping layer, The gradients of the output mapping layer bias parameters are described above, which characterize the local change direction of the loss function near the current parameter point. Indicates the first The bias parameters of the output mapping layer are output before the next iteration.

[0050] In the In this iteration, based on the gradient of the current loss function: (17) Calculate the first and second moments of the gradient with respect to the weight parameters and bias parameters, respectively.

[0051] The first moment is used to characterize the exponentially weighted average of the gradient, and its calculation formula is as follows: (18) in, This is the first-order moment decay coefficient, used to control the degree to which historical gradient information is preserved.

[0052] To eliminate the initial stage factors The introduced deviation is corrected for the first moment, resulting in: (19) The second moment is used to characterize the exponentially weighted average of the squared gradient, and its calculation formula is as follows: (20) in, The second-order moment attenuation coefficient, This represents the element-wise square of the gradient.

[0053] Similarly, after correcting for the deviation of the second moment, we obtain: (twenty one) Based on this, a standardized update direction vector is constructed: (twenty two) in, This is a numerically stable term used to avoid the denominator being zero.

[0054] Secondly, LAMB introduces a "layer trust ratio" mechanism, which adaptively scales the learning step size of different layer parameters by comparing the norm of the current weight parameters with the norm of their candidate update directions. This mechanism effectively alleviates training instability caused by inconsistent parameter scales across layers. Based on the learning rate and the corrected update direction, the parameters are then updated. and The iterative update is performed. The calculation formula is as follows: The LAMB optimization algorithm introduces a hierarchical trust ratio in weight updates: (twenty three) in, Indicates the first Output the weight parameters of the mapping layer before the next iteration. This is a standardized update direction vector constructed based on the first and second moments.

[0055] Based on this, complete the parameter update: (twenty four) (25) in, and They represent the first The weight parameters of the output mapping layer before and after each iteration; and They represent the first The bias parameters of the output mapping layer before and after each iteration; and The gradient of the loss function with respect to the weights and biases indicates the direction of the current parameter update. and These represent the first moment estimate of the gradient and its bias correction result, respectively, used to characterize the expected trend of gradient change; and These represent the second-order moment estimate of the gradient and its bias correction result, respectively, used to characterize the fluctuation range of the gradient; This is a standardized update direction vector constructed based on the first and second moments; and These are the exponential decay coefficients of the first and second moments, respectively; This is the global learning rate, used to control the step size for parameter updates; The hierarchical trust ratio is used to adaptively adjust the actual learning step size of different layers according to the parameter scale. This is a numerical stability term used to prevent the denominator from being zero and to improve computational stability.

[0056] In the two-stage adaptive prediction model for typhoons, first-order optimization is applied to the output projection layer using LAMB. This layer is responsible for mapping the representation at the decoding end of the pre-trained Lag-Llama model to the real power point prediction. To achieve efficient parameter fine-tuning, all parameters of the pre-trained Lag-Llama model are frozen, and gradient updates are performed only on the output projection layer. This allows the model to quickly adapt to specific task requirements, avoiding the high computational cost and catastrophic forgetting problem caused by full parameter fine-tuning.

[0057] The second stage involves adaptive fine-tuning for extreme weather. After training on a regular weather dataset, the two-stage adaptive prediction model for typhoons begins adaptive parameter fine-tuning on samples from the typhoon dataset. Addressing the characteristics of short-term wind speed jumps and power spikes during typhoons, the model's adaptability to extreme, non-stationary weather processes is enhanced. In terms of optimization strategy, the second stage continues the hybrid optimization paradigm: zero-order optimization is used for the dynamic time-series fusion module, and first-order gradient optimization is used for the output mapping layer. This enables the model to quickly capture the dynamic characteristics of the wind speed-power relationship under limited typhoon sample conditions, improving prediction accuracy in extreme weather scenarios while maintaining its original generalization ability.

[0058] S4. Input the data obtained in S1 into the trained two-stage adaptive prediction model of typhoon obtained in S3 to obtain the predicted wind power value.

[0059] In one specific embodiment, the effectiveness of the embodiments of this application is verified through experiments. All comparative experiments are based on the official dataset of the 5th "Goldwind Cup" Energy Innovation Challenge finals and the supplementary typhoon-specific dataset. To ensure the repeatability of the experiments and the reliability of the results, all data undergo a unified and rigorous preprocessing process. Specifically, the raw data is first parsed and sorted according to the time column, and two years of complete observation records are selected to ensure that the data has sufficient timeliness and representativeness. Subsequently, the feature columns (including key meteorological variables such as wind speed, wind direction, temperature, and humidity) and the target power column are automatically identified, and missing values ​​are filled using time-based linear interpolation combined with forward and backward imputation strategies to ensure the continuity of the time series and to avoid introducing significant bias. In the typhoon-specific data, the typhoon_flag field is further annotated according to the provided typhoon impact period dictionary to distinguish between regular weather and typhoon weather samples. Normalization employs a dual-scaler strategy: regular weather samples are standardized using a global MinMaxScaler, while typhoon weather samples are normalized using a separate MinMaxScaler. This preserves the extreme numerical distribution characteristics during typhoon seasons while avoiding excessive smoothing of the normalization parameters for typhoon samples by regular samples. The target power is uniformly normalized using a single MinMaxScaler to ensure consistent physical meaning after inverse normalization. The normalization formula is as follows: (26) in, These are actual observed values. The maximum value of the observations in the entire dataset. It represents the minimum value of the observations in the entire dataset.

[0060] After normalization, hourly, daily, and monthly features are extracted as temporal covariates. A supervised sequence incorporating historical power is constructed using a sliding window approach. The input sequence simultaneously includes historical meteorological features, historical power, and temporal features. The prediction target is the normalized power value at multiple time steps. The training and test sets are strictly divided in an 8:2 ratio, ensuring that the test set includes the complete typhoon impact period. Finally, all sequences are batch-processed and converted to PyTorch tensors to fully utilize GPU parallelism for accelerated training and inference. The above preprocessing procedure remains consistent across different wind farms, terrains, and weather scenarios, ensuring the fairness and reliability of all subsequent comparative experimental results.

[0061] The specific experiment includes the following parts: First, a two-stage training of the TTPAPM model was completed based on data from a specific offshore wind farm (station0), evaluating its predictive performance under both normal and typhoon weather conditions. Next, the model's zero-shot capability was verified in a cross-terrain setting, i.e., the model obtained from offshore wind power predictions was directly applied to a hilly wind farm (station1) without any retraining or fine-tuning. Finally, some baseline models were selected to compare the model's superiority. This embodiment uses Mean Absolute Error (MAE). Root Mean Square Error (RMSE) and the coefficient of determination ) As the main evaluation index for model prediction accuracy.

[0062] (27) (28) (29) in, To predict the length of time, For the first The actual observed power for each time period For the first Predicted power for each time period, This represents the average of the actual observed power over all time periods.

[0063] To verify the wind power prediction performance of the proposed method in the extreme weather scenario of typhoons, this embodiment selects offshore wind farms as an independent test set to systematically evaluate the generalization and transfer effects of the two-stage training strategy. The first stage involves pre-training the model on a general offshore wind power dataset to learn the general wind speed-power mapping and temporal characteristics. The second stage implements adaptive fine-tuning based on offshore typhoon data to enhance the model's ability to represent high fluctuations and non-stationary disturbances. The test results of the two stages are as follows: Figure 8 and Figure 9 As shown.

[0064] Figure 8 and Figure 9 The model's prediction performance under normal weather and typhoon conditions is presented, and the results verify the effectiveness of the proposed method. Table 1 shows the model's prediction performance under normal weather conditions. It is 0.55%. It is 0.87%. A score of 0.95 indicates that the two-stage adaptive prediction model for typhoons, while keeping the main parameters of the Lag-Llama model frozen, has achieved high-precision fitting through efficient training and optimization of external modules, without significant lag or systematic bias. In the second stage, after small-scale parameter optimization of the DTFM module and output mapping layer under typhoon samples, the model achieved [high accuracy] during typhoon events. =0.89%, =1.32%, The accuracy is 0.89. Although the error is higher than that under normal circumstances, it can still maintain a low level of error under the strong non-stationarity and high-frequency disturbances caused by typhoons, indicating that the model proposed in this application has good robustness and generalization ability, and improves the ability to characterize sudden power transitions and short-term fluctuations.

[0065] Table 1. Two-stage forecast results indicators for wind farms using TTPAPM

[0066] To evaluate the zero-shot generalization ability of TTPAPM, this embodiment selects hilly wind farms with significantly different terrain conditions as external validation datasets. TTPAPM trained on the offshore wind farm dataset is used directly for prediction without any retraining or fine-tuning. To ensure the interpretability and comparability of the experiment, a correlation analysis between wind speed and wind power is first conducted for both offshore and hilly wind farms, using the correlation stability of Pearson and Spearman rank correlation coefficients.

[0067] like Figure 10As shown, the wind speed-power relationship between offshore wind farms and hilly wind farms exhibits significant differences. Scatter fitting results indicate that both scenarios display typical S-shaped power curve characteristics, but offshore wind farms have a wider wind speed distribution range and tend to saturate power output in the medium-to-high wind speed range; while hilly wind farms show obvious nonlinear fluctuations in the lower wind speed range, indicating that they are more significantly affected by turbulence intensity and topographic disturbance under complex terrain. Correlation analysis shows that the Pearson linear correlation coefficient between offshore and hilly wind power is low (r≈0.11–0.12), indicating that the linear correlation between the two is limited under typhoon conditions; Spearman rank correlation analysis further shows that the wind speed and power at the two sites are almost uncorrelated in terms of monotonicity (r≈0.09–0.10). These results demonstrate that there are significant differences in the wind speed-power relationship between different types of wind farms, and traditional transfer methods that rely on strong correlation or monotonicity assumptions are difficult to guarantee prediction accuracy. Especially when historical observation data for the target scenario is lacking, insufficient cross-scenario power correlation will further amplify the uncertainty of zero-sample prediction. To verify the zero-shot capability of the proposed model, the model obtained from the two-stage offshore wind power project was used to perform zero-shot predictions directly on samples from hilly wind farms under normal and typhoon weather conditions without any retraining or fine-tuning. The results are as follows: Figure 11 and Figure 12 As shown.

[0068] Figure 11 and Figure 12 The zero-sample prediction performance of the model under normal weather and typhoon weather conditions in hilly wind farms is presented respectively. As shown in Table 1, under normal weather conditions, the model's various evaluation indicators are... =1.24%, =1.45%, =0.85, with the error remaining at a low level, indicating that TTPAPM can accurately capture the dynamic characteristics and amplitude changes of power under conditions of lack of target supervision, demonstrating good cross-domain generalization ability. Compared with the prediction results of offshore wind farms, the model's cross-scenario generalization performance under typhoon weather is somewhat reduced due to the influence of terrain roughness and turbulence intensity in hilly wind farms. =0.32%, =0.68%, =0.71), but the key error indicators are still controlled within a relatively low range, indicating that TTPAPM still maintains a certain accuracy under the cross-terrain zero-sample setting and has the ability to adapt to scene differences.

[0069] To verify the performance of the proposed model in improving wind power prediction accuracy during typhoon weather, Informer, transformer, autoformer, and FEDformer networks were introduced as comparative models to verify the wind power prediction effects of different models under normal and typhoon weather conditions. The prediction results are as follows: Figure 13 , Figure 14 As shown in Table 2.

[0070] Table 2 Evaluation Indicators for Offshore Wind Power Prediction under 5 Models

[0071] The comparison results show that, under normal weather conditions, TTPAPM achieved the best forecast accuracy. , and The percentages were 0.55%, 0.87%, and 0.95%, respectively. Compared to the Informer model, which performed the second best, TTPAPM... It decreased by another 0.25 percentage points. It decreased by 0.66 percentage points. It improved by 0.09; and compared to the weakest performer, Autoformer, its performance was... and The decrease was approximately 50%. An increase of 0.15. Despite more challenging typhoon weather conditions, TTPAPM maintained its lead. , and They were 0.89%, 1.32%, and 0.90%, respectively. It is on par with Transformer, which is the lowest among all. and All are optimal. Autoformer's prediction performance deteriorates significantly under typhoon conditions. For example... Figure 13 and Figure 14 As shown, especially in the extreme peak and trough ranges, TTPAPM tracks the true power peak more accurately, while Transformer, FEDformer, and Informer overestimate or underestimate to varying degrees, and Autoformer also exhibits significant shape distortion. In the trough and plateau phases, TTPAPM generally falls within the fluctuation range, while Transformer and FEDformer show stronger randomness, Informer exhibits significant lag, and Autoformer displays a "sawtooth" oscillation. These results indicate that TTPAPM not only has a smaller average error but is also more sensitive to large errors. It also maintains a leading position, and is able to track power fluctuations more stably, demonstrating that TTPAPM has stronger robustness and generalization ability under extreme wind conditions.

[0072] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely illustrative of the principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the present invention as claimed. The scope of protection of this invention is defined by the appended claims and their equivalents.

Claims

1. A method for predicting offshore wind power based on a pre-trained time-series large model, characterized in that, Includes the following steps: S1. Acquire offshore wind power data and numerical weather prediction data; S2. Construct a two-stage adaptive prediction model for typhoons. The two-stage adaptive prediction model for typhoons includes a channel SE module, a dynamic temporal fusion module, a stitching projection layer, a pre-trained temporal base model, and an output mapping layer connected in sequence. The channel temporal fusion module includes a parallel temporal convolution module and a dynamic module decomposition module. S3. A two-stage hybrid optimization training strategy is adopted to train the two-stage adaptive prediction model for typhoons. In the first stage, the input module is trained on regular weather data, and in the second stage, the dynamic time series fusion module and the output mapping layer are fine-tuned on typhoon data. S4. Input the data obtained in S1 into the trained two-stage adaptive prediction model of typhoon obtained in S3 to obtain the predicted wind power value.

2. The offshore wind power prediction method based on a pre-trained time-series large model according to claim 1, characterized in that, The channel SE module employs a multi-scale dynamic channel attention enhancement module, and the implementation process is as follows: The multi-scale dynamic channel attention enhancement module transforms the input feature map into a spatiotemporal representation through multi-scale convolutional branches. The multi-scale convolutional branches include parallel temporal branches and simulated periodic branches. The temporal branches capture pure temporal dependencies, while the simulated periodic branches reshape the sequence into a two-dimensional image pattern. The time-series branch output and the simulation periodic branch output are dynamically compressed, and the weighted global description vector is calculated by combining the typhoon sign as a conditional input. Channel recalibration is performed, and the two-layer fully connected network is extended into a residual cascade structure to ensure stable gradient propagation. The calculation process of the multi-scale dynamic channel attention enhancement module is as follows: (1) (2) (3) in, For time step, For the number of channels, Indicates the first One channel, For the first Global characteristics of each channel For learnable physical constraint parameters, and Here are the weight matrices for the two fully connected layers. It is the sigmoid activation function. express Activation function Indicates a time index. Indicates the first The time step, the first The original input feature values ​​in each feature channel For the first The typhoon status indicator variables corresponding to each time step This is a channel-level global feature vector. Here is the channel attention weight vector. For the channel attention weight vector corresponding to the first The weighting coefficients of each channel The output feature value is modulated by the multi-scale dynamic channel attention mechanism.

3. The offshore wind power prediction method based on a pre-trained time-series large model according to claim 2, characterized in that, The implementation process of the temporal convolution module is as follows: The weighted features output by the channel SE module Divide into blocks, cut into A fixed-length overlapping block ; Each overlapping block Projected into a higher-dimensional space via a linear compression-mapping structure Generate local timing input ; For local timing input A bidirectional, multi-scale causal temporal convolutional structure is performed, followed by weighted modeling of the bidirectional convolutional output, feature reshaping, and nonlinear enhancement. Additionally, a typhoon identifier variable is introduced, and learnable coefficients are used to perform the modeling. The attention output is modulated by disaster intensity, and finally the convolution enhancement result is added to the original local input through residual connection to obtain the output of the temporal convolution module; The calculation process of the temporal convolution module is as follows: (4) (5) (6) in, For weighted features The first one obtained by dividing in the time dimension A locally overlapping time sequence block, For local timing input Intermediate feature representation after applying bidirectional multi-scale dilated causal convolution. These are the final output features of the temporal convolution module. The weights of the compression layer, For the bias of the compression layer, For the weights of the mapping layer, For the bias of the mapping layer, For the set of dilation rates of dilated causal convolutions, This represents a three-layer dilated causal convolutional group that operates in parallel forward and backward directions. This is a multi-head self-attention mechanism. This is a one-dimensional convolution operation. The learnable typhoon enhancement coefficient. This is an indicator variable for typhoon conditions.

4. The offshore wind power prediction method based on a pre-trained time-series large model according to claim 3, characterized in that, The implementation process of the dynamic module decomposition module is as follows: Weighted features of the SE module output of the channel Application layer normalization is obtained by stabilizing the feature distribution. For the normalized feature map, a sliding window expansion operation is used to extract a feature map of length [length missing] over the time dimension. subsequence of: (7) in, To output the shape, The number of sliding windows. This indicates the operation of expanding the sliding window. Weighted features of the SE module output of the channel The feature matrix obtained after applying layer normalization, This represents the dimension index in which the sliding window expansion operation is performed. Indicates the length of the sliding window. This indicates the step size of the sliding window in the time dimension; The dimensions are rearranged using the permute operation to construct a multidimensional Hankel matrix: (8) in, For Hankel matrix, Indicates the permute operation; right Perform low-rank singular value decomposition and take the first-order singular value. Each singular value defines a delay coordinate. Construct the shifted data and then use the last coordinate. As an external force channel, the rest of the front Each is a system state. Least squares estimation: (9) in, This represents the hidden state vector of the system at the next time step. Represents the system state transition matrix. Indicates the first The system's hidden state vector at each time step Represents the vector of external force coupling coefficients. Indicates the first The external force channel variables corresponding to each time step Indicates a discrete-time index; For the hidden state vector of the system With the corresponding external force channel variable Concatenate the vectors to generate dynamic modular decomposition embedding vectors. Before stitching and projection, the channels are aligned with those of the temporal convolution module, and a linear mapping is applied: (10) in, For dynamic module decomposition embedding, it represents the dynamic feature representation after linear mapping. Represents dynamic embedding vectors The transpose form, The linear mapping weight matrix represents the embedding of dynamic module decomposition into a unified feature space. This represents the bias term in the linear mapping process.

5. The offshore wind power prediction method based on a pre-trained time-series large model according to claim 4, characterized in that, The stitching projection layer is used to fuse the output of the temporal convolution module and the output of the dynamic module decomposition module and map them to the input dimension of the pre-trained temporal base model. The calculation process is as follows: (11) (12) in, Indicates will With The joint feature representation formed by concatenating features along the feature dimension. This represents the feature concatenation operation function. Indicates batch size, This represents the final fused feature representation output by the stitching projection layer. This indicates a random deactivation operation. Presentation layer normalization operation, This represents the linear mapping weight matrix of the spliced ​​projection layer. This indicates the bias term corresponding to the splicing projection layer.

6. The offshore wind power prediction method based on a pre-trained time-series large model according to claim 5, characterized in that, The pre-trained temporal base model adopts the Lag-Llama model, which is based on the Transformer architecture with only a decoder. It uses lag features and time covariates as inputs and obtains a general temporal representation through pre-training.

7. The offshore wind power prediction method based on a pre-trained time-series large model according to claim 6, characterized in that, The output mapping layer maps the hidden representation of the pre-trained temporal base model to the predicted wind power value. The calculation process is as follows: (13) in, Indicates the first The sample, the first The model outputs predicted values ​​at each prediction time step. The model represents the first time. The sample, the first The high-dimensional hidden state vector corresponding to each time step This represents the linear mapping weight matrix of the output mapping layer. This represents the bias term of the output mapping layer.

8. The offshore wind power prediction method based on a pre-trained time-series large model according to claim 7, characterized in that, The two-stage hybrid optimization training strategy includes: In the first stage, a regular weather dataset is used as input. The zero-order optimization method is used to update the parameters of the channel SE module and the dynamic time series fusion module, and the first-order optimization method is used to update the output mapping layer. The second stage uses the typhoon dataset as input, employs a zero-order optimization method to update the parameters of the dynamic time series fusion module, and a first-order optimization method to update the output mapping layer.

9. The offshore wind power prediction method based on a pre-trained time-series large model according to claim 8, characterized in that, The zero-order optimization method is calculated using the following formula: (14) in, To estimate the gradient, For the model's parameter vector, For the number of queries, The random perturbation step size, For random perturbation direction, and For loss function Loss values ​​in different perturbation directions; right Update: (15) in, Update the step size for the parameters. Indicates the first The parameter vector of the model at the next iteration. Indicates the first In the next iteration, the corresponding parameters The loss function value.

10. The offshore wind power prediction method based on a pre-trained time-series large model according to claim 9, characterized in that, The first-order optimization method includes the following steps: Using the standard backpropagation mechanism, the gradients of the output mapping layer weights and biases are calculated respectively: (16) in, To output the gradient of the weight parameters of the mapping layer, To output the gradient of the bias parameters of the mapping layer, For the first The bias parameters of the output mapping layer are output before the next iteration; In the In this iteration, based on the gradient of the current loss function: (17) Calculate the first and second moments of the gradient with respect to the weight parameters and bias parameters, respectively. The first moment is used to characterize the exponentially weighted average of the gradient, and its calculation formula is as follows: (18) in, This is the first-order moment decay coefficient, used to control the degree to which historical gradient information is preserved; After correcting for the deviation of the first moment, we get: (19) The second moment is used to characterize the exponentially weighted average of the squared gradient, and its calculation formula is as follows: (20) in, The second-order moment attenuation coefficient, This represents the element-wise square of the gradient; After correcting for the deviation of the second moment, we get: (21) Construct a standardized update direction vector: (22) in, It is a numerically stable term; Introducing hierarchical trust ratio in weight updates: (23) in, For the first Output the weight parameters of the mapping layer before the next iteration. This is a standardized update direction vector constructed based on the first and second moments; Based on this, complete the parameter update: (24) (25) in, For the first The weight parameters of the mapping layer are output after the next iteration. To indicate the first The bias parameters of the output mapping layer are output after the next iteration. This represents the global learning rate.