Multi-step Wind Power Prediction Method Based on Transformer Network with Multimodal and Multitask

By introducing a multimodal multitasking Transformer network (M2TNet) in wind power power prediction, combining three-stage modeling strategies and maximum information coefficient variable selection, the existing models have solved the shortcomings in multi-step prediction accuracy and time complexity, and achieved more efficient wind power prediction.

CN114118569BActive Publication Date: 2025-06-20NINGBO LIDOU INTELLIGENT TECH CO LTD +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111403178.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-24
Publication Date
2025-06-20
Estimated Expiration
2041-11-24

AI Technical Summary

Technical Problem

Existing wind power power prediction models have potential to improve prediction accuracy and time complexity in multi-step prediction, especially when processing multi-source heterogeneous data, it is difficult to effectively integrate and unify these data.

Method used

A multi-step prediction method for wind power based on Transformer network (M2TNet) based on multimodal multitasking is proposed. By designing a deep M2TNet prediction model, a three-stage modeling strategy is adopted: feature extraction, feature fusion and prediction terminal, combining maximum information coefficient variable selection and multimodal multitasking learning, to improve the prediction accuracy of the model and reduce the time complexity.

Benefits of technology

It effectively improves the accuracy of multi-step prediction of wind power, reduces the time complexity of the model, can better process multi-source heterogeneous data, and provides more reliable power grid operation support.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114118569B_ABST
    Figure CN114118569B_ABST
Patent Text Reader

Abstract

The present invention discloses a multi-step wind power prediction method based on a Transformer network for multi-modal multi-tasks, belonging to the field of power system planning, including: designing a deep M2TNet prediction model; collecting K types of multi-source heterogeneous data and performing variable selection using the maximum information coefficient analysis to select S types of optimal variables; normalizing the optimal variables and dividing the training set and the test set; determining the hyperparameters of the M2TNet model by means of artificial experience and grid search in combination with the training effect of the model; saving the parameter model with the best training effect and performing multi-step wind power prediction using the data of the test set; performing anti-normalization on the prediction results to obtain the wind power prediction values, and analyzing the prediction results in combination with evaluation indicators. It solves the problem of multi-step wind power prediction for multi-source heterogeneous data and improves the performance of the existing model, which can not only improve the prediction accuracy of the existing model but also reduce its time complexity.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of power system planning, and more specifically, relates to a multi-step wind power prediction method based on a multi-modal multi-task Transformer network. Background Art

[0002] Wind energy is one of the most promising clean energies. Wind power generation is one of the main forms of wind energy utilization and has received increasing attention from researchers in recent years. Wind energy has the characteristics of randomness and volatility, which poses a certain threat to the safe and stable operation of the power grid. Therefore, wind power prediction is an essential link for wind farms to connect to the grid. Wind farm power prediction can provide an effective basis for power generation, dispatching, operation and maintenance, etc. When large-scale wind energy is used for power generation, the research on wind power prediction has important practical significance for ensuring the reliable, stable and economic operation of the power grid. Compared with single-step prediction, multi-step prediction can better reflect the general situation and is more widely used in real life. From the perspective of machine learning, multi-step prediction can be transformed into a multi-task learning problem of multiple single-step predictions. From the data-driven perspective, the research on multi-step prediction generally uses multi-source heterogeneous data of wind power and numerical weather prediction (NWP), but it is challenging to respect the differences and unity of these data. In addition, the existing prediction models have great potential for improvement in terms of prediction accuracy and time complexity. Summary of the Invention

[0003] Aiming at the above defects or improvement requirements of the existing technology, the present invention proposes a multi-step wind power prediction method based on a multi-modal multi-task Transformer network (M2TNet) prediction model, which solves the problem of multi-step prediction of wind power with multi-source heterogeneous data and improves the performance of the existing models, not only improving the prediction accuracy of the existing models but also reducing their time complexity.

[0004] To achieve the above object, according to one aspect of the present invention, there is provided a multi-step wind power prediction method based on a multi-modal multi-task Transformer network, including:

[0005] Designing a deep M2TNet prediction model;

[0006] Collecting K types of multi-source heterogeneous data and performing variable selection using the maximum information coefficient analysis, and taking S types of selected variables as the input of the M2TNet prediction model, where S ≤ K;

[0007] Normalizing the S types of selected variables to obtain sample data, and dividing the sample data into a training set and a test set;

[0008] Determine the hyperparameters of the M2TNet model by means of manual experience and grid search and in combination with the effect of model training;

[0009] Save the parameter model with the best training effect, and use the data of the test set for multi-step prediction of wind power;

[0010] Perform inverse normalization on the prediction results to obtain the wind power prediction value, and analyze the prediction results in combination with evaluation indicators.

[0011] In some alternative embodiments, the M2TNet prediction model has a feature extraction layer, a feature fusion layer, and a prediction terminal layer; among them, the feature extraction layer adopts a multi-branch structure with multiple Transformer units in parallel; the feature fusion layer is an ordinary fully connected network; the prediction terminal layer is a regression layer.

[0012] In some alternative embodiments, design a deep M2TNet prediction model, including:

[0013] The first stage: Use multiple feature extractors to extract high-order features from various historical features respectively, where the Transformer unit is used as the feature extractor;

[0014] The second stage: Adopt a multi-modal learning strategy to fuse the high-order features from different historical features to form a unified feature containing richer information;

[0015] The third stage: Adopt a multi-task learning strategy to predict the wind power sequence based on the unified feature.

[0016] In some alternative embodiments, the M2TNet prediction model is described as: where, y t is the wind power sequence within a certain future time period at time t, x t represents the input vector, f(·) represents the implicit function of the prediction model, represents the sub-vector of the historical feature i corresponding to time t, h t represents the unified feature, represents the high-order feature corresponding to the historical feature i, g i (·) represents the implicit function of the feature extractor corresponding to the historical feature i (i = 1, 2,..., S) in the first stage, v(·) and u(·) respectively represent the implicit functions of the functions in the second stage and the third stage, and S represents the number of historical features.

[0017] In some alternative embodiments, collect K types of multi-source heterogeneous data and perform variable selection using the maximum information coefficient analysis, and use S types of selected variables as the input of the M2TNet prediction model, including:

[0018] The mutual information is used to measure the correlation between the wind power and the historical variables of other heterogeneous data, and the MIC value is calculated and sorted in descending order. The top S variables are used as the preferred variables.

[0019] In some alternative embodiments, the hyperparameters of the M2TNet model are determined by means of manual experience and grid search in combination with the effect of model training, including:

[0020] All Transformers adopt the same structural parameters; the number of units in each hidden layer of the Transformer remains consistent. Therefore, the hyperparameters to be determined in M2TNet include: the number of encoding layers of each Transformer, the number of neurons in each hidden layer, the size of the learning rate, the size of Epochs, the size of Batch Size, and the size of Dropout; the optimizer of M2TNet adopts RMSProp, and its loss function is mse;

[0021] The hyperparameters of the M2TNet model are determined by means of manual experience and grid search in combination with the effect of model training.

[0022] In some alternative embodiments, by and are used as evaluation metrics, where y k and are the actual value and the predicted value of the wind power at the k-th moment within the error statistical time interval respectively, n is the total number of time periods in this interval, and C N is the installed capacity of the wind farm to be predicted.

[0023] According to another aspect of the present invention, there is provided a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the method described in any one of the above are implemented.

[0024] Generally speaking, compared with the prior art, the above technical solution conceived by the present invention can achieve the following beneficial effects:

[0025] A method for multi-step prediction of wind power based on a Transformer network with multi-modal and multi-task capabilities uses the maximum information coefficient to perform variable selection on multi-source heterogeneous data, thereby achieving the goals of controlling the model scale, reducing computational complexity, and improving model performance. A three-stage modeling strategy is proposed to solve the problem of multi-step prediction of wind power using multi-source heterogeneous data. A deep M2TNet prediction model is designed, which has a feature extraction layer, a feature fusion layer, and a prediction terminal layer. Among them, the feature extraction layer adopts a multi-branch structure with multiple Transformer units, and this structural design enables each Transformer unit to separately process each type of heterogeneous data in a targeted manner. In addition, the advantages of Transformer in mining feature information and computational efficiency are conducive to improving the performance of the overall model. The multi-step prediction problem is transformed into multiple single-step prediction tasks. In the feature fusion layer and the prediction terminal layer, the M2TNet model adopts multi-modal and multi-task learning strategies. Multi-modal learning is beneficial for integrating information between different modalities output by multiple Transformer units to form high-level features for realizing complex tasks. Multi-task learning is beneficial for simultaneously performing multiple single-step prediction tasks and achieving knowledge sharing between different tasks. Therefore, both multi-modal and multi-task learning can improve the learning performance of the model again. The wind power multi-step prediction model based on M2TNet will improve the prediction accuracy of existing models and reduce their time complexity. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1 FIG. is a schematic flowchart of a method for multi-step prediction of wind power based on a Transformer network with multi-modal and multi-task capabilities provided by an embodiment of the present invention;

[0027] Figure 2 FIG. is a schematic diagram of a deep M2TNet prediction model provided by an embodiment of the present invention;

[0028] Figure 3 FIG. is a schematic diagram of a Transformer feature extractor provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0029] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0030] As Figure 1The following is a schematic flow diagram of a wind power multi-step prediction method based on a multi-modal multi-task Transformer network (M2TNet) provided by an embodiment of the present invention, including the following steps:

[0031] (1) Design a deep M2TNet prediction model;

[0032] As Figure 2 shown, this model has a feature extraction layer, a feature fusion layer, and a prediction terminal layer. Among them, the feature extraction layer consists of S feature extractors, such as Figure 3 shown, the Transformer unit is used as the feature extractor to form; the feature fusion layer is a layer of ordinary fully connected network; the prediction terminal layer is a regression layer.

[0033] (2) Collect K types of multi-source heterogeneous data and use Maximal Information Coefficient (MIC) analysis for variable selection, and use S (S≤K) types of preferred variables as the input of the M2TNet model;

[0034] (3) Normalize the S types of preferred variables to obtain sample data, map the sample data to between [0,1], and divide the entire sample data into a training set and a test set according to a ratio of 8:2;

[0035] (4) Determine the hyperparameters of the M2TNet model by means of artificial experience and grid search and in combination with the effect of model training;

[0036] (5) Save the parameter model with the best training effect, and use the data of the test set for multi-step prediction of wind power;

[0037] (6) Denormalize the prediction result to obtain the wind power prediction value, and analyze the prediction result in combination with evaluation indicators.

[0038] According to the above technical solution, the specific design idea of step (1) is as follows:

[0039] (1.1) First, a mathematical model for wind power prediction needs to be established. Assume that the research object is a certain wind farm, and given S types of multi-source heterogeneous historical data, the wind power sequence y t =(y t+p , y t+p+1 ,…, y t+q ) within a certain future time period at time t is predicted. Among them, y t is the wind power value at time t, and p and q are the minimum and maximum prediction steps respectively. This prediction model can be described as:

[0040]

[0041] In Equation (1): x t represents the input vector, θ and f(·) respectively represent the parameter vector and the implicit function of the prediction model, represents the sub-vector corresponding to historical feature i at time t.

[0042] can be defined as:

[0043]

[0044] In Equation (2): n i represents the dimension of, that is, the sequence length of historical feature i, is the value of feature i at time t.

[0045] Multi-source heterogeneous data generally includes wind power and numerical weather prediction (NWP) features (such as: wind direction, wind speed, temperature, atmospheric pressure, air density, etc.). These data have different physical natures, come from different measurement subjects, and also vary greatly in dimension and numerical range, belonging to typical heterogeneous data.

[0046] When using multi-source heterogeneous data for wind power prediction, two problems need to be addressed. The first is how to respect the differences between different data and handle each historical feature specifically. The second is how to grasp the unity between different historical features so that they can complement each other and jointly provide useful information for wind power prediction. The present invention proposes a three-stage modeling strategy to solve the above two problems. In the first stage, multiple feature extractors are used to extract high-order features from various historical features respectively. In the second stage, the high-order features from different historical features are fused to form a unified feature containing richer information. In the third stage, the wind power sequence is predicted based on the unified feature. The wind power prediction problem can be further described as:

[0047]

[0048] In Equation (3), h t represents the unified feature, represents the high-order feature corresponding to historical feature i, g i (·) represents the implicit function of the feature extractor corresponding to historical feature i in the first stage, and v(·) and u(·) respectively represent the implicit functions of the second and third stage functions.

[0049] The above three-stage modeling strategy integrates various historical features in the same learning framework and has the potential to provide richer information for wind power prediction.

[0050] (1.2) In the first stage of the three-stage modeling strategy proposed in (1.1), the Transformer unit shown in Figure 3 is used as a feature extractor. There is only one encoder in the Transformer unit, and this encoder is composed of an input layer, a position encoding layer, and n identical stacked encoding layers. The input layer maps the input historical feature i sequence to a vector with a dimension of d model through a fully connected network. This step is crucial for the model to adopt the multi-head attention mechanism. The position encoding layer encodes the sequential information in the time series data through sine and cosine functions to obtain a position encoding vector, and then adds the elements of the d model -dimensional input vector and the position encoding vector to obtain a time series vector with position information. This vector is fed into n encoder layers. Each encoder layer consists of two sub-layers, namely a self-attention sub-layer and a fully connected feed-forward sub-layer, and each sub-layer is followed by a normalization layer. Finally, a d model -dimensional vector generated by the last encoding layer undergoes a linear mapping operation to output the high-order features of the historical feature i.

[0051] (1.3) In the M2TNet network shown in Figure 2 , a multi-branch structure with multiple Transformer units in parallel is adopted. Each branch can process each type of heterogeneous data separately. Such a design is beneficial to the flexible configuration of the entire model structure, reduces the number of network connections, and is expected to capture the complex laws hidden in various heterogeneous data. For the second stage of the three-stage modeling strategy proposed in (1.1), a multi-modal learning strategy is adopted, which can integrate the information between different modalities output from each Transformer unit to form higher-level features that are more conducive to the realization of complex tasks, thereby helping the learning machine achieve better performance.

[0052] (1.4) From the perspective of machine learning, single-step prediction can be regarded as single-task learning, and multi-step prediction can be regarded as multi-task learning, that is, multiple single-step prediction tasks. For the third stage of the three-stage modeling strategy proposed in (1.1), a multi-task learning strategy is adopted. For the values at each sampling moment of the sequence y t to be predicted, there is a certain correlation between them. Especially for the values that are closer in time, the degree of correlation may be higher. Integrating multiple single-value prediction tasks (single-step wind power prediction) into a sequence prediction task (multi-step wind power prediction) may achieve knowledge sharing between different tasks, thereby improving the prediction performance. Therefore, the output end of the M2TNet model adopts a multi-dimensional output structure to simultaneously output the values at each moment in the time series.

[0053] According to the above technical solution, the specific method of step (2) is as follows:

[0054] The mutual information is used to measure the correlation between the wind power and the historical variables of other heterogeneous data, and the MIC value is calculated and sorted in descending order. The top S variables are the selected variables.

[0055] According to the above technical solution, the specific design idea of step (4) is as follows:

[0056] To reduce the parameter search space of M2TNet, the following simplifications are taken: ① All transformers adopt the same structural parameters; ② The number of units in each hidden layer of the transformer remains the same. After simplification, there are 6 hyperparameters to be determined in M2TNet: ① The number of encoding layers of each transformer, denoted as n, selected from Ψ n ={2, 3, 4, 5}; ② The number of neurons in each hidden layer, denoted as h, selected from the set Ψ h ={16, 32, 64, 128}; ③ The size of the learning rate, denoted as l, selected from the set Ψ l ={0.001, 0.005, 0.01, 0.015, 0.02, 0.025, 0.03}; ④ The size of Epochs, denoted as e, selected from the set Ψ e ={8, 16, 32, 64, 128}; ⑤ The size of Batch Size, denoted as b, selected from the set Ψ b ={16, 32, 64, 128}; ⑥ The size of Dropout, denoted as d, selected from the set Ψ d ={0.1, 0.3, 0.5, 0.7, 0.9}. In addition, the optimizer of M2TNet is RMSProp, and its loss function is mse. The machine grid search space is set based on manual experience. Once the setting work is completed, machine automatic optimization can be realized, thus reducing the manual workload.

[0057] According to the above technical solution, the model performance evaluation indexes of step (6) are as follows:

[0058] The evaluation indexes of the ultra-short-term wind power multi-step prediction model of the present invention are from two aspects: the error of the model prediction result and the model time complexity. From the perspective of the model time complexity, the model training time and the test time are used as evaluation indexes. From the perspective of the error of the model prediction result, the root mean square error (RMSE) E RMSE and the mean absolute error (MAE) E MAE are selected as evaluation indexes, and the specific definitions are as follows:

[0059]

[0060]

[0061] Where: y k and are respectively the actual and predicted wind power values at the k-th moment within the error statistical time interval, n is the total number of time periods in this interval, and C N is the installed capacity of the wind farm to be predicted.

[0062] A multi-step wind power prediction method based on a multi-modal multi-task Transformer network according to the present invention. The maximum information coefficient is used to perform variable selection on multi-source heterogeneous data, so as to achieve the purpose of controlling the model scale, reducing the computational complexity, and improving the model performance. A three-stage modeling strategy is proposed to solve the problem of multi-step wind power prediction using multi-source heterogeneous data. A deep M2TNet model is designed, which has a feature extraction layer, a feature fusion layer, and a prediction terminal layer. Among them, the feature extraction layer adopts a multi-branch structure with multiple Transformer units, and this structural design enables each Transformer unit to specifically process each type of heterogeneous data separately. In addition, the advantages of Transformer in mining feature information and computational efficiency are conducive to improving the performance of the overall model. The multi-step prediction problem is transformed into multiple single-step prediction tasks. In the feature fusion layer and the prediction terminal layer, the M2TNet model adopts multi-modal and multi-task learning strategies. Multi-modal learning is conducive to integrating the information between different modalities output by multiple Transformer units to form high-level features for realizing complex tasks. Multi-task learning is conducive to simultaneously performing multiple single-step prediction tasks and realizing knowledge sharing between different tasks. Therefore, both multi-modal and multi-task learning can improve the learning performance of the model again. The wind power multi-step prediction model based on M2TNet will improve the prediction accuracy of the existing model and reduce its time complexity.

[0063] It should be noted that according to the needs of implementation, each step / component described in this application can be split into more steps / components, or two or more steps / components or partial operations of steps / components can be combined into new steps / components to achieve the purpose of the present invention.

[0064] Those skilled in the art can easily understand that the above are only the preferred embodiments of the present invention, and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present invention should be included in the protection scope of the present invention.

Claims

1. A multi-step wind power prediction method based on a Transformer network for multi-modal multi-task, characterized in that, Including: The M2TNet prediction model with design depth; Collect K kinds of multi-source heterogeneous data and use the maximum information coefficient analysis for variable selection, and use S kinds of optimal variables as the input of the M2TNet prediction model, where S ≤ K; Normalize the S kinds of optimal variables to obtain sample data, and divide the sample data into a training set and a test set; Determine the hyperparameters of the M2TNet model by means of artificial experience and grid search and in combination with the training effect of the model; Save the parameter model with the best training effect, and use the data of the test set for multi-step prediction of wind power; Denormalize the prediction result to obtain the wind power prediction value, and analyze the prediction result in combination with evaluation indicators; The M2TNet prediction model has a feature extraction layer, a feature fusion layer and a prediction terminal layer; among them, the feature extraction layer adopts a multi-branch structure with multiple Transformer units in parallel; the feature fusion layer is an ordinary fully connected network; the prediction terminal layer is a regression layer; The M2TNet prediction model with design depth, including: The first stage: Use multiple feature extractors to extract high-order features from various historical features respectively, where the Transformer unit is used as the feature extractor; The second stage: Adopt a multi-modal learning strategy to fuse the high-order features from different historical features to form a unified feature containing richer information; The third stage: Adopt a multi-task learning strategy to predict the wind power sequence based on the unified feature; Determine the hyperparameters of the M2TNet model by means of artificial experience and grid search and in combination with the training effect of the model, including: All Transformers adopt the same structural parameters; the number of units in each hidden layer of the Transformer remains the same, so the hyperparameters to be determined in M2TNet include: the number of encoding layers of each Transformer, the number of neurons in each hidden layer, the size of the learning rate, the size of Epochs, the size of Batch Size, and the size of Dropout; the optimizer of M2TNet adopts RMSProp, and its loss function is mse; Determine the hyperparameters of the M2TNet model by means of artificial experience and grid search and in combination with the training effect of the model.

2. The method according to claim 1, characterized in that, The M2TNet prediction model is described as: Among them, y t is the wind power sequence within a certain future time period at time t, and x t represents the input vector, and f(·) represents the implicit function of the prediction model. represents the sub-vector corresponding to historical feature i at time t, and h t represents the unified feature, and r t (i) represents the high-order feature corresponding to historical feature i, and g i (·) represents the implicit function of the feature extractor corresponding to historical feature i (i = 1, 2,..., S) in the first stage, v(·) and u(·) respectively represent the implicit functions of the functions in the second and third stages, and S represents the number of historical features.

3. The method according to claim 1 or 2, characterized in that, Collect K kinds of multi-source heterogeneous data and use the maximum information coefficient analysis for variable selection, and use S kinds of optimal variables as the input of the M2TNet prediction model, including: Use mutual information to measure the correlation between wind power and other heterogeneous data historical variables and calculate the MIC value, sort them in descending order, and use the first S variables as the optimal variables.

4. The method according to claim 3, characterized in that, By and are used as evaluation indicators, where y k and are the actual and predicted wind power values at the k-th moment within the error statistical time interval respectively, n is the total number of time periods in this interval, and C N is the installed capacity of the wind farm to be predicted.

5. A computer-readable storage medium, on which a computer program is stored, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Wind power prediction method and system for optimizing deep Transformer network

    CN112653142A