Multi-task learning-based medium-and-long-term power prediction method for multi-stage wind power cluster

Through the Transformer structure of the multi-task learning model and sparse attention mechanism, the problem of insufficient accuracy and efficiency of long-term power prediction in wind power clusters is solved, coordinated prediction between wind power plants is realized, prediction accuracy and stability are improved, and strong support for the stable operation of the power grid.

CN120454023APending Publication Date: 2025-08-08CHINA THREE GORGES UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510509572.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

Existing wind power power prediction technologies have problems of insufficient accuracy and efficiency in wind power clustering and medium- and long-term prediction, especially ignoring the synergy and spatial and temporal correlation between wind power plants, and traditional rolling prediction strategies lack flexibility, which affects prediction performance.

Method used

The Transformer structure using a multi-task learning model combined with a sparse attention mechanism is adopted. Through the shared information layer and cluster-level expert module working together, the simultaneous prediction of the power of multi-wind power stations is realized, and an independent loss function is set during the rolling prediction process, the model weight is optimized, and the prediction accuracy and stability are improved.

Benefits of technology

It significantly improves the accuracy and stability of long-term power prediction in wind power clusters, provides strong support for the optimal scheduling and operation of wind power station groups, and improves the power generation efficiency of wind farms and the stability of the power grid.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120454023A_ABST
    Figure CN120454023A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-task learning-based medium-and-long-term power prediction method for a multi-stage wind power cluster, and belongs to the technical field of power system scheduling. The invention aims to solve the problem of insufficient precision and efficiency in long-term power prediction in a wind power cluster and overcome the limitations of insufficient utilization of renewable energy sources, poor system load balance and the like in the prior art. A multi-task learning model is applied to the field for the first time, and simultaneous prediction of power of a plurality of wind power stations is realized through cooperation of a shared information layer and a cluster-level expert module; a Transform structure based on a sparse attention mechanism is adopted, long-distance dependence between wind power stations is captured, and the feature extraction capability is improved; a splicing type rolling prediction strategy is provided, the integrity of historical data is kept, an independent loss function is set for each rolling stage, and prediction errors are accurately measured and optimized; according to the method, the accuracy and stability of long-term power prediction in the wind power cluster are remarkably improved, powerful support is provided for optimal scheduling and operation of the wind power station group, and the method has wide application value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of renewable energy power generation prediction technology, and in particular to a multi-stage wind power cluster medium- and long-term power prediction method based on multi-task learning. Background Art

[0002] With the continuous optimization of the global energy mix and the growing acceptance of sustainable development, wind power, as a clean, renewable energy source, is experiencing rapid growth in both installed capacity and power generation worldwide. However, wind power generation is inherently intermittent and volatile, and its power output is significantly affected by meteorological conditions, posing significant challenges to the stable operation and dispatch of power systems. Therefore, accurate wind power forecasting has become a crucial tool for ensuring grid security and improving wind power absorption capacity.

[0003] In the field of wind power forecasting technology, traditional methods mainly focus on short-term or ultra-short-term forecasts for a single wind farm. Although these methods can meet the needs of short-term scheduling to a certain extent, their limitations are becoming increasingly prominent in the face of the development trend of wind power clustering and base-based development. First, the power forecast of a single wind farm fails to fully consider the synergy between multiple wind power stations in the region, ignoring the spatial correlation and time lag between wind farms, resulting in limited forecast accuracy. Second, traditional methods often use a single forecast model, which makes it difficult to simultaneously capture the changing patterns of wind power at different time scales, limiting further improvement in forecast accuracy.

[0004] Although there have been some explorations in the existing technology for the medium- and long-term prediction of wind power, there are still many shortcomings. For example, CN116454875A discloses a method and system for regional wind farm medium-term power probability prediction based on cluster division. This method clusters the wind farms using a subtractive clustering algorithm, and uses the LightGBM (Light Gradient Boosting Machine) algorithm to establish a medium-term power prediction model for each cluster. Finally, the probability density distribution of the power prediction error is obtained through the non-parametric kernel density estimation method. However, this method may not fully consider the complex correlations between wind farms during the cluster division process, resulting in inaccurate cluster division results, which in turn affects the prediction accuracy. In addition, this method may face challenges in computational complexity and efficiency when processing large-scale wind power cluster data.

[0005] On the other hand, CN118199046A proposes a method for predicting the power of multiple wind turbines based on a twin neural network. This method uses a twin neural network to capture the complex correlations between different wind turbines and automatically adjusts the weight coefficients of the loss function of each wind turbine through homoscedastic uncertainty, thereby achieving accurate prediction of the power of multiple wind turbines. Although this method performs well in capturing the correlations between wind turbines, it is mainly suitable for short-term or ultra-short-term predictions, and its effectiveness for medium- and long-term predictions needs further verification. At the same time, this method may require a lot of computing resources and time costs when constructing and training the twin neural network.

[0006] More critically, existing wind power forecasting technologies still have shortcomings in rolling forecast strategies. Traditional rolling forecasting methods typically use a fixed-length historical window. As the number of rolling cycles increases, some historical samples are overwritten by the previous round of forecasts, resulting in a decrease in long-term forecast accuracy. More seriously, these methods share the same model weights throughout the entire rolling forecast process, failing to fully account for the varying reliance on historical information at different stages, hindering further improvements in forecast performance.

[0007] In summary, existing wind power forecasting technologies have many shortcomings when addressing the needs of wind power clustering and medium- and long-term forecasting, especially in terms of cluster division accuracy, forecasting model applicability, rolling forecasting strategies, computational efficiency, and resource consumption. This invention is proposed in this context, aiming to achieve accurate forecasting of the medium- and long-term power of wind power clusters by introducing a multi-task learning model and a multi-stage rolling forecasting strategy. It also optimizes the cluster division method and forecasting model structure, improving forecast accuracy and efficiency, and providing strong support for the optimized scheduling and operation of wind power plant clusters. Summary of the Invention

[0008] The technical problem to be solved by the present invention is to provide a multi-stage medium- and long-term power prediction method for wind power clusters based on multi-task learning, so as to solve the problems of insufficient accuracy and efficiency in the field of wind power prediction, especially in the medium- and long-term power prediction of wind power clusters. Specifically, the short-term or ultra-short-term prediction methods of traditional single wind farms can no longer meet the requirements of modern power grid dispatching for wind power prediction accuracy and timeliness; the existing technology fails to fully consider the synergy between multiple wind power stations in the region, ignores the spatial correlation and time lag between wind farms, resulting in limited prediction accuracy; at the same time, the traditional method lacks flexibility in rolling prediction strategies, fails to fully consider the differences in dependence on historical information at different stages, and affects the further improvement of prediction performance.

[0009] To solve the above technical problems, the technical solution adopted by the present invention is a multi-stage wind power cluster medium- and long-term power prediction method based on multi-task learning, comprising the following steps: Step 1: Identify outliers and fill missing values in the power station’s historical data; Step 2: Based on the power, wind speed and other characteristics, use The clustering algorithm divides the power plants into clusters; Step 3: Input the preprocessed and clustered data into the proposed model to complete the medium- and long-term power forecast of each power station.

[0010] In a preferred solution, the specific steps in Step 1 include: Step 1.1: Standardize the historical wind speed and power data of multiple wind power stations so that data of different dimensions have the same scale; Step 1.2: Build an isolation forest model, determine outliers by calculating the average path length of samples in the isolation forest, and set a threshold based on the anomaly score to identify and remove abnormal data; Step 1.3: For the missing data identified, the adjacent value interpolation method is used to fill in the missing data to ensure the integrity and consistency of the data.

[0011] In the preferred solution, the outlier determination in Step 1.2 is performed using the isolation forest algorithm, as follows: (1) (2) (3) Where, For the The harmonic number of positive integers is used to calculate the expected path length; is the normalization factor, Indicates the total number of samples in the dataset; Representation sample average path length in isolated forests; Representative samples expected path lengths in multiple isolated trees; Representative samples Anomaly score.

[0012] In the preferred solution, the step 2 is adopted The clustering algorithm divides the power plants into clusters. Specifically, it divides wind power plants with similar characteristics into multiple clusters. The specific steps are as follows: Step 2.1: Standardize the historical wind speed and power data of multiple wind power stations; Step 2.2: Traverse multiple Value, for each Value Application The algorithm calculates the corresponding intra-cluster squared error and plots curve, select the elbow inflection point as the optimal number of clusters; Step 2.3: Initialization Clustering, select the initial cluster center, calculate the Euclidean distance of each wind power station to all cluster centers, and assign it to the nearest cluster; Step 2.4: Calculate and update the center of each cluster and reallocate wind power stations to the new cluster center until convergence.

[0013] In the preferred embodiment, the The clustering algorithm calculates the intra-cluster squared error as follows: (4) (5) (6) Where, represents the number of clusters, Indicates the The feature vector of the sample points, Indicates belonging to a cluster The sample set, is the cluster center, Representation sample With cluster center The square of the Euclidean distance, represents the time step, Indicates that in the wind speed sequence, The cluster in The cluster center coordinates on the feature dimension, Indicates that in the power sequence, the The cluster in The cluster center coordinates on the feature dimension.

[0014] In a preferred solution, the specific steps of Step 3 include: Step 3.1: Build a multi-task learning model, including a shared information layer and an expert module. The shared information layer is used to extract global features between wind power stations. The output is residually connected with the initial information of each cluster and then input into the corresponding expert module. Step 3.2: The expert module consists of multiple prediction layers, each of which is responsible for power and wind speed prediction at different stages and has an independent loss function. Step 3.3: Combine the prediction results of each stage through rolling prediction and comprehensively optimize the model weights.

[0015] In the preferred solution, the shared information layer in Step 3.1 adopts a Transformer structure based on a sparse attention mechanism to model the collaborative relationship between wind power stations and extract global features between each wind power station; the normalization layer of the Transformer structure is replaced by a dynamic hyperbolic tangent activation function (DyT, Dynamic Tanh) layer to enhance the expressive power of the model.

[0016] In the preferred solution, the input of the expert module in Step 3.2 is spliced by the input and output of the previous layer to make full use of historical data. Each prediction layer is responsible for power and wind speed prediction at different stages, and is equipped with an independent loss function to measure the prediction error.

[0017] In the preferred solution, in the step of splicing the prediction results of each stage through a rolling prediction method, an independent loss function is set for each rolling stage, and the loss weight of each stage can be adjusted according to actual conditions. The model parameters are adjusted according to the gradient of the total loss function through the back propagation algorithm to optimize the prediction performance.

[0018] In a preferred solution, the total loss function is the weighted sum of all main task loss functions, the loss function of each main task is the weighted sum of its subtask loss functions, and the model parameters are updated by minimizing the total loss function.

[0019] The multi-stage wind power cluster medium- and long-term power prediction method based on multi-task learning provided by the present invention has the following beneficial effects: 1. The present invention solves the problem of insufficient accuracy and efficiency in the field of wind power forecasting, especially in the medium- and long-term power forecasting of wind power clusters. Specifically, the traditional short-term or ultra-short-term forecasting methods for a single wind farm are no longer able to meet the requirements of modern power grid dispatching for wind power forecast accuracy and timeliness. Existing technologies fail to fully consider the synergy between multiple wind power stations in a region, ignoring the spatial correlation and time lag between wind farms, resulting in limited forecast accuracy. At the same time, traditional methods lack flexibility in rolling forecasting strategies and fail to fully consider the differences in reliance on historical information at different stages, which affects the further improvement of forecast performance.

[0020] 2. The present invention uses the isolation forest algorithm to identify and eliminate outliers in the historical data of the wind power station, and uses adjacent values to fill in missing data to ensure the integrity and consistency of the data.

[0021] 3. The present invention combines the isolation forest algorithm and Clustering algorithms are used to preprocess and divide wind power station data, optimize data structure, and improve the accuracy and adaptability of prediction models.

[0022] 4. This invention is based on the power and wind speed characteristics of the wind power station and adopts The clustering algorithm divides wind farms with similar characteristics into multiple clusters so that the model can learn and extract the collaborative characteristics of similar wind farms.

[0023] 5. The present invention adopts a rolling prediction method to splice the prediction results of each stage, and optimizes the model weight by integrating the task losses of each cluster to ensure that the predictions of different clusters at different stages are both independent and efficient.

[0024] 6. The optimized rolling forecast strategy of the present invention uses a splicing method to maintain the integrity of historical data and avoid error accumulation. At the same time, it breaks down the medium- and long-term forecast tasks into multiple short-term tasks, and improves the stability and reliability of the forecast through gradual rolling calculations.

[0025] 7. The present invention sets an independent loss function for each rolling stage and can adjust the loss weight of each stage according to actual conditions to further optimize the prediction accuracy and stability.

[0026] 8. The information processing and prediction output of the model proposed in this invention are based on a combination of a Transformer neural network with a sparse attention mechanism and a linear layer, ensuring efficient feature extraction and improving prediction accuracy.

[0027] 9. The present invention adopts a Transformer structure based on a sparse attention mechanism in the shared information layer and cluster-level expert module to capture long-distance dependencies between wind power stations and improve feature extraction capabilities.

[0028] 10. The present invention introduces a sparse attention mechanism into the Transformer structure, focusing on key wind power stations, improving computational efficiency and reducing model complexity, significantly improving the accuracy and stability of medium- and long-term power forecasts of wind power clusters, and providing strong support for the optimized scheduling and operation of wind power station clusters.

[0029] 11. The present invention extracts the collaborative relationship between wind power stations through a shared information layer, effectively captures the correlation between multiple wind power stations in the same area, and improves the overall power prediction accuracy.

[0030] 12. The spliced rolling prediction strategy and independent loss function design of the present invention improve the stability and reliability of the prediction. By adjusting the loss weights of each stage, the prediction errors of different stages can be accurately measured and optimized.

[0031] 13. The present invention introduces the dynamic hyperbolic tangent activation function (DyT) into the normalization layer of the Transformer structure. By introducing learnable parameters, the expressiveness and generalization capabilities of the model are enhanced, enabling the model to better adapt to the complex and changeable wind power environment.

[0032] 14. In practical applications, the method of the present invention can more accurately predict the medium- and long-term power output of wind power clusters, provide important guarantees for the stable operation of the power system and the efficient absorption of wind power, improve the power generation efficiency of wind farms and the stability of the power grid, and provide strong support for the optimized scheduling and operation of wind farm groups.

[0033] 15. The present invention combines data preprocessing with clustering methods to optimize data structure and improve the accuracy and adaptability of the prediction model.

[0034] 16. The present invention achieves simultaneous prediction of the power of multiple wind power stations by introducing a multi-task learning model and a sparse attention mechanism, significantly improving the prediction efficiency and accuracy. The method of the present invention is not only applicable to wind power prediction, but can also be extended to other new energy power prediction fields such as photovoltaic power prediction, and has broad application value. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] The present invention will be further described below with reference to the accompanying drawings and embodiments: Figure 1 This is the overall framework diagram for the operation of the method of the present invention; Figure 2 This is a prediction flow chart of Example 3 of the present invention; Figure 3 This is a flowchart of outlier identification and missing value filling in Example 3 of the present invention; Figure 4 This is a flow chart of clustering a wind power station cluster using the K-Means algorithm according to embodiment 3 of the present invention; Figure 5 This is a diagram showing the operating framework of the multi-task learning model in Example 3 of the present invention. DETAILED DESCRIPTION The technical solutions of the present invention are further described below with reference to the embodiments and accompanying drawings: Example 1 like Figure 1 As shown, this embodiment provides a multi-stage wind power cluster medium- and long-term power prediction method based on multi-task learning, and the specific steps are as follows: Step 1: Data Preprocessing (1) Data standardization: Eliminate the impact of data of different dimensions on subsequent analysis and make the data have a unified scale.

[0036] The historical wind speed and power data of multiple wind farms are normalized using the Z-score normalization method. For example, if there is wind speed data from five wind farms, the mean and standard deviation of the wind speed data for these five wind farms are first calculated, and then the wind speed data for each wind farm is normalized.

[0037] (2) Outlier identification and removal: Identify and remove outliers in the data set to improve data quality.

[0038] Construct an isolation forest model, setting the number of decision trees in the forest (e.g., 100), the random seed (e.g., 42), and the subsample size (e.g., 256); For each decision tree, a feature (such as wind speed or power) is randomly selected from the sample set, and a cutoff value is randomly selected within the value range of the feature for partitioning. The above steps are recursively performed until a single sample cannot be split any further (i.e., a leaf node contains only one sample), forming a complete isolated tree structure. Calculate the average path length of each sample in the isolation forest, that is, the average depth of the sample; because outliers are more easily isolated, their path length is usually shorter. The specific calculation formula is as follows: (1) (2) (3) Calculate the anomaly score using a predefined anomaly scoring function and set an anomaly score threshold (e.g., 0.6) to identify and remove abnormal data. For example, in the isolation forest model, the average path length of a certain sample is 1 and the anomaly score is 0.95, which exceeds the set threshold of 0.6. Therefore, it is identified as an outlier and is removed.

[0039] (3) Missing value filling: Fill in the missing values in the data set to ensure the integrity and consistency of the data.

[0040] For identified missing data, we use the adjacent value interpolation method to fill in the missing data. Specifically, we use the average of the two valid values before and after the missing value as the filling value. For example, if the wind speed data of a wind power station is missing on a certain day, we use the average of the wind speed data from the previous and next day as the filling value.

[0041] Step 2: Wind power station clustering (1) Data standardization: Same as data standardization in step 1 to ensure the consistency of cluster analysis.

[0042] The historical wind speed and power data of multiple wind power stations are standardized using the Z-score normalization method.

[0043] (2) Determine the optimal number of clusters: Find the optimal number of clusters to balance the similarity within the cluster and the difference between the clusters.

[0044] Using the elbow rule to determine the optimal number of clusters , traverse different value (such as from 2 to 10), and calculate the corresponding intra-cluster squared error ( ), the specific calculation is as follows: (4) (5) (6) draw Curve, select the elbow inflection point as the optimal number of clusters. The elbow inflection point usually appears at the position where the curve begins to flatten. Through calculation, it is found that when hour, The curve has an elbow, so choose as the optimal number of clusters.

[0045] (3) Clustering: Divide wind power plants into multiple clusters for subsequent multi-task learning.

[0046] Random selection Sample points are used as the initial cluster centers; Calculate the Euclidean distance of each wind power station sample to all cluster centers and assign it to the nearest cluster; Recalculate the cluster center of the new cluster (i.e., the mean of all samples in the cluster) and reallocate wind power stations to the new cluster center based on the minimum distance; Repeat the above steps until convergence (i.e., the change in the old and new cluster centers is less than a set threshold, such as 10−4, or the ownership of the power station no longer changes). After multiple iterations, the five wind power plants were divided into three clusters, and the wind power plants in each cluster had similar wind speed and power characteristics.

[0047] Step 3: Build a multi-task learning model (1) Shared information layer: Extract shared information between wind power stations to achieve information sharing across wind power stations.

[0048] All preprocessed data are input into the shared information layer of the multi-task learning model. This layer adopts the Transformer structure based on the sparse attention mechanism to extract global features between wind power stations through operations on query, key and value matrices.

[0049] (2) Residual connection and expert module: Combine shared information and the initial information of each cluster to make predictions and improve prediction accuracy.

[0050] The output of the shared information layer is residually connected with the initial information of each cluster and input into the corresponding expert module; Each wind power cluster corresponds to an expert module, which contains multiple prediction layers (such as 3 layers). Each layer is responsible for power and wind speed prediction at different stages and has an independent loss function (such as root mean square error). , Root MeanSquare Error), as follows: (7) Where, is the sample size, is the true value, is the predicted value.

[0051] The shared information layer extracts wind speed correlation features between wind farms and performs a residual connection with the initial wind speed information of a cluster before inputting them into the corresponding expert module. The first layer of the expert module is responsible for predicting wind speed and power in the short term (e.g., one hour), the second layer for the medium term (e.g., one day), and the third layer for the long term (e.g., one week).

[0052] (3) Rolling prediction and optimization: The prediction results of each stage are spliced together through rolling prediction to optimize the model weights.

[0053] A rolling prediction method is used to gradually splice the prediction results of each stage. An independent loss function is set for each rolling stage, and the loss weight of each stage can be adjusted according to the actual situation. The details are as follows: The loss function of each main task is the weighted sum of the loss functions of its subtasks, that is: (8) Where, It is the main task The number of subtasks under It is The loss function of each subtask is is the weight of the corresponding subtask loss; The total loss function is the weighted sum of all main task loss functions, that is: (9) Where, is the number of main tasks, It is The loss function of the main task, is the weight of the main task loss; Through the back-propagation algorithm, the model parameters (such as the weights of the shared information layer and the expert module) are adjusted according to the gradient of the total loss function to optimize the prediction performance.

[0054] During the rolling forecast process, the prediction results of the first stage serve as one of the inputs of the second stage, which in turn serves as one of the inputs of the third stage, and so on. The weight of the loss function at each stage can be adjusted based on the size of the prediction error to improve the overall prediction accuracy.

[0055] This embodiment extracts collaborative relationships between wind farms through a shared information layer, effectively improving the accuracy and stability of medium- and long-term power forecasts for wind farm clusters. The optimization of the rolling forecast strategy ensures the integrity of historical data and avoids error accumulation. The structural design of the multi-task learning model enables information sharing and collaborative forecasting across wind farms, further improving forecast performance.

[0056] Example 2 In another preferred embodiment, based on Example 1, this embodiment further optimizes the structure and training process of the multi-task learning model. The specific steps are as follows: Steps 1 to 2: Same as in Example 1, data preprocessing and wind power station clustering are performed.

[0057] Step 3 (Optimized version): Build an optimized multi-task learning model (1) Dimension expansion layer: Before the shared information layer, a dimension expansion layer is added. Assume that the input data dimensions are (N, T, d), where N is the number of wind farms, T is the time step, and d is the feature dimension. The dimension expansion layer maps the data to a higher-dimensional space (N, T, d′) to enhance feature representation.

[0058] (2) Shared information layer: A Transformer encoder based on sparse attention mechanism is used as the shared information layer, and its core calculation includes operations on query, key and value matrices to extract global features between wind power stations.

[0059] (3) Expert module optimization: The input of each expert module is composed of the concatenation of the input and output of the previous layer to fully utilize historical data. The expert module uses a deeper network structure with more prediction layers to improve the model's predictive capabilities.

[0060] (4) Loss function and training optimization: Each prediction layer has an independent loss function, using the root mean square error ( ) as a loss function to measure the error between the predicted value and the true value; During model training, a dynamic learning rate adjustment strategy is adopted to automatically adjust the learning rate according to the training progress to accelerate model convergence and improve training results; Early Stopping is introduced. When the loss function on the validation set no longer decreases within a certain number of rounds, training is terminated early to prevent overfitting.

[0061] This example further improves the model's prediction accuracy and generalization capabilities by adding a dimensionality expansion layer and optimizing the expert module structure. The introduction of dynamic learning rate adjustment and early stopping effectively enhances the model's training efficiency and stability.

[0062] Example 3 In another preferred embodiment, based on embodiments 1 and 2, this embodiment provides a multi-stage wind power cluster medium- and long-term power prediction method based on multi-task learning, such as Figure 2 , the specific steps are as follows: S1. Identify outliers and fill missing values in the power station’s historical data; S2, according to the characteristics of power, wind speed, etc., use The clustering algorithm divides the power plants into clusters; S3. Input the pre-processed and clustered data into the proposed model to complete the medium- and long-term power forecast of each power station.

[0063] Specifically, this method first uses the isolation forest algorithm to identify outliers in the historical data of wind power stations and remove abnormal samples. At the same time, the adjacent value interpolation method is used to fill in the missing data to ensure the integrity and consistency of the data. A clustering algorithm uses power and wind speed as key features to classify wind farms, grouping those with similar characteristics into the same cluster. Each cluster corresponds to a primary task in the multi-task learning model. Finally, the preprocessed data is fed into the multi-task learning model, which, combined with a Transformer architecture using a sparse attention mechanism, extracts correlations between wind farms and achieves accurate medium- and long-term power forecasting. This method effectively improves forecast accuracy and stability, overcoming issues such as information loss, rolling error accumulation, and inadequate modeling of wind farm synergy relationships in traditional methods for medium- and long-term forecasting.

[0064] The above steps S1-S3 are described in detail below: See Figure 3 As shown, Figure 3 This is a schematic diagram of outlier identification and missing value filling. The present invention uses the isolation forest algorithm to detect outliers in the historical wind speed and power data of the wind power station, and uses the adjacent value interpolation method to fill the missing data to ensure data integrity and consistency. The specific steps are as follows: S1.1. Historical wind speed of each wind power station and power The data is standardized to eliminate the impact of data of different dimensions on anomaly detection results. The standard score (Z-score) normalization method is used in the standardization process: (10) Where, is the original data, and are the mean and standard deviation of the feature, respectively. is the standardized data; S1.2. Set the number of decision trees in the forest , select random seed, set subsample size ; S1.3. For each tree, randomly select a feature from the sample set each time (such as wind speed or power ), and randomly select a cutoff value within the value range of the feature Divide and recursively execute the above steps until a single sample cannot be split any further (i.e., a leaf node contains only one sample), forming a complete isolated tree structure; S1.4. For each tree, randomly select a feature from the sample set each time (such as wind speed or power ), and randomly select a cutoff value within the value range of the feature Divide and recursively execute the above steps until a single sample cannot be split any further (i.e., a leaf node contains only one sample), forming a complete isolated tree structure; S1.5. For each sample , calculate its average path length , which is the average depth of the sample in the isolation forest. Since outliers are more easily isolated, their path length is usually short. Anomaly scoring function: (11) Where, is the normalization factor, which is approximately: ,in For the harmonic number, which is Representative samples The anomaly score of is in the range of [0, 1]. The closer the score is to 1, the more likely the sample is an outlier. S1.5. After calculating the anomaly score, set the anomaly score threshold ,like , it is determined to be an outlier and removed from the data set. For missing values, Filling is done by interpolation of the nearest neighbor mean; S1.6. Organize and obtain a preprocessed data set that has removed outliers and filled in missing data.

[0065] See Figure 4 As shown, Figure 4This is a schematic diagram of the wind power station clustering process. Clustering algorithm is used to classify multiple wind power stations based on wind speed. or power As a feature quantity, wind power stations are divided into several clusters to improve the accuracy of wind power prediction. The specific steps are as follows: S2.1. Historical wind speed of each wind power station or power The data is normalized to eliminate the impact of different dimensional data on anomaly detection results. The normalization process also uses the Z-score normalization method; S2.2. Determine the optimal number of clusters using the elbow rule , traverse different Values, and calculate the corresponding intra-cluster squared error: S2.3, Random Selection Sample points as initial cluster centers , calculate each wind power station sample The Euclidean distance to all cluster centers is calculated and assigned to the nearest cluster. The distance calculation formula is: (12) Where, Indicates the A sample of wind power stations, Indicates the The center of the cluster, represents the Euclidean distance between the sample and the cluster center, Indicates the The set of samples contained in a cluster; S2.4. Recalculate the cluster center of the new cluster. The center update formula is as follows: (13) Where, For the The cluster centers after rounds of iterations, is the number of samples in the cluster; S2.5. Calculate the distances from all samples to the new cluster center and redistribute clusters based on the minimum distance. S2.6. Repeat S2.4 and S2.5 and calculate the change in the old and new cluster centers. Stop the iteration if the convergence condition is met or when the sample attribution no longer changes. The convergence condition is as follows: (14) Where, Indicates the The cluster in The center in the iteration, Indicates the The cluster in The center in the iteration, represents the Euclidean distance between the new and old cluster centers, is the set convergence threshold.

[0066] S2.7, finally get Wind power station clusters are constructed, and each cluster serves as a task unit of the multi-task learning model to provide grouped data for subsequent wind power prediction.

[0067] See Figure 5 As shown, Figure 5 This is a schematic diagram of the multi-task model framework. The multi-task learning model proposed in the present invention is used for medium- and long-term power prediction of wind power stations, mainly including a normalization layer, a dimensionality expansion layer, a shared information layer, a residual connection, an expert module and a rolling prediction mechanism.

[0068] First, the power and wind speed data of multiple wind power stations are input into the normalization layer and normalized using the Z-score method to ensure that data of different dimensions have the same scale: After normalization, the data is dimensionally expanded. Assuming the input data dimensions are (N, T, d), where N is the number of wind farms, T is the time step, and d is the feature dimension, the dimension expansion layer maps it to a high-dimensional space (N, T, d′) to enhance feature expression capabilities. After normalization, the clustered dataset is fed into the shared information layer, which consists of a Transformer encoder with a sparse attention mechanism. Its core computations include: (15) Where, is a normalization function, which is used to standardize the attention weight. , , are query, key, and value matrices respectively, is the transpose of the key matrix, For feature dimensions, a sparse attention mechanism is used to make the model focus on key wind power stations and improve computational efficiency; After the shared information is processed, its output is residually connected with the initial information of each cluster: , and sent to the corresponding expert module. Each expert module consists of multiple information extraction and prediction layers. The structure of each layer is the same as the Transformer structure of the shared information layer, and cooperates with the linear layer to complete the stage power prediction; Within the expert module, each prediction layer is responsible for power and wind speed prediction at different stages. The input of the prediction layer is the concatenation of the input and output of the previous layer. The specific formula is: (16) Where, is the current input, The prediction results of the current stage are concatenated and input into the next prediction layer; The total loss function is the weighted sum of all main task loss functions. The loss function of each main task is the weighted sum of its subtask loss functions. The model updates its parameters by minimizing the total loss function. Through the rolling prediction method, the prediction results of each stage are spliced together, and finally the medium- and long-term power prediction results of multiple wind power of the multi-task learning model are output.

[0069] In summary, the multi-stage medium- and long-term power prediction method for wind power clusters based on multi-task learning provided by the present invention proposes an innovative solution to the problem of insufficient accuracy and efficiency in the field of wind power prediction, especially in the medium- and long-term power prediction of wind power clusters: the short-term or ultra-short-term prediction of traditional single wind farms is difficult to meet the dispatching needs of modern power grids, the existing technology ignores the synergy and spatial and temporal correlation between wind power stations, and the traditional rolling prediction strategy lacks flexibility, which affects the prediction performance. The present invention applies the multi-task learning model to the medium- and long-term power prediction of wind power clusters for the first time, and realizes the simultaneous prediction of the power of multiple wind power stations through the collaborative work of the shared information layer and the expert module, which significantly improves the prediction efficiency and accuracy; adopts the Transformer structure based on the sparse attention mechanism to capture the long-distance dependency between wind power stations and improve the feature extraction capability; proposes a rolling prediction strategy and sets an independent loss function for each stage, accurately measures and optimizes the prediction error, and enhances the prediction stability and reliability; introduces the dynamic hyperbolic tangent activation function DyT in the normalization layer of the Transformer structure, and enhances the model expression and generalization capability through learnable parameters; uses the isolation forest algorithm to identify and eliminate outliers and combines A clustering algorithm rationally divides wind farms, optimizes data structure, and improves the accuracy and adaptability of the prediction model. A shared information layer enables cross-farm information sharing, addressing the problem of insufficient prediction for a single wind farm. To address the problem of decreased long-term prediction accuracy due to a reduction in historical samples in traditional rolling prediction, a spliced rolling prediction strategy is proposed to maintain the integrity of historical data and avoid error accumulation. Independent loss functions are designed for each rolling stage, and weights are adjusted to make the model better adaptable to the prediction needs of different stages. A sparse attention mechanism and DyT are creatively introduced into the Transformer architecture to improve computational efficiency, reduce model complexity, and enhance the model's adaptability to complex wind power environments, providing strong support for the optimized scheduling and operation of wind farm clusters. The present invention has important application value in the field of regional renewable energy power forecasting. Compared with traditional methods, it can identify outliers and fill missing values in data from multiple wind farms. By optimizing the data structure based on a clustering method and improving prediction accuracy and stability through a shared information layer and cluster-level expert modules, the proposed method addresses the shortcomings of existing methods in regional prediction timescales. The proposed method is applicable to wind farm clusters of different sizes, improving the scalability and adaptability of regional renewable energy power forecasting. Furthermore, the proposed method can be extended to photovoltaic power forecasting, enabling a wider range of renewable energy power forecasting applications.

Claims

1. A multi-stage wind power cluster medium- and long-term power forecasting method based on multi-task learning, characterized by: The following steps are involved: Step 1: Identify outliers and fill missing values in the power station’s historical data; Step 2: Based on the power, wind speed and other characteristics, use The clustering algorithm divides the power plants into clusters; Step 3: Input the preprocessed and clustered data into the proposed model to complete the medium- and long-term power forecast of each power station.

2. The multi-stage wind power cluster medium- and long-term power prediction method based on multi-task learning according to claim 1 is characterized in that: The specific steps in Step 1 include: Step 1.1: Standardize the historical wind speed and power data of multiple wind power stations so that data of different dimensions have the same scale; Step 1.2: Build an isolation forest model, determine outliers by calculating the average path length of samples in the isolation forest, and set a threshold based on the anomaly score to identify and remove abnormal data; Step 1.3: For the missing data identified, the adjacent value interpolation method is used to fill in the missing data to ensure the integrity and consistency of the data.

3. The multi-stage wind power cluster medium- and long-term power prediction method based on multi-task learning according to claim 2 is characterized in that: In Step 1.2, the isolation forest algorithm is used to determine the outliers, as follows: (1); (2); (3); Where, For the The harmonic number of positive integers is used to calculate the expected path length; is the normalization factor, Indicates the total number of samples in the dataset; Represents a sample average path length in isolated forests; Representative samples expected path lengths in multiple isolated trees; Representative samples Anomaly score.

4. The multi-stage wind power cluster medium- and long-term power prediction method based on multi-task learning according to claim 1 is characterized in that: The method of using Step 2 is as follows The clustering algorithm divides power plants into clusters. Specifically, wind power plants with similar characteristics are divided into multiple clusters. The specific steps are as follows: Step 2.1: Standardize the historical wind speed and power data of multiple wind power stations; Step 2.2: Traverse multiple Value, for each Value Application The algorithm calculates the corresponding intra-cluster squared error and plots curve, select the elbow inflection point as the optimal number of clusters; Step 2.3: Initialization Clustering, select the initial cluster center, calculate the Euclidean distance of each wind power station to all cluster centers, and assign it to the nearest cluster; Step 2.4: Calculate and update the center of each cluster and reallocate wind power stations to the new cluster center until convergence.

5. The multi-stage wind power cluster medium- and long-term power prediction method based on multi-task learning according to claim 4 is characterized in that: described The clustering algorithm calculates the intra-cluster squared error as follows: (4); (5); (6); Where, represents the number of clusters, Indicates the The feature vector of the sample points, Indicates belonging to a cluster The sample set, is the cluster center, Represents a sample With cluster center The square of the Euclidean distance, represents the time step, Indicates that in the wind speed sequence, The cluster in The cluster center coordinates on the feature dimension, Indicates that in the power sequence, the The cluster in The cluster center coordinates on the feature dimension.

6. The multi-stage wind power cluster medium- and long-term power forecasting method based on multi-task learning according to claim 1 is characterized in that: The specific steps of Step 3 include: Step 3.1: Build a multi-task learning model, including a shared information layer and an expert module. The shared information layer is used to extract global features between wind power stations. The output is residually connected with the initial information of each cluster and then input into the corresponding expert module. Step 3.2: The expert module consists of multiple prediction layers, each of which is responsible for power and wind speed prediction at different stages and has an independent loss function. Step 3.3: Combine the prediction results of each stage through rolling prediction and comprehensively optimize the model weights.

7. The multi-stage wind power cluster medium- and long-term power prediction method based on multi-task learning according to claim 6 is characterized in that: The shared information layer in Step 3.1 adopts a Transformer structure based on a sparse attention mechanism to model the collaborative relationship between wind power stations and extract global features between wind power stations. The normalization layer of the Transformer structure is replaced by a dynamic hyperbolic tangent activation function layer to enhance the expressive power of the model.

8. The multi-stage wind power cluster medium- and long-term power forecasting method based on multi-task learning according to claim 6 is characterized by: The input of the expert module in Step 3.2 is composed of the input and output of the previous layer to make full use of historical data. Each prediction layer is responsible for power and wind speed prediction at different stages, and has an independent loss function to measure the prediction error.

9. The multi-stage wind power cluster medium- and long-term power forecasting method based on multi-task learning according to claim 6 is characterized by: In the step 3.3 of splicing the prediction results of each stage by rolling prediction, an independent loss function is set for each rolling stage, and the loss weight of each stage can be adjusted according to the actual situation. The model parameters are adjusted according to the gradient of the total loss function through the back propagation algorithm to optimize the prediction performance.

10. The multi-stage wind power cluster medium- and long-term power forecasting method based on multi-task learning according to claim 9 is characterized by: The total loss function is the weighted sum of all main task loss functions, and the loss function of each main task is the weighted sum of its subtask loss functions. The model parameters are updated by minimizing the total loss function.