Enhanced power generation prediction

A hybrid machine learning approach using TiDE and GraphCast models enhances wind speed forecasting, addressing the variability of renewable energy sources and improving grid control by optimizing thermal generator dispatch and reducing emissions.

WO2026050153A1PCT designated stage Publication Date: 2026-03-05X DEVELOPMENT LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/043338
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-08-29
Filing Date
2025-08-25
Publication Date
2026-03-05

AI Technical Summary

Technical Problem

The integration of renewable energy sources into the electric grid is challenging due to the operational difficulties of managing their highly variable power generation, which is not directly controllable by operators, necessitating advanced forecasting methods to ensure stable generation and demand satisfaction.

Method used

A hybrid methodology using machine learning techniques, including a Time-series Dense Encoder (TiDE) model for short-term wind speed forecasting and a GraphCast graph neural network-based model for medium-term weather forecasting, combined with a fixed effects ordinary least squares regression model to predict thermal generator output, enhances the accuracy and reliability of wind speed forecasts.

Benefits of technology

The hybrid model improves the coordination and control of renewable energy integration by providing precise short- and medium-term wind speed predictions, optimizing thermal generator dispatch and reducing overall system-level emissions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025043338_05032026_PF_FP_ABST
    Figure US2025043338_05032026_PF_FP_ABST
Patent Text Reader

Abstract

This disclosure describes forecasting required power output for a plurality of thermal generators. Including training a Time-series Dense Encoder (TiDE) model on a corpus of historical weather data for a particular geographic region, the trained TiDE model forecasts weather for the geographic region. Training a machine learning model on historical power output and weather data for a particular wind farm within the geographic region, wherein the trained machine learning model predicts the power output out of the wind farm for a given weather condition. Forecasting, using the TiDE model, the weather in the geographic region. Predicting the power output of the wind farm by providing the forecasted weather to the machine learning model. Forecasting a required power output for a plurality of thermal generators based on the predicted power output of the wind farm by calculating the output of a fixed effects ordinary least squares (OLS) regression model.
Need to check novelty before this filing date? Find Prior Art

Description

Attorney Docket No.: 43374-0851WO1 ENHANCED POWER GENERATION PREDICTION

[0001] This disclosure generally relates to predicting the power output of wind turbines and applying that prediction to electrical grid control. BACKGROUND

[0002] Modern power grids use a combination of renewable sources and thermal sources. Renewable sources often produce an output that is dependent on environmental factors that are not directly controllable by operators. Such factors can include, but are not limited to, windspeed, air density, solar intensity, cloud cover, or precipitation. Thermal sources often have an output that is controllable, thus grid operators adjust the output of various thermal sources such that the combination of thermal and renewable sources can satisfy demand at any given time period. SUMMARY

[0003] In general, this disclosure involves a system, medium, and method for forecasting required power output for a plurality of thermal generators. This includes training a Time-series Dense Encoder (TiDE) model on a corpus of historical weather data for a particular geographic region, the trained TiDE model forecasts weather for the particular geographic region. Training a machine learning model on a corpus of historical power output and weather data for a particular wind farm within the particular geographic region, wherein the trained machine learning model predicts the power output out of the wind farm for a given weather condition. Forecasting, using the TiDE model, the weather in the particular geographic region. Predicting the power output of the wind farm by providing the forecasted weather to the machine learning model. Forecasting a required power output for a plurality of thermal generators based on the predicted power output of the wind farm by calculating the output of a fixed effects ordinary least squares (OLS) regression model.

[0004] Systems and methods optionally include one or more of the following features.

[0005] In some instances, the TiDE model includes a multi-layer perceptron (MLP) based encoder-decoder architecture.

[0006] In some instances, the TiDE model includes an encoder with four residual blockseach containing fully connected layers of 128 neurons, and a decoder with four residual blocks each containing fully connected layers of 128 neurons.

[0007] In some instances, the TiDE model is trained using a mean square error loss function and an Adaptive Moment Estimation (Adam) optimization.Attorney Docket No.: 43374-0851WO1

[0008] In some instances, the fixed effects OLS regression model is a non-logarithmic statistical model. In some instances, calculating output of the non-logarithmic statistical modelincludes identifying a value for Gt in the equation ^^௧ ൌ ^^ ^ ^^^^^௧ ^ ^^ଶ^^௧ ^ ^^ଷ^^௧ ^^^ସ^^^௧ െ ^^௧ି^^ ^ ^^ହ^^^௧ െ ^^௧ି^^ ^ ^^௧,^ ^ ^^^^௧ ^ ^^௧,^ ^ ^^.

[0009] In some instances, the weather datawind direction,

[0010] The details of one or more implementations of the subject matter of this specification are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages of the subject matter will become apparent from the description, the drawings, and the claims. DESCRIPTION OF DRAWINGS

[0011] FIG.1A shows short-term mean hourly RMSE and normalized RMSE of wind magnitude predictions

[0012] FIG. 1B shows medium-term mean hourly RMSE and normalized RMSE of wind magnitude predictions

[0013] FIG.2 shows mean generation of various power sources per hour by year.

[0014] FIG.3 shows normalized power generation by hour of day for solar and wind generation.

[0015] FIG. 4 shows normalized root mean square error difference in forecasts for 12 hours, 2 days, and 10 days respectively.

[0016] FIG.5 is a flowchart illustrating an example process for forecasting weather.

[0017] FIG.6 illustrates a schematic diagram of an example computing system. DETAILED DESCRIPTION

[0018] Replacing fossil fuels with renewable energy for electricity generation has a direct impact on climate change, since electricity generation globally contributes to roughly a third of annual carbon emissions. However, the integration of renewable energy sources into the electric grid is challenging due to the operational difficulties of managing their power generation, which is highly variable compared to fossil fuel sources, delaying the availability of clean energy. Chile has taken a globally leading role in clean energy, based on its excellent renewable energy resources, emergence as a leading destination for solar and wind developers, and ambitious clean energy goals for the future. This makes it a particularly interesting andAttorney Docket No.: 43374-0851WO1 important area of study for renewable energy forecasting. While its electricity demand is relatively modest (6.5% of Central and South America’s total in 2022), it is growing rapidly, both in aggregate (up from 40 TWh in 2000 to 88 TWh in 2022) and per-capita (70% over the same period).

[0019] In response to the need for enhanced coordination and control, electric grid Independent System Operators (ISOs) recognize the need for precise renewable generation and net demand forecasts. Historically, thermal plants provided stable generation with predictable net demand patterns. However, increased penetration of renewables necessitates advanced forecasting due to significant generation ramps from wind and solar. Accurate renewable generation forecasts are critical for subsequent thermal generator dispatch, which rely on forecast accuracy, security constraints and contingency requirements. Among renewable sources, wind power prediction is particularly challenging due to its temporal and spatial variability, and oftentimes lack of sufficient high-quality data, not only for wind power but for wind speed itself.

[0020] This disclosure contributes to improving the coordination and control of Chile’s electric power system amidst growing renewable energy integration by enhancing the reliability and accuracy of wind speed forecasts. The disclosed approach diverges from traditional forecast methods by applying machine learning techniques. It introduces a novel hybrid methodology for short- and medium-term wind speed prediction, designed as an input to the ISO’s unit commitment and economic dispatch models in the day-ahead and week-ahead grid operations and market. For short-term forecasts ML models specifically optimized for capturing the intricate temporal dependencies in hourly wind speed data are developed and described. For medium-term forecasts, an ML model based on GraphCast—a graph neural network-based machine learning weather prediction (MLWP) model for medium-range weather forecasts is described.

[0021] Data and Methods:

[0022] A dataset is developed with thermal power plants that includes hourly generation by fuel source obtained from Coordinador Eléctrico Nacional (CEN), Chile’s national ISO. CEN also records historical hourly generation from wind, solar, hydro and geothermal plants. In addition to thermal and renewable energy generation, hourly country- level demand is recorded. For short and medium-term wind magnitude forecasts, ERA5, a global reanalysis dataset produced by the European Centre for Medium-Range Weather Forecasts (ECMWF) can be used. GraphCast is initially trained on ERA5 data spanning 1979- 2016. The ECMWF’s high-resolution operational forecasting system (HRES) provides real-Attorney Docket No.: 43374-0851WO1 time weather predictions using the most recent observations and advanced numerical weather prediction models. HRES-fc0 is a dataset derived from HRES, consisting of the initial state of each high-resolution ensemble forecast generated by the ECMWF. This dataset is used as the benchmark for evaluating our hybrid model.

[0023] Statistical models. a fixed effects ordinary least squares (OLS) regression formulation is used to model the change in operational behavior of thermal generators in response to a marginal increase in generation from wind and solar. The formulation is a non- logarithmic form of the statistical model:

[0024] ^^௧ ൌ ^^ ^ ^^^^^௧ ^ ^^ଶ^^௧ ^ ^^ଷ^^௧ ^ ^^ସ^^^௧ െ ^^௧ି^^ ^ ^^ହ^^^௧ െ ^^௧ି^^ ^ ^^௧,^ ^^^^^௧ ^ ^ ^^and wind generation, respectively; Dt is the electricity demand, and the terms between the brackets denote the hourly wind and solar ramp. ηmand ηadenote the month and year fixed effects. Xtis a set of control variables that includes the hydro and geothermal generation, and imports.The parameters that are of interest are the coefficients for solar, β1 and wind, β2. These coefficients represent the change in thermal generation in response to a marginal unit increase in generation.

[0026] Short-term forecasts. The Time-series Dense Encoder (TiDE) model represents a state-of-the-art approach for short-term wind speed forecasting. This model employs a Multi- layer Perceptron (MLP) based encoder-decoder architecture, optimized to handle time-series data with multiple covariates and complex, non-linear dependencies. The encoder processes historical wind speed data and relevant covariates, transforming them into a latent space representation through multiple residual blocks, which are crucial for capturing intricate patterns in the data. The decoder then reconstructs future wind speeds from this latent representation, allowing the model to make accurate predictions up to 48 hours ahead.

[0027] Medium-term forecasts using GraphCast. GraphCast is a neural network architecture designed to predict weather states by taking the two most recent states of Earth’s weather - the current time and six hours earlier - and forecasting the next state six hours ahead. The model uses a 0.25-degree latitude-longitude grid to represent a single weather state, corresponding to roughly 28 km by 28 km resolution at the equator. Implemented as a neural network architecture based on Graph Neural Networks (GNNs) in an "encode-process-decode" configuration, GraphCast consists of 36.7 million parameters. To improve the forecast accuracy of GraphCast, a novel multi-stage finetuning process is employed. This involvedAttorney Docket No.: 43374-0851WO1 autoregressive fine-tuning on HRES-fc0 data to account for potential distribution shifts between reanalysis data (ERA5) and real-time operational data. The model is further customized through location-based weighting, emphasizing accuracy within a bounding box encompassing the Chilean wind farms, and wind magnitude and power weighting, directly optimizing for these specific metrics. Finally, iterative fine-tuning is combined with linear regression post-processing to further enhance the model’s accuracy. The base GraphCast model is first fine-tuned autoregressively on five years (2016-2020) of HRES-fc0 data. During this stage, we used the original GraphCast loss function, a spatially-weighted mean squared error (MSE) calculated across all variables and pressure levels:

[0028] ^^^^^^ൌ ^^^^^ ^^ ^^ ^^^^ ^^ ^ଶ training0.25◦ cells in 0.25-degree latitude-longitude grid, J is the set of variables and pressure levels, sjrepresents the inverse variance of time differences for variable j, wj is the loss weight for variable j, ai is the area of grid cell i, normalized to unit mean over the grid, ^^^^,^,ௗబାఛis the model’s prediction for variable j at grid cell i and lead time τ from initialization d0, and xi,j,d0+τ is the corresponding target value from the HRES-fc0 data.

[0030] The loss function is also modified by incorporating a location-based weighting factor mi, which up-weights the error contributions from grid cells within a specified bounding box by a factor ωl, while grid cells outside the bounding box are left unweighted.

[0031] Results

[0032] Marginal response of thermal power plants. As indicated by analyzing the coefficients of the statistical model, on average, fluctuations or ramps in thermal generation are more sensitive to increases in wind generation compared to solar. Looking at the coefficients β1 and β2 of the model, a 1 GWh increase in wind generation displaces 0.95 GWh of thermal generation, whereas the same increase in solar generation displaces only 0.67 GWh of thermal generation, which is 30% less than the displacement from wind. The response of individual thermal generators can also be modeled to marginal changes in wind and solar generation by isolating hourly generation in the fixed-effects model. Using the coefficients for solar and wind generation, each power plant can be classified as either ‘solar-following’ or ‘wind-following’ determined based on the absolute value of the larger, statistically significant coefficient. Of the 128 thermal plants, 91 are wind-following, while the remaining 37 largely respond to utility-Attorney Docket No.: 43374-0851WO1 scale solar generation. By optimizing day- and week-ahead dispatch, we can ensure wind- following plants are consistently operated at higher capacity factors with a lower baseline emissions intensity resulting in overall lower system-level emissions.

[0033] Operational wind speed forecasts. In medium and long-term weather forecasts, the current top deterministic operational system in the world is the European Center for Medium-Range Weather Forecasts (ECMWF)’s High RESolution forecast (HRES). GraphCast is a graph neural network forecasting model that outperforms the baseline HRES forecasts on 90.3% of 1380 location targets with an overall skill improvement of 7 to 14% across global locations. However, the vanilla GraphCast model for Chile proves to be more accurate than HRES forecasts only following a lead time of 100 hours (approximately 4 days). Before this cross-over point, GraphCast’s normalized Root Means Square Error (RMSE) relative to HRES ranges from 0 to 1.15 shown in the SI. To overcome the limitations of GraphCast at short lead times that we observe in Chile, we categorize the forecasting problem into two distinct temporal regimes using a hybrid model. For the purposes of making a direct comparison of our methods with HRES, we use the same quarter-degree mesh that is used in HRES. However, we focus on those mesh points that are closest to the actual 52 wind farm locations in Chile. Thus wind speed forecasts are developed for 24 unique quarter-degree interpolated locations in Chile.

[0034] Short-term forecasts using transformers. We evaluate the performance of the TiDE model in predicting short-term wind speeds using the RMSE metric, averaged hourly across all 24 locations. FIG. 1A shows the mean hourly averaged RMSE values for wind magnitude over a 48-hour prediction horizon for both TiDE 102 and HRES 106. The RMSE for wind magnitude predictions (left panel of the figure) increases gradually over the 48-hour forecast period for both models; for HRES the data available is at 6-hour intervals. Unlike other auto-regressive models, the TiDE model predicts the full 48-hour window as a single output, and the optimization function optimizes for the total loss over the full 48-hour window. For a short window, this direct optimization over the full window performs better than the auto- regressive models, and this can be proved with our 4-21% improvement over HRES as shown in the right hand panel 104 in FIG.1A.

[0035] Medium-term forecasts using GraphCast. For medium-term forecasts, we find that weighting training inputs by location, wind magnitude independently, and a combination of both resulted in iterative improvements in skill and RMSE. FIG.1B summarizes the RMSE for three fine-tuned GraphCast variants at quarter-degree interpolated wind farm locations. At 96-hour lead times, these weighted models outperform HRES predictions by 10%, 13%, and 17% respectively. The location and wind-weighted GraphCast has greater wind velocityAttorney Docket No.: 43374-0851WO1 forecasting skill than HRES for periods after a lead time of 30 hours at various pressure levels which roughly correspond to those of wind farm elevations. When further improved by incorporating HRES inputs as a covariate, the crossover point for wind magnitude predictions improves to 2.5 days.

[0036] FIGS 1A and 1B illustrate Mean hourly RMSE and Normalized RMSE of wind magnitude predictions averaged across all interpolated wind farm locations for 2021. (a) Short- term forecasts (0-2 days). The left column represents mean hourly wind magnitude RMSE values over a 48-hour forecast horizon, while the right column represents normalized RMSE considering HRES as the baseline using TiDE forecasts at six-hour intervals. (b) Medium-term forecasts (2-10 days). The left column represents wind magnitude RMSE values over an 8 day forecast horizon commencing at 2 days and at sixhour time intervals. The solid lines represent three alternate GraphCast models considered in the analysis. The right column represents normalized RMSE considering HRES as the baseline at six-hour intervals.

[0037] Methods

[0038] We use a fixed effects ordinary least squares (OLS) regression formulation to model the change in operational behavior of thermal generators in response to a marginal increase in generation from wind and solar. The formulation is a non-logarithmic form of a statistical model.

[0039] ^^௧ ൌ ^^ ^ ^^^^^௧ ^ ^^ଶ^^௧ ^ ^^ଷ^^௧ ^ ^^ସ^^^௧ െ ^^௧ି^^ ^ ^^ହ^^^௧ െ ^^௧ି^^ ^ ^^௧,^ ^^^^^௧ ^ ^^௧,^ ^ ^^

[0040] where t denotes the current timestep; Gt, St, Wt represent the thermal, solar, and wind generation, respectively; Dtis the electricity demand, and the terms between the brackets denote the hourly wind and solar ramp. ηm and ηa denote the month and year fixed effects. Xt is a set of control variables thathydro and geothermal generation, and imports.

[0041] In our formulation, the parameters that are of interest are the coefficients for solar, β1and wind, β2. These coefficients represent the change in thermal generation in response to a marginal unit increase in generation from solar and wind, respectively. For example, if generation from wind increases by 1 MWh, then β1and β2will represent the MWh change in generation from thermal plants in Chile. These coefficients thus measure the causal change in the dependent variable under the assumption that daily wind and solar production (represented by W and S) is uncorrelated with the error term ^ after controlling for demand and time fixed effects. This assumption is valid for wind at the hourly level and can thus be extended into the present form of the statistical model.Attorney Docket No.: 43374-0851WO1

[0042] Short-term wind forecasting

[0043] The Time-series Dense Encoder (TiDE) model represents a state-of-the-art approach for short-term wind speed forecasting. This model employs a Multi-layer Perceptron (MLP) based encoder-decoder architecture, optimized to handle time-series data with multiple covariates and complex, non-linear dependencies. The encoder processes historical wind speed data and relevant covariates, transforming them into a latent space representation through multiple residual blocks, which are crucial for capturing intricate patterns in the data. The decoder then reconstructs future wind speeds from this latent representation, allowing the model to make accurate predictions up to 48 hours ahead.

[0044] The hyperparameters of the TiDE model are chosen to enhance its predictive performance. Key hyperparameters include the number of layers in the encoder and decoder, the hidden layer sizes, and dropout rates to prevent overfitting. Specifically, the encoder and decoder each consist of four residual blocks, with each block containing fully connected layers of 128 neurons. Dropout is applied with a rate of 0.1 to ensure regularization. The model is trained using the Mean Squared Error (MSE) loss function, which measures the average squared difference between predicted and actual wind speeds. The Adam optimizer, with a learning rate of 0.001, is utilized to update the model’s weights efficiently during training. An early stopping mechanism with a patience of 200 epochs is implemented to halt training when the validation loss ceases to improve, thus avoiding overfitting.

[0045] The TiDE model’s performance is evaluated using the Root Mean Squared Error (RMSE) metric, averaged hourly across multiple locations in Chile. To handle the time-series nature of the data, the model incorporates various datetime attributes as additional covariates, such as the hour of the day, day of the week, and month, which are encoded using which are encoded using dense layers in a Multi-Layer Perceptron (MLP)-based architecture. Additionally, the model is equipped to use a Scaler transformer to normalize the input data, ensuring that all features have zero-mean unit variance. This preprocessing step is essential for maintaining the stability and performance of the neural network. The inclusion of future covariates, such as temperature, pressure, and various wind parameters at different altitudes, enhances the model’s ability to capture the underlying weather patterns, leading to more accurate short-term wind speed forecasts.

[0046] Let y1:L represent the historical wind speeds, x1:L+H the covariates over the look-back and forecast horizon, and a the static attributes. The encoding step maps the past covariates and wind speeds into a dense representation e:

[0047] ^^௧ ൌ ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^௧^Attorney Docket No.: 43374-0851WO1

[0048] ^^ ൌ ^^^^^^^^^^^^^^^^^^:^; ^^^:^;^^^

[0049] This encoding e is then utilized by the decoder to predict future wind speeds.

[0050] The decoding process in the TiDE model involves mapping the encoded hidden representations into future predictions. The dense decoder processes the encoding e to produce an intermediate representation, which is further refined by the temporal decoder to generate the final forecast. The process is described by the following equations:

[0051] ^^ ൌ ^^^^^^^^^^^^^^^^^^

[0052] ^^ ൌ ^^^^^^ℎ^^^^^^^^^^

[0053] ^^^^ା௧ ൌ ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^௧; ^^^ା௧^

[0054] where g is the output of the dense decoder, D is the reshaped matrix of decoded vectors, and dtis the decoded vector for time t. The temporal decoder combines this decoded vector with the projected covariates xL+t to generate the final prediction ^^^^ା௧.

[0055] The proposed method involves three major steps: data decomposition, individual prediction using randomized algorithms, and results ensemble.

[0056] Data Decomposition: We employ an effective multi-scale analysis technique such as Ensemble Empirical Mode Decomposition (EEMD) to decompose the original time series data into n intrinsic mode functions (IMFs) and a residue. This decomposition ensures that the data complexity is reduced, making it easier to model.

[0057] ^^௧ ൌ ∑^ ^ୀ^ ^^^,௧ ^ ^^௧using Randomized Algorithms: For each decomposedcomponent, we apply a randomized algorithm like ELM, RVFL, or RKS to predict the future values. These algorithms operate on the principle of randomization to achieve faster computation and improved generalization.

[0059] ^̂^^,௧ ൌ ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ℎ^^^^^^,௧^

[0060] Results Ensemble: The final forecast is obtained by aggregating the predictions of all decomposed components and the residue.

[0061] ^^^௧ ൌ ∑^ ^ୀ^ ^̂^^,௧ ^ ^̂^௧forecasting function f can be defined as:

[0063] ^^: ^^^^^^ே :^^, ^^^^^ே :^ାு^, ^^^^^ே ^→ ^∑^^^ ^ୀ^ ^̂^^ ^ ^̂^ே ା^:^ାு,^ ^ା^:^ାு,^H is the forecast horizon. The accuracy of the prediction is measured by the Root Mean Squared Error (RMSE), defined as:Attorney Docket No.: 43374-0851WO1 ே ே

[0065] RMSE: ^^^^^^ା^:^ାு^^ୀ^, ^∑^ ^ୀ^ ^̂^ ^^ା^:^ାு,^ ^ ^̂^^ ^ା^:^ାு,^^^ୀ^ ^ ൌweather states by taking the two most recent states of Earth’s weather - the current time and six hours earlier - and forecasting the next state six hours ahead. The model uses a 0.25-degree latitude- longitude grid to represent a single weather state, corresponding to roughly 28 km by 28 km resolution at the equator. Implemented as a neural network architecture based on Graph Neural Networks (GNNs) in an "encode-process-decode" configuration, GraphCast consists of 36.7 million parameters.

[0069] The encoder within GraphCast employs a single GNN to map variables, normalized to zero-mean unit variance, represented as node attributes on the input grid to learned node attributes on an internal multi-mesh representation. This multi-mesh is spatially homogenous and defined by refining a regular icosahedron iteratively six times, leading to a high-resolution graph with 40,962 nodes and a flat hierarchy of edges of varying lengths. The processor then uses 16 unshared GNN layers to perform learned message-passing on the multi- mesh, enabling efficient local and long-range information propagation with minimal message- passing steps. The decoder maps the final processor layer’s learned features back to the latitude-longitude grid using a single GNN layer, predicting the output as a residual update to the most recent input state, with output normalization to achieve unit-variance on the target residual.

[0070] For fine-tuning GraphCast to optimize wind magnitude and power forecasts specifically for Chilean wind farms, a multi-stage fine-tuning process was employed. This involved autoregressive fine-tuning on HRES-fc0 data to account for potential distribution shifts between reanalysis data (ERA5) and real-time operational data. The model was further customized through location-based weighting, emphasizing accuracy within a bounding box encompassing the Chilean wind farms, and wind magnitude and power weighting, directly optimizing for these specific metrics. Iterative fine-tuning combined with linear regression post-processing enhanced the model’s accuracy, making it a robust tool for predicting wind patterns in Chile’s diverse and complex terrain.

[0071] Finetuning GraphCastAttorney Docket No.: 43374-0851WO1

[0072] GraphCast demonstrates strong global weather forecasting skills, but our objective is to optimize wind magnitude and power predictions specifically for Chilean wind farms. To achieve this, we customized the publicly available GraphCast model, pre-trained on ERA5 reanalysis data, using a multi-stage fine-tuning process. This section elaborates on each stage and provides the mathematical formulation of the modified loss function.

[0073] Operational Data Fine-tuning

[0074] Recognizing the potential distribution shift between reanalysis data (ERA5) and real-time operational data (HRES-fc0), we first fine-tuned the base GraphCast model autoregressively on five years (2016- 2020) of HRES-fc0 data. This fine-tuning involved several steps. Firstly, HRES-fc0 data was pre-processed to align with the GraphCast input format. This included variable selection, focusing on relevant variables directly related to wind dynamics, such as eastward and northward wind components at various pressure levels. Additionally, the data was converted to the same spatial resolution as GraphCast (0.25-degree latitude-longitude grid).

[0075] The autoregressive fine-tuning process involved 30,000 steps where the model predicted the weather state six hours ahead for up to a ten-day lead time. In each training step, the model’s output from the previous timestep was fed back as input, simulating a real-time forecasting scenario. This recursive process allowed the model to learn temporal dependencies and adapt to the dynamics of operational data. During this stage, we used the original GraphCast loss function, a spatially-weighted mean squared error (MSE) calculated across all variables and pressure levels:

[0076] ^^^^^^ൌ ^^^^ଶtimes in training set, Ttrain is the number of autoregressive steps used for training, G0.25◦ is the set of grid cells in the 0.25-degree latitude-longitude grid, J is the set of variables and pressure levels, sjrepresents the inverse variance of time differences for variable j, wj is the loss weight for variable j, aiis the area of grid cell i, normalized to unit mean over the grid, ^^^^,^,ௗబାఛis the model’s prediction for variable j at grid cell i and lead time τ fromis the corresponding target value from the HRES-fc0 data.

[0078] Location and Wind-Magnitude Weighting

[0079] To focus the model’s predictive skill on Chilean wind farms, we applied location-based weighting and wind magnitude weighting. The location-based weightingAttorney Docket No.: 43374-0851WO1 emphasized the accuracy within a bounding box encompassing the Chilean wind farms, while the wind magnitude weighting directly optimized for predicting wind weather variables. This involved adjusting the weights in the loss function to prioritize errors in these regions and variables, thereby fine-tuning the model to the specific operational characteristics of the Chilean wind farms.

[0080] The location-based weighting modified the loss function as follows:

[0081] ^^ ൌ^^^^^^ ^^^^^^ ^^ ^^ ^^^^ ^^ ^ଶ ^^ ^loss terms to optimize the model for predicting wind magnitude (wm) and power (wp). Wind magnitude was calculated from the predicted eastward (u ) and northward (v ) wind components at 1010m 10m meters:

[0086] ^^^^ ൌ ^^^ ଶ^^^ ^ ^^^ଶ^^using a simplified cubic relationship:

[0088] ^^^^ ൌ

[0089] The additional loss terms were incorporated as follows:

[0090] ^^^ଶ ଶ௪^^ௗ ൌ ^^^^^ ^ ∑^ఢீ ^^^^^^௪ ^^^ െ ^^ ^ ^ ^^ ^^^ െ ^^ ^ ൧^ ^,^ ^,^ ௪^ ^,^ ^,^బ.మఱ°ீబ.మఱ°themodel’s predicted u10m and v10m at grid cell i, and ωwm and ωwp are the weighting factors for wind magnitude and power.

[0092] Iterative Fine-tuning and Configuration Selection

[0093] An iterative fine-tuning process was employed to refine the model further. Various configurations were tested, adjusting hyperparameters and re-evaluating performance to identify the optimal setup. The configuration achieving the lowest RMSE for Chilean wind farm locations was chosen as the final fine-tuned model.

[0094] To enhance accuracy, we applied a per-lead-time bias correction using linear regression. This involved training separate linear regression models for each six hour lead time (up to 10 days), variable (e.g., u10m, v10m, wind magnitude, and power), and grid cell withinAttorney Docket No.: 43374-0851WO1 the bounding box. The last year of HRES-fc0 data was used for training these models. The correction formula is given by:

[0095] ^^^^^^^^^௧^ௗ ൌ ^^ ^ ^^^^^^^௪

[0096] where ^^^^^^^^^௧^ௗis the bias-corrected prediction, ^^^^^௪is the raw GraphCast prediction, and α and β are the learned intercept and slope, respectively. These models were applied to the fine-tuned GraphCast predictions to generate the final bias-corrected wind forecasts.

[0097] HRES Initialized Inputs

[0098] To further enhance the predictive accuracy of our location and wind-magnitude weighted GraphCast model, we investigated the integration of real-time forecasts from the ECMWF’s High-Resolution Forecasting System (HRES) as additional input features. Building on existing work in this field, our "HRES forcing" approach leverages the inherent short-term predictive skill of HRES to potentially inform and improve the medium-range wind speed forecasts generated by the GraphCast model.

[0099] Implementation of HRES forcing involves augmenting the input layer of our customized GraphCast model to accommodate the HRES forecast variables. The weights associated with these new inputs are initialized as zeros, preserving the model’s pre-existing functionality, which is derived from pre-training on ERA5 and operational fine-tuning on historical HRES-fc0 data. A subsequent singlestep autoregressive fine-tuning process, utilizing a reduced number of steps (2,000 compared to the standard 30,000), allows the model to learn how to effectively incorporate the real-time HRES information while minimizing disruption to the previously established, regionally-tailored weights.

[0100] FIG. 2 shows mean hourly generation and standard deviation for wind 202, solar 204, and thermal 206 power plants in Chile from 2019 to 2023. Each row corresponds to a different year, and each column represents a different technology (Wind, Solar, and Thermal). The solid lines indicate the mean generation in gigawatt-hours (GWh), and the shaded areas represent the standard deviation across all hours. The x-axis denotes the hour of the day, ranging from 0 to 23. This figure illustrates the temporal variation in power generation for each technology over the years, highlighting both the average generation and the variability within each day.

[0101] FIG. 3 shows hourly solar and wind generation (2019-2023). Solar (left) 302 and Wind (right) 304. The colored bars represent mean generation normalized by installed capacity by source and year. The bars indicate the standard deviation of hourly normalized generation by source.Attorney Docket No.: 43374-0851WO1

[0102] FIG. 4 shows Normalized RMSE difference of GraphCast’s 10u forecasts relative to HRES, by location, at 12 hours 402, 2 days 404, and 10 day 406 lead times. Blue indicates that GraphCast has greater skill than HRES, Red that HRES has greater skill. Here "10u" refers to the u-component of wind at an altiitude corresponding to 10m

[0103] FIG. 5 is a flowchart illustrating an example process 500 for forecasting weather. It will be understood that process 500 may be performed, for example, by any suitable system, environment, software, and hardware, or a combination of systems, environments, software, and hardware as appropriate. The operations shown in process 500 may not be exhaustive and other operations can be performed as well before, after, or in between any of the illustrated operations. Further, some of the operations may be performed simultaneously or in a different order than shown in FIG. 5. In some implementations, some of the operations may be performed by a computer or multiple computers. For example, one or more of a computation system 600 of FIG.6, appropriately programmed, can perform the process 500.

[0104] At 502, a TiDE model is trained on historical weather data. This can include historical wind speed data and relevant covariates. The TiDE model can be trained to transform the historical data into a latent space representation through multiple residual blocks. The decoder of the TiDE model then reconstructs future wind speeds from this latent representation, allowing the model to make accurate predictions up to 48 hours ahead. In some implementations, the TiDE model is a MLP based encoder-decoder architecture with four residual blocks in both the encoder and the decoder. Each set of blocks containing fully connected layers of 128 neutrons. In some implementations, training is conduced using Adaptive moment Estimation (Adam) optimization. In some implementations, other training techniques are used. For example the TiDE model can be trained using stochastic gradient descent (SGD), cross-entropy, RMSProp, or other techniques. In some implementations, the historical weather data includes wind speed data, wind direction data, precipitation data, and temperature data.

[0105] At 504, a machine learning model is trained on historical power outputs and weather data for a particular wind farm. In some implementations, the machine learning model is GraphCast—a graph neural network-based machine learning. Other machine learning models are further possible. The machine learning model can be trained to predict a power output for a particular wind farm based on provided weather data.

[0106] At 506, The TiDE model is used to forecast weather in a region. In some implementations the region includes the particular wind farm.Attorney Docket No.: 43374-0851WO1

[0107] At 508, the forecasted weather is provided to the machine learning model to predict the power output of the wind farm for a period of time during the forecast.

[0108] At 510, the predicted power output can then be used to forecast the required output for a plurality of thermal generators based on the predicted power output. Thermal generators can include, but are not limited to geothermal, coal, natural gas, oil, or other power generators. Forecasting the required power output can be accomplished using a fixed effects ordinary least squares (OLS) regression model. The OLS regression model can, in some implementations, be a non-logarithmic statistical model, where the output is Gtin the equation^^௧ ൌ ^^ ^ ^^^^^௧ ^ ^^ଶ^^௧ ^ ^^ଷ^^௧ ^ ^^ସ^^^௧ െ ^^௧ି^^ ^ ^^ହ^^^௧ െ ^^௧ି^^ ^ ^^௧,^ ^ ^^^^௧ ^ ^^௧,^ ^ ^^.600.the implementations described herein. For example, the system 600 may be included in computing devices of the one or more online components and / or the one or more offline components. The system 600 includes a processor 610, a memory 620, a storage device 630, and an input / output device 640, which are interconnected using a system bus 650. The processor 610 is capable of processing instructions for execution within the system 600. In some implementations, the processor 610 is a single-threaded processor. The processor 610 is a multi-threaded processor. The processor 610 is capable of processing instructions stored in the memory 620 or on the storage device 630 to display graphical information for a user interface on the input / output device 640.

[0110] The memory 620 stores information within the system 600. In some implementations, the memory 620 is a computer-readable medium. The memory 620 can be a volatile memory unit or a non-volatile memory unit. The storage device 630 is capable of providing mass storage for the system 600. The storage device 630 is a computer-readable medium. The storage device 630 may be a floppy disk device, a hard disk device, an optical disk device, or a tape device. The input / output device 640 provides input / output operations for the system 600. The input / output device 640 includes a keyboard and / or pointing device. The input / output device 640 includes a display unit for displaying graphical user interfaces.

[0111] Implementations of the subject matter and the functional operations described in this specification can be implemented in digital electronic circuitry, in tangibly embodied computer software or firmware, in computer hardware, including the structures disclosed in this specification and their structural equivalents, or in combinations of one or more of them. Implementations of the subject matter described in this specification can be implemented asAttorney Docket No.: 43374-0851WO1 one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible non transitory storage medium for execution by, or to control the operation of, data processing apparatus. The computer storage medium can be a machine- readable storage device, a machine-readable storage substrate, a random or serial access memory device, or a combination of one or more of them. Alternatively, or in addition, the program instructions can be encoded on an artificially generated propagated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal, that is generated to encode information for transmission to suitable receiver apparatus for execution by a data processing apparatus.

[0112] The term “data processing apparatus” refers to data processing hardware and encompasses all kinds of apparatus, devices, and machines for processing data, including by way of example a programmable processor, a computer, or multiple processors or computers. The apparatus can also be, or further include, special purpose logic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application specific integrated circuit). The apparatus can optionally include, in addition to hardware, code that creates an execution environment for computer programs, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them.

[0113] A computer program, which may also be referred to or described as a program, software, a software application, an app, a module, a software module, a script, or code, can be written in any form of programming language, including compiled or interpreted languages, or declarative or procedural languages; and it can be deployed in any form, including as a stand alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A program may, but need not, correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data, e.g., one or more scripts stored in a markup language document, in a single file dedicated to the program in question, or in multiple coordinated files, e.g., files that store one or more modules, sub programs, or portions of code. A computer program can be deployed to be executed on one computer or on multiple computers that are located at one site or distributed across multiple sites and interconnected by a data communication network.

[0114] The processes and logic flows described in this specification can be performed by one or more programmable computers executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows canAttorney Docket No.: 43374-0851WO1 also be performed by special purpose logic circuitry, e.g., an FPGA or an ASIC, or by a combination of special purpose logic circuitry and one or more programmed computers.

[0115] Computers suitable for the execution of a computer program can be based on general or special purpose microprocessors or both, or any other kind of central processing unit. Generally, a central processing unit will receive instructions and data from a read only memory or a random-access memory or both. The essential elements of a computer are a central processing unit for performing or executing instructions and one or more memory devices for storing instructions and data. The central processing unit and the memory can be supplemented by, or incorporated in, special purpose logic circuitry. Generally, a computer will also include, or be operatively coupled to receive data from or transfer data to, or both, one or more mass storage devices for storing data, e.g., magnetic, magneto optical disks, or optical disks. However, a computer need not have such devices. Moreover, a computer can be embedded in another device, e.g., a mobile telephone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a Global Positioning System (GPS) receiver, or a portable storage device, e.g., a universal serial bus (USB) flash drive, to name just a few.

[0116] Computer readable media suitable for storing computer program instructions and data include all forms of non volatile memory, media and memory devices, including by way of example semiconductor memory devices, e.g., EPROM, EEPROM, and flash memory devices; magnetic disks, e.g., internal hard disks or removable disks; magneto optical disks; and CD ROM and DVD-ROM disks.

[0117] To provide for interaction with a user, implementations of the subject matter described in this specification can be implemented on a computer having a display device, e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, for displaying information to the user and a keyboard and a pointing device, e.g., a mouse or a trackball, by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback, e.g., visual feedback, auditory feedback, or tactile feedback; and input from the user can be received in any form, including acoustic, speech, or tactile input. In addition, a computer can interact with a user by sending documents to and receiving documents from a device that is used by the user; for example, by sending web pages to a web browser on a user’s device in response to requests received from the web browser.

[0118] Implementations of the subject matter described in this specification can be implemented in a computing system that includes a back end component, e.g., as a data server, or that includes a middleware component, e.g., an application server, or that includes a frontAttorney Docket No.: 43374-0851WO1 end component, e.g., a client computer having a graphical user interface, a web browser, or an app through which a user can interact with an implementation of the subject matter described in this specification, or any combination of one or more such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication, e.g., a communication network. Examples of communication networks include a local area network (LAN) and a wide area network (WAN), e.g., the Internet.

[0119] The computing system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship with each other. In some implementations, a server transmits data, e.g., an HTML page, to a user device, e.g., for purposes of displaying data to and receiving user input from a user interacting with the device, which acts as a client. Data generated at the user device, e.g., a result of the user interaction, can be received at the server from the device.

[0120] While this specification contains many specific implementation details, these should not be construed as limitations on the scope of what may be claimed, but rather as descriptions of features that may be specific to particular implementations. Certain features that are described in this specification in the context of separate implementations can also be implemented, in combination, in a single implementation. Conversely, various features that are described in the context of a single implementation can also be implemented in multiple implementations, separately, or in any sub-combination. Moreover, although previously described features may be described as acting in certain combinations and even initially claimed as such, one or more features from a claimed combination can, in some cases, be excised from the combination, and the claimed combination may be directed to a sub- combination or variation of a sub-combination.

[0121] As used in this disclosure, the terms “a,” “an,” or “the” are used to include one or more than one unless the context clearly dictates otherwise. The term “or” is used to refer to a nonexclusive “or” unless otherwise indicated. The statement “at least one of A and B” has the same meaning as “A, B, or A and B.” In addition, the phraseology or terminology employed in this disclosure, and not otherwise defined, is for the purpose of description only and not of limitation. Any use of section headings is intended to aid reading of the document and is not to be interpreted as limiting; information that is relevant to a section heading may occur within or outside of that particular section.Attorney Docket No.: 43374-0851WO1

[0122] As used in this disclosure, the term “about” or “approximately” can allow for a degree of variability in a value or range, for example, within 10%, within 5%, or within 1% of a stated value or of a stated limit of a range.

[0123] As used in this disclosure, the term “substantially” refers to a majority of, or mostly, as in at least about 50%, 60%, 70%, 80%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.9%, 99.99%, or at least about 99.999% or more.

[0124] Values expressed in a range format should be interpreted in a flexible manner to include not only the numerical values explicitly recited as the limits of the range, but also the individual numerical values or sub-ranges encompassed within that range as if each numerical value and sub-range is explicitly recited. For example, a range of “0.1% to about 5%” or “0.1% to 5%” should be interpreted to include about 0.1% to about 5%, as well as the individual values (for example, 1%, 2%, 3%, and 4%) and the sub-ranges (for example, 0.1% to 0.5%, 1.1% to 2.2%, 3.3% to 4.4%) within the indicated range. The statement “X to Y” has the same meaning as “about X to about Y,” unless indicated otherwise. Likewise, the statement “X, Y, or Z” has the same meaning as “about X, about Y, or about Z,” unless indicated otherwise.

[0125] Particular implementations of the subject matter have been described. Other implementations, alterations, and permutations of the described implementations are within the scope of the following claims as will be apparent to those skilled in the art. While operations are depicted in the drawings or claims in a particular order, such operations are not required to be performed in the particular order shown or in sequential order, or that all illustrated operations be performed (some operations may be considered optional), to achieve desirable results. In certain circumstances, multitasking or parallel processing (or a combination of multitasking and parallel processing) may be advantageous and performed as deemed appropriate.

[0126] Moreover, the separation or integration of various system modules and components in the previously described implementations are not required in all implementations, and the described components and systems can generally be integrated together or packaged into multiple products.

[0127] Accordingly, the previously described example implementations do not define or constrain the present disclosure. Other changes, substitutions, and alterations are also possible without departing from the spirit and scope of the present disclosure.

[0128] The foregoing description of the specific implementations can be readily modified and / or adapted for various applications. Therefore, such adaptations andAttorney Docket No.: 43374-0851WO1 modifications are intended to be within the meaning and range of equivalents of the disclosed implementations, based on the teaching and guidance presented herein.

[0129] The breadth and scope of the present disclosure should not be limited by any of the above-described example implementations but should be defined only in accordance with the following claims and their equivalents. Accordingly, other implementations also are within the scope of the claims.

[0130] Although the disclosed inventive concepts include those defined in the attached claims, it should be understood that the inventive concepts can also be defined in accordance with the following embodiments.

[0131] In addition to the embodiments of the attached claims and the embodiments described above, the following numbered embodiments are also innovative.

[0132] Embodiment 1 is a computer implemented method, the method comprising: training a Time-series Dense Encoder (TiDE) model on a corpus of historical weather data for a particular geographic region, wherein the trained TiDE model forecasts weather for the particular geographic region; training a machine learning model on a corpus of historical power output and weather data for a particular wind farm within the particular geographic region, wherein the trained machine learning model predicts the power output out of the wind farm for a given weather condition; forecasting, using the TiDE model, the weather in the particular geographic region; predicting the power output of the wind farm by providing the forecasted weather to the machine learning model; and forecasting a required power output for a plurality of thermal generators based on the predicted power output of the wind farm by calculating the output of a fixed effects ordinary least squares (OLS) regression model.

[0133] Embodiment 2 is the method of embodiment 1, wherein the TiDE model is a Multi-layer Perceptron (MLP) based encoder-decoder architecture.

[0134] Embodiment 3 is the method of embodiment 2, wherein the TiDE model comprises an encoder with four residual blocks each containing fully connected layers of 128 neurons, and a decoder with four residual blocks each containing fully connected layers of 128 neurons.

[0135] Embodiment 4 is the method of any of embodiments 1 through 3, wherein the TiDE model is trained using a mean square error loss function and an Adaptive Moment Estimation (Adam) optimization.

[0136] Embodiment 5 is the method of any of embodiments 1 through 4, wherein the fixed effects OLS regression model is a non-logarithmic statistical model.Attorney Docket No.: 43374-0851WO1

[0137] Embodiment 6 is the method of embodiment 5, wherein calculating output of the non-logarithmic statistical model comprises identifying a value for Gt in the equation ^^௧ൌ^^ ^ ^^^^^௧ ^ ^^ଶ^^௧ ^ ^^ଷ^^௧ ^ ^^ସ^^^௧ െ ^^௧ି^^ ^ ^^ହ^^^௧ െ ^^௧ି^^ ^ ^^௧,^ ^ ^^^^௧ ^ ^^௧,^ ^ ^^.

[0138] Embodiment 7 is the method of any of embodiments 1 through 6, wherein the

[0139] Embodiment 8 is a computer program carrier encoded with a computer program, the program comprising instructions that are operable, when executed by data processing apparatus, to cause the data processing apparatus to perform operations comprising: training a Time-series Dense Encoder (TiDE) model on a corpus of historical weather data for a particular geographic region, wherein the trained TiDE model forecasts weather for the particular geographic region; training a machine learning model on a corpus of historical power output and weather data for a particular wind farm within the particular geographic region, wherein the trained machine learning model predicts the power output out of the wind farm for a given weather condition; forecasting, using the TiDE model, the weather in the particular geographic region; predicting the power output of the wind farm by providing the forecasted weather to the machine learning model; and forecasting a required power output for a plurality of thermal generators based on the predicted power output of the wind farm by calculating the output of a fixed effects ordinary least squares (OLS) regression model.

[0140] Embodiment 9 is the method of embodiment 8, wherein the TiDE model is a Multi-layer Perceptron (MLP) based encoder-decoder architecture.

[0141] Embodiment 10 is the method of embodiment 9, wherein the TiDE model comprises an encoder with four residual blocks each containing fully connected layers of 128 neurons, and a decoder with four residual blocks each containing fully connected layers of 128 neurons.

[0142] Embodiment 11 is the method of any of embodiments 8 through 10, wherein the TiDE model is trained using a mean square error loss function and an Adaptive Moment Estimation (Adam) optimization.

[0143] Embodiment 12 is the method of any of embodiments 8 through 11, wherein the fixed effects OLS regression model is a non-logarithmic statistical model.

[0144] Embodiment 13 is the method of embodiment 12, wherein calculating output of the non-logarithmic statistical model comprises identifying a value for Gt in the equation ^^௧ൌ^^ ^ ^^^^^௧ ^ ^^ଶ^^௧ ^ ^^ଷ^^௧ ^ ^^ସ^^^௧ െ ^^௧ି^^ ^ ^^ହ^^^௧ െ ^^௧ି^^ ^ ^^௧,^ ^ ^^^^௧ ^ ^^௧,^ ^ ^^.Attorney Docket No.: 43374-0851WO1

[0145] Embodiment 14 is the method of any of embodiments 8 through 13, wherein the weather data comprises wind speed, wind direction, precipitation, and temperature.

[0146] Embodiment 15 is a system comprising: one or more computers and one or more storage devices storing instructions that are operable, when executed by the one or more computers, to cause the one or more computers to perform the method of any one of embodiments 1 to 14.

[0147] The foregoing description is provided in the context of one or more particular implementations. Various modifications, alterations, and permutations of the disclosed implementations can be made without departing from scope of the disclosure. Thus, the present disclosure is not intended to be limited only to the described or illustrated implementations but is to be accorded the widest scope consistent with the principles and features disclosed herein.

[0148] In other words, although this disclosure has been described in terms of certain embodiments and generally associated methods, alterations and permutations of these embodiments and methods will be apparent to those skilled in the art. Accordingly, the above description of example embodiments does not define or constrain this disclosure. Other changes, substitutions, and alterations are also possible without departing from the spirit and scope of this disclosure.

Claims

Attorney Docket No.: 43374-0851WO1 WHAT IS CLAIMED IS:

1. A computer-implemented method comprising: training a Time-series Dense Encoder (TiDE) model on a corpus of historical weather data for a particular geographic region, wherein the trained TiDE model forecasts weather for the particular geographic region; training a machine learning model on a corpus of historical power output and weather data for a particular wind farm within the particular geographic region, wherein the trained machine learning model predicts the power output out of the wind farm for a given weather condition; forecasting, using the TiDE model, the weather in the particular geographic region; predicting the power output of the wind farm by providing the forecasted weather to the machine learning model; and forecasting a required power output for a plurality of thermal generators based on the predicted power output of the wind farm by calculating the output of a fixed effects ordinary least squares (OLS) regression model.

2. The method of claim 1, wherein the TiDE model is a Multi-layer Perceptron (MLP) based encoder-decoder architecture.

3. The method of claim 2, wherein the TiDE model comprises an encoder with four residual blocks each containing fully connected layers of 128 neurons, and a decoder with four residual blocks each containing fully connected layers of 128 neurons.

4. The method of any of claims 1 through 3, wherein the TiDE model is trained using a mean square error loss function and an Adaptive Moment Estimation (Adam) optimization.

5. The method of any of claims 1 through 4, wherein the fixed effects OLS regression model is a non-logarithmic statistical model.

6. The method of claim 5, wherein calculating output of the non-logarithmic statisticalmodel comprises identifying a value for Gt in the equation ^^௧ ൌ ^^ ^ ^^^^^௧ ^ ^^ଶ^^௧ ^ ^^ଷ^^௧ ^^^ସ^^^௧ െ ^^௧ି^^ ^ ^^ହ^^^௧ െ ^^௧ି^^ ^ ^^௧,^ ^ ^^^^௧ ^ ^^௧,^ ^ ^^.Attorney Docket No.: 43374-0851WO1 7. The method of any of claims 1 through 6, wherein the weather data comprises wind speed, wind direction, precipitation, and temperature.

8. A non-transitory, computer-readable medium storing one or more instructions executable by a computer system to perform operations comprising: training a Time-series Dense Encoder (TiDE) model on a corpus of historical weather data for a particular geographic region, wherein the trained TiDE model forecasts weather for the particular geographic region; training a machine learning model on a corpus of historical power output and weather data for a particular wind farm within the particular geographic region, wherein the trained machine learning model predicts the power output out of the wind farm for a given weather condition; forecasting, using the TiDE model, the weather in the particular geographic region; predicting the power output of the wind farm by providing the forecasted weather to the machine learning model; and forecasting a required power output for a plurality of thermal generators based on the predicted power output of the wind farm by calculating the output of a fixed effects ordinary least squares (OLS) regression model.

9. The medium of claim 8, wherein the TiDE model is a Multi-layer Perceptron (MLP) based encoder-decoder architecture.

10. The medium of claim 9, wherein the TiDE model comprises an encoder with four residual blocks each containing fully connected layers of 128 neurons, and a decoder with four residual blocks each containing fully connected layers of 128 neurons.

11. The medium of any of claims 8 through 10, wherein the TiDE model is trained using a mean square error loss function and an Adaptive Moment Estimation (Adam) optimization.

12. The medium of any of claims 8 through 11, wherein the fixed effects OLS regression model is a non-logarithmic statistical model.Attorney Docket No.: 43374-0851WO1 13. The medium of claim 12, wherein calculating output of the non-logarithmicstatistical model comprises identifying a value for Gt in the equation ^^௧ ^^ ^ ^ ^^ଶ^^௧ ^^^ ^ ^ ^ ^ ^^.

14. The medium of any of claims 8 through 13, wherein the weather data comprises wind speed, wind direction, precipitation, and temperature.

15. A computer-implemented system, comprising: one or more computers; and one or more computer memory devices interoperably coupled with the one or more computers and having tangible, non-transitory, machine-readable media storing one or more instructions that, when executed by the one or more computers, perform one or more operations comprising: training a Time-series Dense Encoder (TiDE) model on a corpus of historical weather data for a particular geographic region, wherein the trained TiDE model forecasts weather for the particular geographic region; training a machine learning model on a corpus of historical power output and weather data for a particular wind farm within the particular geographic region, wherein the trained machine learning model predicts the power output out of the wind farm for a given weather condition; forecasting, using the TiDE model, the weather in the particular geographic region; predicting the power output of the wind farm by providing the forecasted weather to the machine learning model; and forecasting a required power output for a plurality of thermal generators based on the predicted power output of the wind farm by calculating the output of a fixed effects ordinary least squares (OLS) regression model.

16. The system of claim 15, wherein the TiDE model is a Multi-layer Perceptron (MLP) based encoder-decoder architecture.Attorney Docket No.: 43374-0851WO1 17. The system of claim 16, wherein the TiDE model comprises an encoder with four residual blocks each containing fully connected layers of 128 neurons, and a decoder with four residual blocks each containing fully connected layers of 128 neurons.

18. The system of any of claims 15 through 17, wherein the TiDE model is trained using a mean square error loss function and an Adaptive Moment Estimation (Adam) optimization.

19. The system of any of claims 15 through 18, wherein the fixed effects OLS regression model is a non-logarithmic statistical model.

20. The system of claim 19, wherein calculating output of the non-logarithmic statistical model comprises identifying a value for Gt in the equation ^^௧ ൌ ^^ ^ ^^^^^௧ ^ ^^ଶ^^௧ ^^^ଷ^^௧ ^ ^^ସ^^^௧ െ ^^௧ି^^ ^ ^^ହ^^^௧ െ ^^௧ି^^ ^ ^^௧,^ ^ ^^^^௧ ^ ^^௧,^ ^ ^^.