Weather forecasting using diffusion neural networks

By employing a diffusion neural network to process surface graph representations, the system generates probabilistic weather forecasts that are more accurate, physically consistent, and computationally efficient than existing deterministic ML models, addressing the limitations of current medium-range weather forecasting techniques.

WO2025133387A1PCT designated stage expired Publication Date: 2025-06-26DEEPMIND TECH LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
PCT/EP2024/088369
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-22
Filing Date
2024-12-23
Publication Date
2025-06-26

AI Technical Summary

Technical Problem

Current machine learning (ML) forecast models for medium-range weather forecasting are largely deterministic, unable to estimate uncertainty or predict probabilities, and lack physical consistency, especially at longer lead times.

Method used

The use of a diffusion neural network to generate probabilistic weather forecasts by processing a mesh representation of a graph of a surface, enabling the prediction of multiple weather variables over long periods and high resolutions, while being computationally efficient and reducing latency.

Benefits of technology

This approach produces more skillful and calibrated forecasts than traditional ensemble methods, maintaining physical consistency and significantly reducing the time required to generate forecasts, from hours on supercomputers to mere seconds on hardware accelerators.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2024088369_26062025_PF_FP_ABST
    Figure EP2024088369_26062025_PF_FP_ABST
Patent Text Reader

Abstract

Methods, systems, and apparatus, including computer programs encoded on computer storage media, for predicting weather using diffusion neural networks.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] WEATHER FORECASTING USING DIFFUSION NEURAL NETWORKS

[0002] CROSS-REFERENCE TO RELATED APPLICATION

[0003] This application claims priority to U.S. Provisional Application No. 63 / 614,461, filed on December 22, 2023. The disclosure of the prior application is considered part of and is incorporated by reference in the disclosure of this application.

[0004] BACKGROUND

[0005] This specification relates to using neural networks to perform weather forecasting.

[0006] Neural networks are machine learning models that employ one or more layers of nonlinear units to predict an output for a received input. Some neural networks include one or more hidden layers in addition to an output layer. The output of each hidden layer is used as input to the next layer in the network, i.e., the next hidden layer or the output layer. Each layer of the network generates an output from a received input in accordance with current values of a respective set of parameters.

[0007] SUMMARY

[0008] This specification describes a system implemented as computer programs on one or more computers in one or more locations that generates predicted weather forecasts. In particular, the system generates a predicted weather forecast by processing a mesh representation of a graph of a surface using a diffusion neural network.

[0009] A weather forecast, as used in this specification, is a prediction of the values of one or more weather properties, e.g., atmospheric properties, surface properties, or both, at one or more locations at a future time.

[0010] Particular embodiments of the subject matter described in this specification can be implemented so as to realize one or more of the following advantages.

[0011] Probabilistic weather forecasting is critical for decision-making in high-impact domains. For example, in many of these domains, e.g., flood forecasting, energy system planning or transportation routing, quantifying the forecast uncertainty and capturing extreme events is essential to guide important cost-benefit trade-offs and mitigation measures. Traditional probabilistic approaches rely on producing ensembles from physics-based models, which sample from the joint distribution of spatio-temporally coherent weather trajectories but are expensive to run. A more computationally-efficient alternative is to use a machine learning (ML) forecast model to generate the ensemble. However, state-of-the-art ML forecast models for medium -range weather are largely trained to produce deterministic forecasts. Despite achieving great deterministic skill scores, they cannot estimate uncertainty or predict probabilities. Moreover, they lack physical consistency, a limitation that grows at longer lead time and further negatively impacts their ability to characterize the underlying joint distribution.

[0012] This specification, on the other hand, describes a ML-based generative model for ensemble weather forecasting. By making use of a diffusion neural network, the described techniques can generate probabilistic forecasts for a large number of weather variables, e.g., as many as 84 weather variables or more, for a long period of time and at high-resolution, e.g., up to 15 days at 1 degree resolution or .25 degree resolution, while being much more computationally-efficient and having significantly lower latency than traditional probabilistic systems. For example, when implementing ensemble members on respective hardware accelerators, the described techniques can generate a 15 day forecast at 1 degree resolution in ~60 seconds per ensemble member, significantly quicker than traditional probabilistic forecasting approaches. Additionally, all of the forecasts in the ensemble can be generated in parallel by deploying respective instances of the diffusion neural network on multiple hardware accelerators, thereby avoiding incurring additional latency.

[0013] Moreover, the described techniques are more skillful than, e.g., ENS, a top operational ensemble forecast, while maintaining good calibration and physically consistent power spectra. For comparison, traditional physics-based ensemble forecasts such as those produced by ENS take hours on a supercomputer with tens of thousands of processors.

[0014] The details of one or more embodiments of the subject matter of this specification are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages of the subject matter will become apparent from the description, the drawings, and the claims.

[0015] BRIEF DESCRIPTION OF THE DRAWINGS

[0016] FIG. 1 shows an example forecasting system.

[0017] FIG. 2 is a flow diagram of an example process for generating a forecast.

[0018] FIG. 3 is a flow diagram of an example process for performing a sampling iteration.

[0019] FIG. 4 shows an example of the operation of the system. FIG. 5 shows an example configuration of the diffusion neural network.

[0020] Like reference numbers and designations in the various drawings indicate like elements.

[0021] DETAILED DESCRIPTION

[0022] FIG. 1 shows an example forecasting system 100. The forecasting system 100 is an example of a system implemented as computer programs on one or more computers in one or more locations, in which the systems, components, and techniques described below can be implemented.

[0023] The system 100 generates predicted weather forecasts 120 from data 102 characterizing a current weather state at a current time step.

[0024] The current weather state is defined on a latitude-longitude grid over a surface of a body, e.g., over some or all of the surface of a celestial object, e.g., over the Earth or a different planet and includes, for each of a plurality of points on the latitude-longitude grid, current weather properties at the point as of the current time step.

[0025] The current weather properties can include properties of the surface at the point, properties of the atmosphere above the surface at the point, or both. For example, surface properties can include any of temperature, wind speed, wind direction, air pressure, precipitation levels, and so on. Atmospheric properties can include, for each of multiple vertical levels, any one of humidity, temperature, wind speed, wind direction, geopotential, vertical wind velocity, and so on.

[0026] The current weather state can also include one or more weather-independent properties for each point, e.g., latitude and longitude, time of day, time of year, and optionally other properties, e.g., solar radiation.

[0027] Table 1 shows one example of the set of surface and atmospheric properties that can be included in the current weather data 120 and one example of the set of vertical levels (“pressure levels”) at which each atmospheric property can be measured.

[0028] Table 1 Table 2 shows another example of the properties that can be included in the current weather data 120.

[0029] Table 2

[0030] In Table 2, the “Type” column indicates whether the variable represents a static property, a time-varying single-level property (e.g. surface variables are included), or a time varying atmospheric property. The “Variable name” and “Short name” columns are the European Centre for Medium-Range Weather Forecasts (ECMWF)’s labels. The “ECMWF Parameter ID” column is ECMWF’s numeric label, and can be used to construct the URL for ECMWF’s description of the variable, by appending it as suffix to the following prefix, replacing “ID” with the numeric code: apps.ecmwf.int / codes / grib / param-db / ?id=ID. The “Role” column indicates whether the variable is something the system both takes as input and predicts, or only uses as input context (the double horizontal line separates predicted from input-only variables, to make the partitioning more visible). In the example of Table 2, the 13 atmospheric pressure levels are taken as input and predicted by the system are: 50, 100, 150, 200, 250, 300, 400, 500, 600, 700, 850, 925, and 1000 hPa.

[0031] A weather forecast, as used in this specification, can include a respective predicted weather state for each future time step in a sequence of time steps that starts at a current time step and includes multiple future time steps that follow the current time step. For example, the weather forecasts generated by the system 100 can be referred to as “medium-range” forecasts, e.g., because the total time interval spanned by the future time steps is between one and two weeks from the current time step.

[0032] The interval between the current time step and the first future time step in the sequence and between any two future time steps in the sequence can generally be any appropriate time interval, e.g., between three and twenty four hours, e.g., six or twelve hours. For example, the system 100 can generate predictions at each of multiple twelve hour intervals into the future starting from the weather data 102 for the current time step.

[0033] A predicted weather state generally includes a prediction of the values of the weather-dependent properties for each point, e.g., a prediction of the atmospheric properties, surface properties, or both, at the future time step.

[0034] In some implementations, the system 100 generates an “ensemble weather forecast” that includes multiple individual weather forecasts that each include respective predicted weather states for each of the future time steps. The individual forecasts in the ensemble can approximate samples from the distribution of future weather scenarios. Thus, the forecasts can model different plausible weather outcomes and help capture the occurrence of extreme events, i.e., that may not be present if the system 100 just generated a single, most likely weather forecast.

[0035] In general, an ensemble weather forecast includes multiple weather forecasts, which can be referred to as a “first” weather forecast and one or more “additional” weather forecasts.

[0036] To generate a predicted weather forecast 120, e.g., one of the forecasts in the ensemble, the system 100 uses a diffusion neural network 110 to iteratively generate the predicted future state at each time step in the forecast.

[0037] That is, the system 100 uses the diffusion neural network 110 to, for any given future time step, perform multiple sampling iterations to “denoise” a noisy representation of the weather state for the future time step in order to generate the predicted future state at the given future time step in the forecast.

[0038] The diffusion neural network 110 is a neural network that, at any given sampling iteration, processes a diffusion input for the sampling iteration that includes (i) the representation of the weather state for the future time step and (ii) a representation of an input weather state for the future time step to generate a denoising output. As will be described in more detail below, this denoising output defines an estimate of the noise component of the representation of the weather state given the input weather state for the future time step.

[0039] For example, the diffusion neural network 110 can process the representations of weather states as a graph representation, e.g., as data representing a graph that has nodes that represent points on the grid. More specifically, in this example the diffusion neural network 110 processes, for each point on the grid, a graph with data derived from the values of properties for the points on the grid.

[0040] This example is described in more detail below with reference to FIG. 5.

[0041] Generating predicted future states using the diffusion neural network 110 will be described in more detail below.

[0042] Once generated, the system 100 or another system can use the weather forecasts, i.e., either a single individual forecast or an ensemble of weather forecasts, for any of a variety of purposes. Some examples now follow.

[0043] For example, the predictions can be used for “everyday” forecasting and the system 100 can generate a user interface presentation that visually displays data characterizing the predicted values of one or more of the weather properties at each of one or more of the future time steps and provide the user interface presentation for viewing by users, e.g., in a user interface of a weather software application running on a user device, on a web page displayed in web browsers running on user devices, or by including the presentation as a “green screen” or other visualization in a streaming video or television broadcast. As a particular example, a user can submit a query specifying a point on the grid, and the system 100 can provide, in response, weather data characterizing predicted weather at the point at one or more of the future time steps. Instead of or in addition to the user interface presentation, the system 100 can generate speech describing the predicted value(s) and provide the speech for playback to users.

[0044] As another example, the predictions can be used for “extreme weather forecasting”, e.g., cyclone tracking, extreme heat, extreme cold, high and low rainfall, and so on. In these examples, the system 100 can generate and provide an alert whenever the predicted value of a given property satisfies a specified threshold (that would indicate that an extreme weather event is occurring) in one or more of the forecasts in the ensemble or whenever the average, maximum, or minimum value of the given property across the forecasts in the ensemble satisfies the threshold.

[0045] In another example application the system 100 may be incorporated into an energy management system. The energy management system may be configured to obtain context weather data characterizing the weather in a real-world location of a renewable energy generation facility such as a wind power or solar energy generation facility at preceding time steps. The context weather data may be processed as described above, and the renewable energy generation facility, e.g., the wind power or solar energy generation facility, may be controlled in response to the predicted weather states characterizing the predicted weather in the real-world location of the renewable energy generation facility at the corresponding future time step(s). For example, in a solar energy farm solar panels may be tilted or covered to protect the panels from predicted weather hydrometeors of greater than a threshold severity; or in a wind farm the energy generation facility may be configured to maximize output based on the predicted weather; or one or more other power generation sources on the same electrical grid as the renewable energy generation facility may be controlled to increase or decrease power from the other power generation sources for grid balancing in response to a predicted power output from the renewable energy generation facility based on the predicted weather. In a related energy management system instead of controlling the renewal energy generation facility the predicted weather data characterizing the predicted weather in the real-world location of the renewable energy generation facility may be used to send a signal to a consumer of power on the electrical grid to which the renewable energy generation facility is connected, for controlling one or more power-consuming devices of the consumer for load balancing, e.g., in a smart grid, in response to a predicted power output from the renewable energy generation facility based on the predicted weather.

[0046] In another example application the system 100 may be incorporated into a flood early-warning system. The flood early-warning system may be configured to obtain context weather data characterizing the weather in a real-world location of an at-risk area at preceding time steps. The context weather data may be processed as described above; and the flood early-warning system may output a warning, in response to the predicted weather states characterizing the predicted weather in the real-world location of the at- risk area at the corresponding future time step. For example the warning may be provided in response to the predicted weather forecasting greater than a threshold level of precipitation. A similar system may be used to warn of potential landslides. The warning may be issued via one or more channels, e.g., television, radio, internet, mobile phones, public transport, or public alarms. In some implementations the warning is a public warning such as an audible and / or visual alarm in the real-world location of the at-risk area that can warn of local danger to life or property. In another example application the system 100 may be incorporated into an air or marine traffic control system. The air or marine traffic control system may be configured to obtain context weather data characterizing the weather in a real-world location of one or more air or marine craft at preceding time steps. The context weather data may be processed as described above; and the air or marine traffic control system may output a signal for controlling the flight patterns or marine craft routes of the one or more aircraft or marine craft in the real-world location in response to the predicted weather states characterizing the predicted weather in the real-world location at the corresponding future time step. The signal may be a warning signal or a routing signal; in some implementations it may be automatically provided to the air or marine craft, e.g., for the purpose of enabling the air or marine craft to take evasive action. In some implementations, the system is an air traffic control system and the real-world location may then be the location of an airport. The signal may be output in response to the predicted weather having greater than a threshold level of severity, e.g., greater than a threshold level of precipitation or hydrometeors, or greater than a threshold level of wind (speed), or greater than a threshold level of a particular wind behavior (e.g., characterized by wind speed and / or direction).

[0047] In another example application the system 100 may be incorporated into an energy management system configured to obtain context weather data characterizing the weather in a real-world location of a building or industrial facility at preceding time steps. The context weather data may be processed as described above; and control shutters, or a ventilation system, or a temperature control system, of the building or industrial facility in response to the predicted weather states characterizing the predicted weather in the real- world location of the building or industrial facility at the corresponding future time step. Such a system may be used to control a temperature or humidity of the building or industrial facility, or to keep it dry, or to protect the building or industrial facility.

[0048] FIG. 2 is a flow diagram of an example process 200 for generating a predicted weather forecast. For convenience, the process 200 will be described as being performed by a system of one or more computers located in one or more locations. For example, a forecasting system, e.g., the forecasting system 100 of FIG.1, appropriately programmed, can perform the process 200.

[0049] For example, the predicted weather forecast generated by performing the process 200 can be one of multiple forecasts in an ensemble of forecasts generated by the system. Thus, the system can perform multiple different iterations of the process 200, e.g., in parallel, starting from the current weather state data in order to generate multiple, potentially different, weather forecasts to be included in the ensemble.

[0050] As a specific example, the system can perform each iteration of the process 200 in parallel using a respective instance of the diffusion neural network deployed on a corresponding set of one or more hardware devices, e.g., sets of one or more hardware accelerators for accelerating machine learning computations, e.g., tensor processing units (TPUs), graphics processing units (GPUs), or other ASICs.

[0051] Thus, the system can effectively leverage parallel processing hardware to improve the realism and usefulness of the forecasts without increasing forecast generation latency.

[0052] The system obtains data characterizing a current weather state for a current time step (step 202). As described above, the current weather state includes, for each of a plurality of points on a latitude-longitude grid over a surface of a body, current weather properties at the point as of the current time step.

[0053] The system generates a weather forecast that includes a respective predicted weather state for each future time step in a sequence of time steps that (i) starts at the current time step and (ii) includes a plurality of future time steps that follow the current time step (step 204).

[0054] As part of generating the weather forecast, the system performs the below steps for each of the future time steps, starting from the first future time step in the sequence and continuing until the last future time step in the sequence.

[0055] The system initializes a representation of a weather state for the future time step (step 206).

[0056] Generally, the representation Z of the weather state includes noisy values that are sampled from a noise distribution.

[0057] In particular, to initialize the representation of the weather state for the future time step, the system can sample noisy values from a noise probability distribution.

[0058] As a particular example of this, the system can sample a respective noise value for each property for each point on the grid from an appropriate distribution, e.g., by sampling (e.g., isotropic) Gaussian noise.

[0059] As another particular example of this, the system can sample (e.g., isotropic) Gaussian noise on a sphere corresponding to the surface and then project the sampled Gaussian noise onto the grid to yield a respective noise value for each property for each point on the grid. Such an approach can allow the noise to have a flatter spherical harmonic power spectrum compared to sampling Gaussian noise on the latitude-longitude grid.

[0060] When the system is generating an ensemble forecast, initializing the representation by sampling introduces diversity into the ensemble, e.g., because different forecasts in the ensemble will be generated from different initialized representations due to the effect of the sampling.

[0061] The system updates the representation of the weather state for the future time step to generate a final representation of the weather state for the future time step (step 208).

[0062] As part of this updating, the system performs a plurality of sampling iterations using the diffusion neural network. That is, the system updates the representation of the weather state at each of the multiple sampling iterations using the diffusion neural network.

[0063] Each sampling iteration z has a corresponding noise level so that the representation is initialized at z = 0 and the final sampling iteration is at V-1, resulting in a final representation ZN.

[0064] Generally, the noise level <Jtdecreases as the iterations increment, so that earlier iterations are associated with higher noise levels while later iterations are associated with lower noise levels, with the index being associated with a noise level of zero, i.e., so that the final representation is “denoised.”

[0065] The system receives an input defining the noise levels associated with each of the iterations, e.g., an input that directly specifies the noise level for each iteration or an input that specifies a schedule for adjusting the noise level between iterations.

[0066] Thus, in the case that each future time step is conditioned on the weather states for the preceding two time steps in the sequence the representation Z+1at iteration z+1 for future time step t can be expressed as: where rerepresents the output of the sampling iteration that is determined using the diffusion neural network. In some cases realso operates on <ji+1.

[0067] Performing a sampling iteration using the diffusion neural network will be described below with reference to FIG. 3.

[0068] The system then generates the respective predicted weather state for the future time step from the final representation of the weather state (step 210), i.e., the representation after the last sampling iteration. For example, the representation can define, for a given property, a delta, i.e., a difference, between the value of the property at the future time step and the value for the property at the preceding time step. The delta can be, e.g., in the output space or in a normalized space.

[0069] As another example, the representation can define, for a given property, a raw or a normalized value of the property at the future time step.

[0070] That is, in some implementations, for some or all of the properties, the representation defines values of the property that are in a normalized space. For example, in the normalized space, each output value can be normalized to have zero mean and unit variance. For example, a respective mean and variance of each output value can be estimated from the corresponding output values at the points on the latitude-longitude grid. The estimated mean may then be subtracted from each of the corresponding output values at the points on the latitude-longitude grid, and the corresponding output values scaled using the estimated variance (e.g. dividing by the corresponding standard deviation).

[0071] In these cases, when the output for a given property is a normalized delta (e.g., a difference between normalized values) the system can first invert the normalization to generate an un-normalized delta and then add the un-normalized value for the property at the preceding time step.

[0072] When the output for a given property is a normalized value, the system can invert the normalization to generate the un-normalized value.

[0073] FIG. 3 is a flow diagram of an example process 300 for performing a sampling iteration. For convenience, the process 300 will be described as being performed by a system of one or more computers located in one or more locations. For example, a forecasting system, e.g., the forecasting system 100 of FIG.1, appropriately programmed, can perform the process 300.

[0074] The system generates a denoising output for the sampling iteration for the sampling iteration (step 302).

[0075] In particular, the system generates the denoising output at least in part by processing a first diffusion input for the sampling iteration that includes (i) the representation of the weather state for the future time step and (ii) a representation of an input weather state for the future time step using a diffusion neural network to generate a first denoising output. The first diffusion input can also include the noise level for the sampling iteration. The first denoising output can, e.g., represent a prediction of the actual weather state for the future time step or a prediction of the noise that has been added to the actual weather state for the future time step to generate the representation of the weather state as of the sampling iteration.

[0076] For the first future time step in the sequence, the input weather state includes the current weather state for the current time step.

[0077] In some implementations, the system uses different representations of the current weather state for different forecasts in the ensemble. For example, each different representation can represent a different estimate of the current weather state.

[0078] The system can obtain these different estimates in a variety of ways.

[0079] For example, the system can obtain different weather states from different external sources that each provide measured or observed weather properties. In other words, the system can receive, as part of the data characterizing the current weather state, the multiple different representations of the current weather state.

[0080] As another example, the multiple different estimates can be generated by the system by applying perturbations to a single initial current weather state. Thus, in this example the system can receive a single representation of the current weather state and generate multiple different representations of the current weather state from the received representation by perturbing the values in the received representation.

[0081] In these implementations, the system leverages both the differences in the representations and the stochasticity involved in the sampling iterations and in the sampling of the noise for the initialized representation of the weather state to cause the individual forecasts to capture diverse weather outcomes.

[0082] In some other implementations, the system uses the same representation of the current weather state for each forecast in the ensemble. In these implementations, the system leverages the stochasticity involved in the sampling iterations and in the sampling of the noise for the initialized representation of the weather state to cause the individual forecasts to capture diverse weather outcomes.

[0083] For each subsequent future time step in the sequence, the input weather state includes the respective predicted weather state for a preceding future time step in the sequence.

[0084] In some implementations, the input weather state for a given future time step can include the weather states (predicted or actual) from two or more preceding time steps in the sequence. That is, in these implementations, for the first future time step in the sequence, the input weather state includes the current weather state and a preceding weather state for a preceding time step that precedes the current time step. For the second future time step in the sequence, the input weather state includes the respective predicted weather state for the first future time step in the sequence and the current weather state. For each future time step that is after the second future time step in the sequence, the input weather state includes the respective predicted weather states for two or more preceding future time steps in the sequence.

[0085] Including this additional context in the input weather state can provide the diffusion neural network more context to more accurately “denoise” the representation for the future time step.

[0086] In some implementations, as part of processing the first diffusion input, the system applies one or more preconditioning functions to the input to the diffusion neural network.

[0087] For example, the first denoising output for a given future time step t at a given iteration can be represented as: wheren(<j) and cnoise((T) are respective preconditioning functions that are applied to the noise level for the given iteration, ferepresents the operations of the diffusion neural network, is the representation of the weather state of the future time step, are the weather states for the two preceding time steps. Examples of preconditioning functions that can be used are described in T. Karras, M. Aittala, T. Aila, and S. Laine. Elucidating the design space of diffusion-based generative models. Advances in Neural Information Processing Systems, 35:26565-26577, 2022.

[0088] In some implementations, the first denoising output is the denoising output.

[0089] In some other implementations, the system also processes a second diffusion input for the sampling iteration that includes the representation of the weather state for the future time step, i.e., the representation that is being denoised, using the diffusion neural network to generate a second denoising output. This second diffusion input is generally an “unconditional” input that does not include the representation of the input weather state for the future time step.

[0090] The system can then generate the denoising output using the first denoising output and the second denoising output. As a particular example, the system can generate a final denoising output by combining the first denoising output and the second denoising output in accordance with a classifier-free guidance weight for the sampling iteration. For example, the system can set the final denoising output equal to (1+w) * the first denoising output - vi’*the additional denoising output, where w is the guidance weight and * denotes multiplication. The guidance weights for the iterations can be received as input by the system.

[0091] The system then updates the representation of the future weather state using the denoising output (step 304), e.g., by, for each sampling iteration other than the last sampling iteration, applying a diffusion sampler using the denoising output.

[0092] For example, the system can determine the output of a denoiser, i.e., an initial estimate of the final representation, from the denoising output and then apply the diffusion sampler (also referred to as a diffusion “solver”) to the output of the denoiser.

[0093] Generally, the system computes the denoiser output by combining the current representation with the denoising output. As a particular example, when the preconditioning functions are used, the system can determine the output of the denoiser as: and cout(a) are respective preconditioning functions that are applied to the noise level for the given iteration. Examples of preconditioning functions that can be used are described in T. Karras, M. Aittala, T. Aila, and S. Laine. Elucidating the design space of diffusion-based generative models. Advances in Neural Information Processing Systems, 35:26565-26577, 2022.

[0094] The diffusion sampler can generally be any appropriate diffusion sampler that maps the denoiser output to the updated representation. For example, the system can use an appropriate first or second order solver. Examples of first order solvers are described in Ho, et al, Denoising Diffusion Probabilistic Models, available at arXiv:2006.11239 and Song, et al, Denoising Diffusion Implicit Models, available at arXiv 2010.02502. Examples of second order solvers are described in T. Karras, M. Aittala, T. Aila, and S. Laine. Elucidating the design space of diffusion-based generative models. Advances in Neural Information Processing Systems, 35:26565-26577, 2022 and C. Lu, Y. Zhou, F. Bao, J. Chen, C. Li, and J. Zhu. DPM-Solver++: Fast solver for guided sampling of diffusion probabilistic models. arXiv preprint arXiv:2211.01095, 2022. In some cases, the system can augment a deterministic first or second order solver with stochasticity in order to add stochasticity to the denoising process and improve the diversity of the ensemble. Examples of augmenting a solver to add additional stochasticity are described in T. Karras, M. Aittala, T. Aila, and S. Laine. Elucidating the design space of diffusion-based generative models. Advances in Neural Information Processing Systems, 35:26565-26577, 2022.

[0095] Prior to using the diffusion neural network to generate forecasts, the system or another training system trains the diffusion neural network on training data.

[0096] For example, the training data can include multiple sequences of target weather data, i.e., multiple sequences that each include target weather data that has been observed or measured at respective time points. Such training data can be obtained from any of a variety of data sources that compile weather that has been historically observed or measured at different points around the surface (for example, the ERA5 archive of the European Centre for Medium-Range Weather Forecasts, ECMWF).

[0097] The diffusion neural network can be trained using any of a variety of objectives that measure how accurately the diffusion neural network can predict weather properties or, more specifically, how accurately the diffusion neural network can generate denoising outputs.

[0098] As one example, the training system can train the diffusion neural network to minimize an objective that measures how well the outputs of the denoiser that are generated using a denoising output generated by the diffusion neural network by processing a noisy version of a ground truth target representation match the noise-free, ground truth target representation.

[0099] As one example, the objective can be: where:

[0100] - 1 indexes the different timesteps in the training set Dtrain,

[0101] - j E J indexes the variable, and for atmospheric variables the pressure level. E.g. J ={zl000, z850, . . . , 2T, MSL},

[0102] - i E G indexes the location (latitude and longitude coordinates) in the grid,

[0103] - Wj is the per-variable-level loss weight, - atis the area of the latitude-longitude grid cell, which varies with latitude, and optionally can be normalized to unit mean over the grid,

[0104] - (cr) is a per-noise-level loss weight,

[0105] - the expectation IE is taken over (5 ~ ptrain, i.e., from a training distribution over different noise levels, e ~ pnoise (•; <J), i.e., from a noise level-dependent noise distribution.

[0106] FIG. 4 shows an example 400 of the operation of the forecasting system 100 when generating a forecast.

[0107] As shown in FIG. 4, the system 100 receives actual weather data A 402 at time point t (and optionally one or more time points that are earlier than f). The system uses the diffusion neural network 110 to generate predicted weather data 404 at time point Z+l. The system then uses the predicted weather 404 (and optionally the actual weather data 402 at time point f) as input to the diffusion neural network 110 to generate predicted weather data 404 at time point t+2. The system can continue iteratively feeding predicted weather states as input to the diffusion neural network 110 across multiple different time points until generating predicted weather data 406 at time point T.

[0108] At each future time point, the system initializes a representation Z and then uses the diffusion neural network 110 to perform N sampling iterations re conditioned on at least the state from the preceding time point. Each sampling iteration updates the representation Z until arriving at the final representation ZNfor the time point. The system then uses the final representation ZNand, when the values in ZNrepresent delta values, the preceding weather state X to generate the predicted weather state X at the future time step.

[0109] The diffusion neural network 110 that is used by the system can generally have any appropriate architecture that allows the diffusion neural network to map from the input weather state and the current representation to generate a denoising output that has the same dimensionality as the current representation.

[0110] One example of the architecture of the diffusion neural network is described below with reference to FIG. 5.

[0111] FIG. 5 shows an example 500 of the configuration of the diffusion neural network

[0112] 110. As shown in the example 500, the diffusion neural network 110 includes an encoder neural network 510, a processor neural network 520, a decoder neural network 530, and an output neural network 540.

[0113] The encoder neural network 510 is configured to process a graph representation 502 of the input weather state.

[0114] In particular, the graph representation represents a graph of the surface that includes a plurality of nodes and a plurality of edges.

[0115] The nodes in the graph include grid nodes that each correspond to one of the points on the latitude-longitude grid and mesh nodes that each correspond to a node in a mesh that is placed around the surface. For example, the mesh can be an icosahedral mesh placed around the surface.

[0116] The edges also include a plurality of mesh edges that each connect a respective pair of mesh nodes in the mesh.

[0117] In some cases, the edges also include grid-to-mesh edges, where each of the grid- to-mesh edges is a unidirectional edge from a respective grid node to a respective mesh node.

[0118] Further optionally, the edges can also include mesh-to-grid edges, where each of the plurality of mesh-to-grid edges is a unidirectional edge from a respective mesh node to a respective grid node.

[0119] Generally, the graph representation 502 includes a respective initial embedding for each node and each edge in the graph, with the respective initial embeddings for the grid nodes being generated from the input weather state.

[0120] For example, the system can generate respective features for each of the nodes and then process the features using an embedding neural network to generate the initial embeddings.

[0121] For example, the system can generate the respective features for each of the grid nodes at least in part from the current weather properties at the point corresponding to the grid node and generate the respective features for each of the mesh nodes based at least in part on the longitude and latitude of the mesh node. In addition to the current weather properties at the point, the features for a given grid node can optionally include analytically computed features at the point and static features at the point.

[0122] As a particular example, the respective features for each of the grid nodes can include respective normalized values for each of one or more of the current weather properties at the point. For example, the values of the grid nodes (i.e., the distribution of values over the grid nodes) can be normalized to have zero mean and unit variance. For example, for each physical variable, the system can compute the per-pressure level mean and standard deviation over a historical period, e.g., the period covered by the training data for the neural network or a different historical period, and use the per-pressure level means and standard deviations to normalize the corresponding values to zero mean and unit variance.

[0123] As another example, the respective features for each of the edges can be based on respective positions of each of the two nodes (e.g., locations on the surface of the body) connected by the edge. For example, for mesh edges, mesh-to-grid edges, and grid-to- mesh edges the features can include the length of the edge and the vector difference between the 3d positions of a sender node for the edge and a receiver node for the edge computed in a local coordinate system of the receiver. In some cases, the system can normalize the edge lengths, e.g., based on the length of the longest edge in the graph.

[0124] For example, each type of node and edge can have a different corresponding embedding neural network and the embedding neural networks can have been trained jointly with the diffusion neural network. As discussed, the types of node may include mesh nodes and grid nodes, and the types of edge may comprise mesh edges and optionally, grid-to-mesh edges and / or mesh-to-grid edges. For example, the mesh nodes and grid nodes may have different corresponding embedding neural networks, and the mesh edges, grid-to-mesh edges and mesh-to-grid edges may have different corresponding embedding neural networks. The embedding neural network(s) can generally have any appropriate neural network architecture. For example, the neural network(s) can be respective multi-layer perceptrons (MLPs). In some cases, as will be described below, the graph representation does not include features for mesh edges. In some cases, the graph representation does not include features for mesh edges and, accordingly, the diffusion neural network does not generate initial embeddings for the mesh edges.

[0125] As described above, the first diffusion input is also conditioned on the representation of the weather state for the future time step as of the sampling iteration. Accordingly, the graph representation 502 can also be conditioned on the representation of the weather state. The system can implement this conditioning in any of a variety of ways. As one example, the system can, prior to generating the initial embeddings of the grid nodes and for each grid node, concatenate the values in the representation of the weather state for the grid node to the features of the grid node. The encoder neural network 510 processes the graph representation 502 to generate a respective embedding 512 for each mesh node. For example, the encoder neural network 510 can be a graph neural network. As a particular example, the encoder neural network 510 can process the respective initial embeddings of each of the nodes and edges in a bipartite subgraph of the graph that includes the grid nodes, the mesh nodes, and the grid-to-mesh edges using a grid-to-mesh graph neural network to update the respective embeddings of at least the mesh nodes.

[0126] Graph neural networks are described in Battaglia et al. arXiv: 1806.01261, for example.

[0127] A bipartite subgraph includes nodes that can be divided into two disjoint (nonoverlapping) sets, in this case the grid nodes and the mesh nodes, with every edge (in this case, grid-to-mesh edge) connecting respective nodes that belong to different ones of the disjoint sets.

[0128] The processor neural network 520 is configured to process at least the respective embeddings 512 for each of the mesh nodes in the mesh to generate a respective updated embedding 522 for each of the mesh nodes in the mesh.

[0129] The processor neural network 520 can generally have any appropriate architecture that can update the embeddings 512.

[0130] As one example, the processor neural network can be a Transformer neural network.

[0131] In some examples, to improve the computational efficiency of the generation of the forecast, the processor neural network can be a sparse Transformer neural network. In particular, the Transformer neural network is referred to as “sparse” because, when applying self-attention, any given mesh node attends to only a proper subset of the other mesh nodes (rather than to all other mesh nodes, as in a dense attention mechanism). In particular, any given mesh node can attend to only the nodes that are within a threshold number of hops on the mesh of the given mesh node. The threshold number can be, e.g., 4, 8, 16, or 32 hops.

[0132] In this example, i.e., when the processor neural network 520 is a Transformer, e.g., a dense or a sparse Transformer, the mesh “edges” are data that identifies which mesh nodes should attend to each other when applying self-attention. That is, two mesh nodes attend to one another when they are connected by a mesh edge. In other word, when the processor neural network 520 is a Transformer, the neural network does not need to obtain initial features of mesh edges or generate initial embeddings of mesh edges.

[0133] The decoder neural network 530 is configured to process the respective updated embedding 522 for each of the mesh nodes in the mesh to generate a respective embedding 532 for each of the points in the grid.

[0134] For example, the decoder neural network 530 can be a graph neural network.

[0135] As a particular example, the decoder neural network 530 can be configured to process respective input embeddings of each of the nodes and edges in a bipartite subgraph of the graph that includes the grid nodes, the mesh nodes, and the mesh-to-grid edges using a mesh-to-grid graph neural network to update the respective input embeddings of each of the grid nodes.

[0136] The output neural network 540 is configured to process the respective embeddings for each of the points in the grid to generate the denoising output. For example, the output neural network 520 can be, e.g., a multi-layer perceptron (MLP) or other feedforward neural network that, for each grid node, processes the embedding of the grid node to generate the values corresponding to the grid node in the denoising output.

[0137] As described above, the input to the diffusion neural network can also include the noise level for the sampling iteration, optionally after applying a pre-conditioning function to the noise level.

[0138] The diffusion neural network 110 can be conditioned on the noise level in any of a variety of ways.

[0139] As a particular example, the diffusion neural network 110 can include, e.g., as part of the processor neural network, the encoder neural network, the decoder neural network, or two or more of the above, one or more conditional layer normalization layers that are conditioned on the noise level. As a particular example, the system can transform log noise levels into a vector of sine / cosine Fourier features, e.g., at 32 frequencies with base period 16, and then pass the vector of features through an encoding neural network, e.g., an MLP, to obtain noise-level encodings. For example, each of the conditional layernorm layers can apply a further linear layer to output replacements for the standard scale and offset parameters of layer norm, conditioned on these noise-level encodings. Conditional layer normalization layers are described in more detail in M. Chen, X. Tan, B. Li, Y. Liu, T. Qin, S. Zhao, and T.-Y. Liu. Adaspeech: Adaptive text to speech for custom voice, 2021 This specification uses the term “configured” in connection with systems and computer program components. For a system of one or more computers to be configured to perform particular operations or actions means that the system has installed on it software, firmware, hardware, or a combination of them that in operation cause the system to perform the operations or actions. For one or more computer programs to be configured to perform particular operations or actions means that the one or more programs include instructions that, when executed by data processing apparatus, cause the apparatus to perform the operations or actions.

[0140] Embodiments of the subject matter and the functional operations described in this specification can be implemented in digital electronic circuitry, in tangibly-embodied computer software or firmware, in computer hardware, including the structures disclosed in this specification and their structural equivalents, or in combinations of one or more of them. Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible non transitory storage medium for execution by, or to control the operation of, data processing apparatus. The computer storage medium can be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or a combination of one or more of them. Alternatively or in addition, the program instructions can be encoded on an artificially generated propagated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal, that is generated to encode information for transmission to suitable receiver apparatus for execution by a data processing apparatus.

[0141] The term “data processing apparatus” refers to data processing hardware and encompasses all kinds of apparatus, devices, and machines for processing data, including by way of example a programmable processor, a computer, or multiple processors or computers. The apparatus can also be, or further include, special purpose logic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application specific integrated circuit). The apparatus can optionally include, in addition to hardware, code that creates an execution environment for computer programs, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them.

[0142] A computer program, which may also be referred to or described as a program, software, a software application, an app, a module, a software module, a script, or code, can be written in any form of programming language, including compiled or interpreted languages, or declarative or procedural languages; and it can be deployed in any form, including as a stand alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A program may, but need not, correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data, e.g., one or more scripts stored in a markup language document, in a single file dedicated to the program in question, or in multiple coordinated files, e.g., files that store one or more modules, sub programs, or portions of code. A computer program can be deployed to be executed on one computer or on multiple computers that are located at one site or distributed across multiple sites and interconnected by a data communication network.

[0143] In this specification, the term “database” is used broadly to refer to any collection of data: the data does not need to be structured in any particular way, or structured at all, and it can be stored on storage devices in one or more locations. Thus, for example, the index database can include multiple collections of data, each of which may be organized and accessed differently.

[0144] Similarly, in this specification the term “engine” is used broadly to refer to a software-based system, subsystem, or process that is programmed to perform one or more specific functions. Generally, an engine will be implemented as one or more software modules or components, installed on one or more computers in one or more locations. In some cases, one or more computers will be dedicated to a particular engine; in other cases, multiple engines can be installed and running on the same computer or computers.

[0145] The processes and logic flows described in this specification can be performed by one or more programmable computers executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows can also be performed by special purpose logic circuitry, e.g., an FPGA or an ASIC, or by a combination of special purpose logic circuitry and one or more programmed computers.

[0146] Computers suitable for the execution of a computer program can be based on general or special purpose microprocessors or both, or any other kind of central processing unit. Generally, a central processing unit will receive instructions and data from a read only memory or a random access memory or both. The essential elements of a computer are a central processing unit for performing or executing instructions and one or more memory devices for storing instructions and data. The central processing unit and the memory can be supplemented by, or incorporated in, special purpose logic circuitry. Generally, a computer will also include, or be operatively coupled to receive data from or transfer data to, or both, one or more mass storage devices for storing data, e.g., magnetic, magneto optical disks, or optical disks. However, a computer need not have such devices. Moreover, a computer can be embedded in another device, e.g., a mobile telephone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a Global Positioning System (GPS) receiver, or a portable storage device, e.g., a universal serial bus (USB) flash drive, to name just a few.

[0147] Computer readable media suitable for storing computer program instructions and data include all forms of non volatile memory, media and memory devices, including by way of example semiconductor memory devices, e.g., EPROM, EEPROM, and flash memory devices; magnetic disks, e.g., internal hard disks or removable disks; magneto optical disks; and CD ROM and DVD-ROM disks.

[0148] To provide for interaction with a user, embodiments of the subject matter described in this specification can be implemented on a computer having a display device, e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, for displaying information to the user and a keyboard and a pointing device, e.g., a mouse or a trackball, by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback, e.g., visual feedback, auditory feedback, or tactile feedback; and input from the user can be received in any form, including acoustic, speech, or tactile input. In addition, a computer can interact with a user by sending documents to and receiving documents from a device that is used by the user; for example, by sending web pages to a web browser on a user’s device in response to requests received from the web browser. Also, a computer can interact with a user by sending text messages or other forms of message to a personal device, e.g., a smartphone that is running a messaging application, and receiving responsive messages from the user in return.

[0149] Data processing apparatus for implementing machine learning models can also include, for example, special-purpose hardware accelerator units for processing common and compute-intensive parts of machine learning training or production, i.e., inference, workloads.

[0150] Machine learning models can be implemented and deployed using a machine learning framework, e.g., a TensorFlow framework. Embodiments of the subject matter described in this specification can be implemented in a computing system that includes a back end component, e.g., as a data server, or that includes a middleware component, e.g., an application server, or that includes a front end component, e.g., a client computer having a graphical user interface, a web browser, or an app through which a user can interact with an implementation of the subject matter described in this specification, or any combination of one or more such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication, e.g., a communication network. Examples of communication networks include a local area network (LAN) and a wide area network (WAN), e.g., the Internet.

[0151] The computing system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. In some embodiments, a server transmits data, e.g., an HTML page, to a user device, e.g., for purposes of displaying data to and receiving user input from a user interacting with the device, which acts as a client. Data generated at the user device, e.g., a result of the user interaction, can be received at the server from the device.

[0152] While this specification contains many specific implementation details, these should not be construed as limitations on the scope of any invention or on the scope of what may be claimed, but rather as descriptions of features that may be specific to particular embodiments of particular inventions. Certain features that are described in this specification in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment can also be implemented in multiple embodiments separately or in any suitable subcombination. Moreover, although features may be described above as acting in certain combinations and even initially be claimed as such, one or more features from a claimed combination can in some cases be excised from the combination, and the claimed combination may be directed to a subcombination or variation of a subcombination.

[0153] Similarly, while operations are depicted in the drawings and recited in the claims in a particular order, this should not be understood as requiring that such operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed, to achieve desirable results. In certain circumstances, multitasking and parallel processing may be advantageous. Moreover, the separation of various system modules and components in the embodiments described above should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.

[0154] Particular embodiments of the subject matter have been described. Other embodiments are within the scope of the following claims. For example, the actions recited in the claims can be performed in a different order and still achieve desirable results. As one example, the processes depicted in the accompanying figures do not necessarily require the particular order shown, or sequential order, to achieve desirable results. In some cases, multitasking and parallel processing may be advantageous.

Claims

CLAIMS1. A method performed by one or more computers, the method comprising: obtaining data characterizing a current weather state for a current time step, wherein the current weather state comprises, for each of a plurality of points on a latitudelongitude grid over a surface of a body, current weather properties at the point as of the current time step; and generating a first weather forecast that comprises a respective predicted weather state for each future time step in a sequence of time steps that starts at the current time step and comprises a plurality of future time steps that follow the current time step, comprising, for each future time step: initializing a representation of a weather state for the future time step; updating the representation of the weather state for the future time step to generate a final representation of the weather state for the future time step, the updating comprising, at each of a plurality of sampling iterations: generating a denoising output for the sampling iteration, comprising processing a first diffusion input for the sampling iteration that comprises (i) the representation of the weather state for the future time step and (ii) a representation of an input weather state for the future time step using a diffusion neural network to generate a first denoising output, wherein: for the first future time step in the sequence, the input weather state comprises a representation of the current weather state for the current time step, and for each subsequent future time step in the sequence, the input weather state comprises the respective predicted weather state for a preceding future time step in the sequence; and updating the representation of the future weather state using the denoising output; and generating the respective predicted weather state for the future time step from the final representation of the respective predicted weather state.

2. The method of claim 1, wherein: for the first future time step in the sequence, the input weather state comprises a preceding weather state for a preceding time step that precedes the current time step; for a second future time step in the sequence, the input weather state comprises the respective predicted weather state for the first future time step in the sequence and the current weather state; and for each future time step that is after the second future time step in the sequence, the input weather state comprises the respective predicted weather states for two or more preceding future time steps in the sequence.

3. The method of any preceding claim, wherein the final representation defines, for each point and for each of one or more first weather properties, a predicted delta of the first weather property between the future time step and the preceding time step in the sequence.

4. The method of claim 3, wherein the final representation specifies, for each point and for each of the one or more first weather properties, a predicted delta between a normalized value of the first weather property at the future time step and a normalized value of the first weather property at the preceding time step in the sequence.

5. The method of any preceding claim, wherein the final representation defines, for each point and for each of one or more second weather properties, a predicted value of the second weather property at the future time step.

6. The method of claim 5, wherein the final representation specifies, for each point and for each of the one or more second weather properties, a normalized value of the second weather property at the future time step.

7. The method of any preceding claim, further comprising: generating one or more additional weather forecasts that each comprise a respective predicted weather state for each future time step in the sequence of time steps, comprising, for each additional weather forecast and for each future time step: initializing a representation of a weather state for the future time step; updating the representation of the weather state to generate a final representation of the weather state, the updating comprising, at each of a plurality of sampling iterations: generating a denoising output for the sampling iteration, comprising processing a first diffusion input for the sampling iteration that comprises (i) the representation of the weather state for the future time step and (ii) a representation of an input weather state for the future time step using a diffusion neural network to generate a first denoising output, wherein: for the first future time step in the sequence, the input weather state comprises a representation of the current weather state for the current time step, and for each subsequent future time step in the sequence, the input weather state comprises the respective predicted weather state for a preceding future time step in the sequence; and updating the representation of the future weather state using the denoising output; and generating the respective predicted weather state for the future time step in the additional weather forecast from the final representation of the respective predicted weather state.

8. The method of claim 7, wherein generating the weather forecast and the one or more additional weather forecasts comprises generating the weather forecast and the one or more additional weather forecasts in parallel using a respective instance of the diffusion neural network deployed on a corresponding set of one or more hardware devices.

9. The method of claim 7 or claim 8, wherein the representations of the current weather state at the current time step are different between the weather forecast and one or more of the additional weather forecasts.

10. The method of any preceding claim, wherein initializing a representation of a weather state for the future time step comprises sampling noisy values from a noise probability distribution.

11. The method of claim 10, wherein initializing a representation of a weather state for the future time step comprises: sampling Gaussian noise on a sphere corresponding to the surface; and projecting the sampled Gaussian noise onto the grid.

12. The method of any preceding claim, wherein the diffusion neural network comprises: an encoder neural network configured to process a graph representation of the input weather state, wherein: the graph representation represents a graph of the surface, the graph comprises a plurality of nodes and a plurality of edges, the plurality of nodes comprises a plurality of grid nodes that each correspond to one of the points on the latitude-longitude grid and a plurality of mesh nodes that each correspond to a node in a mesh that is placed around the surface, the plurality of edges comprises a plurality of mesh edges that each connect a respective pair of mesh nodes in the mesh, and the encoder neural network processes the graph representation to generate a respective embedding for each mesh node; a processor neural network configured to process at least the respective embeddings for each of the mesh nodes in the mesh to generate a respective updated embedding for each of the mesh nodes in the mesh; a decoder neural network configured to process the respective updated embedding for each of the mesh nodes in the mesh to generate a respective embedding for each of the points in the grid; and an output neural network configured to process the respective embeddings for each of the points in the grid to generate the first denoising output.

13. The method of claim 12, wherein the encoder neural network is a graph neural network.

14. The method of claim 12 or 13, wherein the decoder neural network is a graph neural network.

15. The method of any one of claims 12-14, wherein the processor neural network is a Transformer neural network.

16. The method of claim 15, wherein the processor neural network is a sparse Transformer neural network.

17. The method of any one of claims 12-16, wherein the graph representation includes a respective initial embedding for each node and each edge in the graph, and wherein the respective initial embeddings for the grid nodes are generated from the input weather state.

18. The method of any one of claims 12-17, wherein the plurality of edges include a plurality of grid-to-mesh edges, wherein each of the plurality of grid-to-mesh edges is a unidirectional edge from a respective grid node to a respective mesh node.

19. The method of claim 18, when dependent on claim 17, wherein the encoder neural network is configured to process the respective initial embeddings of each of the nodes and edges in a bipartite subgraph of the graph that includes the grid nodes, the mesh nodes, and the grid-to-mesh edges using a grid-to-mesh graph neural network to update the respective embeddings of at least the mesh nodes.

20. The method of any one of claims 12-19, wherein the plurality of edges includes a plurality of mesh-to-grid edges, wherein each of the plurality of mesh-to-grid edges is a unidirectional edge from a respective mesh node to a respective grid node.

21. The method of claim 20, wherein the decoder neural network is configured to process respective input embeddings of each of the nodes and edges in a bipartite subgraph of the graph that includes the grid nodes, the mesh nodes, and the mesh-to-grid edges using a mesh-to-grid graph neural network to update the respective input embeddings of each of the grid nodes.

22. The method of any preceding claim, wherein a time interval between each time step is between three hours and twenty four hours.

23. The method of claim 22, wherein the time interval is twelve hours.

24. The method according to any one of the preceding claims, further comprising generating and providing an alert whenever a given predicted weather property of the first weather forecast satisfies a specified threshold.

25. The method of any one of claims 1-23 when dependent on claim 7, further comprising: generating an ensemble weather forecast that comprises the first weather forecast and the one or more additional weather forecasts; and generating and providing an alert (i) whenever a given predicted weather property satisfies a specified threshold in one or more of the forecasts in the ensemble or (ii) whenever the average, maximum, or minimum value of a given predicted property satisfies a specified threshold.

26. The method according to any one of claims 1-23, wherein the predicted weather states characterize the predicted weather at a real-world location of a renewable energy generation facility, the method further comprising:(i) controlling the renewable energy generation facility in response to the predicted weather at the real-world location; or(ii) sending, to a consumer of power on an electrical grid to which the renewable energy generation facility is connected, a signal in response to a predicted power output from the renewable energy generation facility based on the predicted weather at the real- world location.

27. The method according to any one of claims 1-23, wherein the predicted weather states characterize the predicted weather at a real-world location, the method further comprising outputting a warning signal or a routing signal for controlling the flight patterns or marine craft routes of one or more aircraft or marine craft in response to the predicted weather at the real -world location.

28. The method according to any one of claims 1-23, wherein the predicted weather states characterize the predicted weather at a real-world location of a building or industrial facility, the method further comprising controlling shutters, or a ventilation system, or a temperature control system, of the building or industrial facility in response to the predicted weather at the real-world location.

29. A system comprising: one or more computers; and one or more storage devices communicatively coupled to the one or more computers, wherein the one or more storage devices store instructions that, when executed by the one or more computers, cause the one or more computers to perform operations of the respective method of any one of claims 1-28.

30. One or more non-transitory computer storage media storing instructions that when executed by one or more computers cause the one or more computers to perform operations of the respective method of any one of claims 1-28.

Citation Information

Cited By

  • Landslide surface deformation spatial-temporal feature interpolation method, storage medium and equipment

    CN122244719A