A prediction method based on inverse reinforcement learning
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Pingdu Meteorological Bureau
- Filing Date
- 2026-03-17
- Publication Date
- 2026-07-03
Smart Images

Figure CN122334390A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of Earth science technology, specifically to a prediction method based on inverse reinforcement learning. Background Technology
[0002] Large-scale weather forecasting models such as FourCastNet, Pangu-Weather, FuXi, and GraphCast, as well as large-scale ocean forecasting models such as XiHe and AI-GOMS, have significantly improved the quality of weather and ocean forecasts by leveraging the powerful modeling capabilities of large-scale neural networks. However, the predictive power of large models is not only related to their architecture but also highly dependent on the training method. Existing large models are typically trained by manually selecting a loss function, but this limits the model architecture from realizing its true potential.
[0003] In the field of Earth sciences, reanalysis datasets obtained through data assimilation are the best estimates of the state of the atmosphere or ocean. These datasets constitute long-term, high-dimensional, continuous, and physically consistent spatiotemporal evolution trajectories, and their latent space representations can be directly used as approximations of expert policies. Therefore, the field of Earth sciences such as atmosphere and ocean naturally possesses the exemplary data required for inverse reinforcement learning. However, traditional prediction methods have failed to make full use of this characteristic of reanalysis datasets, and their prediction results are often physically inconsistent. Summary of the Invention
[0004] This application provides a prediction method based on inverse reinforcement learning, which can make full use of the features of the reanalysis dataset and improve the prediction accuracy of the prediction model.
[0005] The technical solution of this application is as follows: A prediction method based on inverse reinforcement learning includes the following steps: S1. Obtain the original spatial sample set of the long-term series, wherein the original spatial sample set contains several physical field data at several time steps; S2. Based on the pre-trained encoder, each original space sample is mapped to the latent space to obtain the corresponding latent space samples and form a latent space sample set for a long time series. S3. Extract several consecutive sub-time series latent space sample sets based on the latent space sample set, wherein the sub-time series latent space sample set includes several consecutive latent space samples. S4. Based on several continuous sub-time series latent space sample sets, an inverse reinforcement learning algorithm is used to train the prediction model, resulting in a trained prediction model, wherein: Each sub-time series latent space sample set is treated as an expert trajectory, and the first expert trajectory is... i The first to the second i +n The nth latent space sample is used as the first... ik The first to the secondi- The action of an expert policy with a state represented by a single latent space sample; The prediction model is used as a policy network to construct a discriminator to distinguish between expert trajectories and prediction trajectories generated by the prediction model; the discriminator is constructed to output implicit reward signals to optimize the policy network so that the prediction trajectory approximates the expert trajectory in statistical properties.
[0006] Furthermore, the original spatial sample set is derived from the reanalysis dataset, including at least one of ERA5, CRA40, and GLORYS12.
[0007] Furthermore, the physical field data includes upper-level variables and surface variables from several pressure layers; The upper-level variables include geopotential height, temperature, zonal wind speed, radial wind speed, and relative humidity; The surface variables include 2m air temperature, 10m zonal wind speed, 10m meridional wind speed, sea level air pressure, and total precipitation.
[0008] Furthermore, in step S4, the prediction model trained using the inverse reinforcement learning algorithm is a pre-trained model, and the pre-training method is as follows: Using latent space samples, the prediction model is pre-trained with mean squared error, mean absolute error, or weighted mean absolute error loss function as the loss function. The parameters of the pre-trained model are then used to initialize the policy network.
[0009] Furthermore, the discriminator is optimized based on least squares loss or Wasserstein distance.
[0010] Furthermore, after the prediction model is trained, predictions are made based on the pre-trained decoder, including the following steps: The original spatial physical field is mapped to the latent space initial field by the pre-trained encoder and input into the trained prediction model to obtain the latent space prediction result. The latent space prediction result is then restored to the original spatial prediction result by the pre-trained decoder.
[0011] Furthermore, the pre-trained encoder and pre-trained decoder are obtained through joint training via an autoencoder structure, and the loss function of the autoencoder includes a reconstruction error term.
[0012] Furthermore, the prediction model employs Transformer, Mamba, or convolutional neural network architectures.
[0013] Due to the adoption of the above technical solution, the beneficial effects of this application are as follows: 1. Existing model training methods rely on point-to-point error metrics, which cannot explicitly or implicitly constrain conservation laws such as mass, momentum, and energy, resulting in predictions that, while numerically close, are physically inconsistent. This application utilizes inverse reinforcement learning to learn the overall distribution characteristics of expert trajectories, enabling the prediction model to automatically internalize complex physical constraints during training, significantly improving the physical consistency of prediction results.
[0014] 2. Existing methods use fixed loss functions such as MSE, which only measure the numerical deviation between a single-step prediction and observation, failing to reflect the rationality of multi-step dynamics. This application utilizes a discriminator to distinguish between expert trajectories and predicted trajectories, with its output serving as an implicit reward signal. This is equivalent to defining a loss function automatically learned from the data: this loss is no longer based on point-to-point error, but rather on the distribution matching degree of the entire expert trajectory. Therefore, the model optimization direction shifts from a single-point perspective to a global line perspective, thus intrinsically satisfying physical consistency.
[0015] 3. This application employs a generative adversarial imitation learning mechanism, coupling reward learning and policy optimization into an adversarial training process: the discriminator dynamically evaluates the authenticity of the trajectory, and the policy network directly updates its parameters accordingly. This approach eliminates the need for explicit recovery of the reward function, significantly reducing computational overhead.
[0016] 4. This application fully utilizes the characteristics of meteorological datasets to facilitate the implementation of the method. Reanalysis datasets such as ERA5 provide long-term, continuous, and physically coordinated real evolutionary sequences, which can be directly used as expert trajectories. Moreover, data from atmospheric and other Earth science fields have strong spatial correlations and can be mapped to a low-dimensional latent space through a pre-trained encoder, thereby achieving inverse reinforcement training of high-dimensional continuous fields while preserving the main dynamic characteristics. Attached Figure Description
[0017] The accompanying drawings, which are provided to further illustrate this application and form part of this application, illustrate exemplary embodiments of this application and are used to explain this application, but do not constitute an undue limitation of this application.
[0018] Figure 1 A flowchart of a prediction method based on inverse reinforcement learning provided in this application. Detailed Implementation
[0019] Based on the background technology described, as shown in the appendix Figure 1 As shown, this application provides a prediction method based on inverse reinforcement learning, including the following steps: S1. Obtain the original spatial sample set of the long-term series, wherein the original spatial sample set contains several physical field data at several time steps.
[0020] The original spatial sample set is derived from the reanalysis dataset, including at least one of ERA5, CRA40, and GLORYS12.
[0021] The physical field data includes upper-air variables from several pressure layers and surface variables. The upper-level variables include geopotential height, temperature, zonal wind speed, radial wind speed, and relative humidity; The surface variables include 2m air temperature, 10m zonal wind speed, 10m meridional wind speed, sea level air pressure, and total precipitation.
[0022] This embodiment obtains long-term physical field data from the ERA5 reanalysis dataset released by the European Centre for Medium-Range Weather Forecasts (ECMWF). The time range is from January 1, 1980 to December 31, 2015, with a temporal resolution of 6 hours. The following variables are selected to constitute the original spatial sample: Upper-altitude variables: Geopotential height, temperature, zonal wind speed, meridional wind speed, and relative humidity were collected at 13 standard pressure levels (1000hPa, 925hPa, 850hPa, 700hPa, 600hPa, 500hPa, 400hPa, 300hPa, 250hPa, 200hPa, 150hPa, 100hPa, 50hPa), for a total of 13×5=65 channels; Surface variable: 2-meter air temperature (T) 2M The system has five channels: 10-meter zonal wind speed (10mU), 10-meter meridional wind speed (10mV), sea level pressure (MSL), and total precipitation (TP).
[0023] After standardization, all variables form a 70-channel three-dimensional tensor sequence with a spatial resolution of 0.25°×0.25°, which serves as the original spatial sample set.
[0024] S2. Based on the pre-trained encoder, each original space sample is mapped to the latent space to obtain the corresponding latent space samples and form a latent space sample set for a long time sequence.
[0025] In this embodiment, a convolutional autoencoder is trained on the original spatial sample set, using reconstruction error as the loss function. After training, the encoder and decoder parameters of the convolutional autoencoder are fixed and used as the pre-trained encoder and decoder for this step and subsequent steps.
[0026] S3. Extract several continuous sub-time series latent space sample sets based on the latent space sample set, wherein the sub-time series latent space sample set includes several continuous latent space samples.
[0027] Several consecutive sub-time series latent space sample sets are represented as Each sub-time series latent space sample set includes several consecutive latent space samples, denoted as... The number of latent space samples in different sub-time series latent space sample sets can be different.
[0028] S4. Based on several continuous sub-time series latent space sample sets, an inverse reinforcement learning algorithm is used to train the prediction model, resulting in a trained prediction model, wherein: Each sub-time series latent space sample set is treated as an expert trajectory, and the first expert trajectory is... i The first to the second i +n The nth latent space sample is used as the first... ik The first to the second i- The action of an expert policy with a state represented by a single latent space sample; The prediction model is used as a policy network to construct a discriminator to distinguish between expert trajectories and prediction trajectories generated by the prediction model; the discriminator is constructed to output implicit reward signals to optimize the policy network so that the prediction trajectory approximates the expert trajectory in statistical properties.
[0029] The prediction model trained using the inverse reinforcement learning algorithm is the pre-trained model. The training method is as follows: Using latent space samples, the prediction model is pre-trained with mean squared error, mean absolute error, or weighted mean absolute error loss function as the loss function, and the obtained parameters are used to initialize the policy network.
[0030] In practice, each consecutive sub-time series latent space sample set As an expert trajectory, each expert trajectory will contain ( ) as a ( Given the expert policy (state, action) as the state, the prediction model is used as the policy network. The parameters of the policy network and discriminator are initialized, and a generative adversarial imitation learning algorithm is used to train the policy network and discriminator. The trained policy network is then used as the trained prediction model. During training, the discriminator uses the expert policy (state, action) to... Or policy network (state, action) pair For input, where, ( ) as a policy network with state ( The input represents the action. The discriminator employs a convolutional neural network structure and is optimized based on least squares loss or Wasserstein distance.
[0031] After the prediction model is trained, weather forecasts are made based on the pre-trained decoder, including the following steps: The original spatial physical field is mapped to the latent space initial field by the pre-trained encoder and input into the trained prediction model to obtain the latent space prediction result. The latent space prediction result is then restored to the original spatial prediction result by the pre-trained decoder.
[0032] Specifically, obtain the original spatial physical field at the initial moment; It is mapped to the latent space initial field by a pre-trained encoder; The initial field of the latent space is input into the trained prediction model to recursively generate the future. N The latent space prediction sequence of the step; Each latent space prediction sequence element is restored to the original spatial physical field through a pre-trained decoder, such as key variables like 500hPa geopotential height, 850hPa temperature, and surface precipitation.
[0033] Optionally, the prediction model employs a Transformer, Mamba, or convolutional neural network architecture, including but not limited to weather prediction models and ocean prediction models.
[0034] For any parts not mentioned in this application, existing technologies may be used or referenced.
[0035] The above description is merely an embodiment of this application and is not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.
Claims
1. A prediction method based on inverse reinforcement learning, characterized in that, Includes the following steps: S1. Obtain the original spatial sample set of the long-term series, wherein the original spatial sample set contains several physical field data at several time steps; S2. Based on the pre-trained encoder, each original space sample is mapped to the latent space to obtain the corresponding latent space samples and form a latent space sample set for a long time series. S3. Extract several continuous sub-time series latent space sample sets based on the latent space sample set, wherein the sub-time series latent space sample sets include several continuous latent space samples. S4. Based on several continuous sub-time series latent space sample sets, an inverse reinforcement learning algorithm is used to train the prediction model, resulting in a trained prediction model, wherein: each sub-time sequence latent space sample set as an expert trajectory, taking the first i to the i+n latent space samples in the expert trajectory as the actions of the expert policy with the first ik to the i- 1latent space sample as the state; The prediction model is used as a policy network to construct a discriminator to distinguish between expert trajectories and prediction trajectories generated by the prediction model; the discriminator is constructed to output implicit reward signals to optimize the policy network so that its prediction trajectory is statistically similar to that of the expert trajectory.
2. The prediction method based on inverse reinforcement learning according to claim 1, characterized in that, The original spatial sample set is derived from the reanalysis dataset, including at least one of ERA5, CRA40, and GLORYS12.
3. The prediction method based on inverse reinforcement learning according to claim 2, characterized in that, The physical field data includes upper-air variables from several pressure layers and surface variables. The upper-level variables include geopotential height, temperature, zonal wind speed, radial wind speed, and relative humidity; The surface variables include 2m air temperature, 10m zonal wind speed, 10m meridional wind speed, sea level air pressure, and total precipitation.
4. The prediction method based on inverse reinforcement learning according to claim 3, characterized in that, In step S4, the prediction model trained using the inverse reinforcement learning algorithm is the pre-trained model. The pre-training method is as follows: Using latent space samples, the prediction model is pre-trained with mean squared error, mean absolute error, or weighted mean absolute error as the loss function to obtain a pre-trained model. The parameters of the pre-trained model are then used to initialize the policy network.
5. The prediction method based on inverse reinforcement learning according to claim 4, characterized in that, The discriminator is optimized based on least squares loss or Wasserstein distance.
6. The prediction method based on inverse reinforcement learning according to claim 1, characterized in that, After the prediction model is trained, predictions are made based on the pre-trained decoder, including the following steps: The original spatial physical field is mapped to the latent space initial field by the pre-trained encoder and input into the trained prediction model to obtain the latent space prediction result. The latent space prediction result is then restored to the original spatial prediction result by the pre-trained decoder.
7. The prediction method based on inverse reinforcement learning according to claim 6, characterized in that, The pre-trained encoder and pre-trained decoder are obtained through joint training of an autoencoder structure, and the loss function of the autoencoder includes a reconstruction error term.
8. The prediction method based on inverse reinforcement learning according to claim 1, characterized in that, The prediction model employs Transformer, Mamba, or convolutional neural network architectures.