A numerical weather prediction method and system based on deep reinforcement learning

The DRL-EnVar algorithm based on deep reinforcement learning dynamically adjusts the mixing ratio of the background error covariance matrix, which solves the problem of insufficient scientific basis for the selection of mixing weights in numerical weather forecasting and achieves higher prediction accuracy and adaptability.

CN119940410BActive Publication Date: 2025-10-10NAT UNIV OF DEFENSE TECH

Patent Information

Application Number
CN202510007362.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-03
Publication Date
2025-10-10
Estimated Expiration
2045-01-03

AI Technical Summary

Technical Problem

Existing numerical weather forecasting methods lack a scientific basis for the selection of mixing weights in the background error covariance matrix, resulting in low efficiency in data sample utilization and inability to flexibly adapt to changes in different weather conditions, affecting forecast accuracy.

Method used

A method based on deep reinforcement learning is adopted to dynamically adjust the mixing ratio of the static background error covariance matrix and the flow-dependent background error covariance matrix through the DRL-EnVar algorithm, and an adaptive mixing parameter strategy model is established to improve the accuracy of numerical weather forecasting.

Benefits of technology

In the case of sparse and noisy observations, it can respond to changes in weather conditions in a timely manner, improve the prediction accuracy and performance of numerical weather forecasts, and the forecast effect is particularly significant during mutation periods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119940410B_ABST
    Figure CN119940410B_ABST
Patent Text Reader

Abstract

The application discloses a numerical weather prediction and forecasting method and system based on deep reinforcement learning. The method comprises the following steps: setting a custom environment of ensemble variational data assimilation; obtaining policy network parameters and value network parameters; obtaining updated policy network parameters based on a deep reinforcement learning method according to the custom environment of ensemble variational data assimilation, the policy network parameters and the value network parameters; determining a hybrid background-error covariance matrix according to the updated policy network parameters, wherein the hybrid background-error covariance matrix is a weighted average of a static covariance matrix and an ensemble covariance matrix; updating adaptive hybrid parameter strategy model training according to the hybrid background-error covariance matrix to obtain a numerical weather prediction and forecasting model; and performing numerical weather prediction and forecasting according to the numerical weather prediction and forecasting model. The application can improve the prediction accuracy of numerical weather prediction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of weather forecast technology, and in particular to a numerical weather forecast method and system based on deep reinforcement learning. Background Art

[0002] Data assimilation was first proposed to address the need for initial values ​​in numerical model integration. It aims to solve the inverse problem in numerical forecast models. Its core idea is to define the most accurate atmospheric motion state possible by fully utilizing all available information. In data assimilation systems, short-term forecast results from numerical models are often used as background field information. However, numerical models essentially approximate real weather systems based on extensive prior knowledge, resulting in systematic errors. This necessitates the introduction of real observational data into the data assimilation analysis loop to correct the background field and thus provide the optimal initial field for the model forecast at the next moment.

[0003] Although the demand for global weather forecasts has intensified in recent years, and countries have established relatively comprehensive meteorological observation networks encompassing ground-based, airborne, and space-based observations, in the actual operational data assimilation process, the volume of global multi-source observational data available at a given moment is far less than the dimensionality of the numerical forecast model state due to a variety of objective and subjective factors. This results in the dominance of the model background field in the assimilation process. ECMWF has conducted a quantitative assessment of the contribution of various types of information to the analysis field. The results show that approximately 15% of the analysis field information obtained through the data assimilation system comes from observational data during the assimilation period, while the remaining 85% comes from the model background field.

[0004] Despite this, the background error covariance matrix B plays a crucial role in data assimilation, particularly in characterizing background field uncertainty, ensuring the effective propagation of a small amount of observational data through the matrix. Specifically, the background error covariance matrix not only effectively propagates observational information but also smoothes observation increments, introduces balancing properties, and constructs flow structures, thereby ensuring the smooth assimilation process. This mechanism is particularly significant when observational data are scarce and noisy. The design and optimization of the background error covariance matrix B, to a certain extent, determines the overall performance of the data assimilation system.

[0005] The background error covariance matrix B, as a statistic, describes the probability distribution function of model prediction errors (i.e., background field errors in data assimilation systems). However, a key challenge in actual modeling is the unavailability of the "true" state, making it impossible to accurately calculate the distribution of the true background error. To overcome this limitation, background errors are typically simulated under specific assumptions. A common operational simulation method is the NMC (National Meteorological Center) method proposed by Parrish and Derber. This method simulates prediction errors by comparing the integrated differences between model forecasts at the same time but different forecast times. Because the NMC method is theoretically concise and easy to implement, it is not limited by the resolution of the observation network and can generate an effective background error covariance matrix B that includes the balance and constraint relationships between different variables. Therefore, it is widely used in global meteorological assimilation systems and has been adopted by most numerical forecast centers.

[0006] However, the NMC method typically calculates background error statistics based on model forecast differences over a considerable period of time (e.g., a year). In this approach, the error variations are primarily due to the natural variability of the model, resulting in a background error covariance matrix B that is climatological, static, and isotropic. However, the actual background error covariance matrix B often exhibits complex properties such as significant anisotropy, flow dependence, and baroclinicity, characteristics that are not fully captured by the NMC method. To address this, researchers proposed an ensemble-variational hybrid assimilation method. This method combines the real-time generated background error covariance matrix Be with the static background error covariance matrix Bs obtained by the NMC method through a linear combination, thereby more accurately reflecting the error characteristics that vary with weather conditions.

[0007] There are various methods for obtaining prediction ensemble samples, the most common of which is the Ensemble Kalman Filter (EnKF) method. In this method, an ensemble is generated simultaneously over a certain number of data assimilation cycles through reasonable perturbations (including initial value perturbations, observation perturbations, and parameterized perturbations). The resulting set of prediction data at the same moment can be used to estimate the prediction error covariance. However, as the number of ensemble members increases, the computational cost increases exponentially. Currently, the number of ensemble members that operational numerical forecast centers can tolerate is approximately 30 to 100, which is far smaller than the dimensionality of the numerical model state variables. As a result, the estimated forecast error covariance Be is usually non-full rank, has significant sampling errors, and is prone to spurious long-range correlations.

[0008] An effective and low-cost approach is to increase the number of EnKF ensemble members through the use of valid time-shifted ensemble members (VTS). VTS expands the ensemble size by adding subset members before and after the central analysis time. These include VTSM, which is used to sample time and / or phase errors, and VTSP, which eliminates spurious covariances through temporal smoothing. Another efficient alternative is the time-lagged ensemble forecast method, which directly uses the historical forecast fields of the variational assimilation system, with forecast results for the same time at different initial times as ensemble samples, greatly reducing computational and storage costs. The uncertainty of the time-lagged ensemble mainly comes from the differences in the initial fields, lateral boundaries, and observational data at different times. Therefore, it can effectively reflect the time-evolving flow-dependent forecast error covariance information.

[0009] However, because the credibility of time-lagged ensemble members is closely related to forecast timeliness, forecasts with longer timeliness often deviate from the true value and cannot be used as valid ensemble members. Therefore, the number of samples in a time-lagged ensemble is typically limited. To supplement the sample size of a time-lagged ensemble, the empirical orthogonal decomposition (EOF) technique can be used to select forecast samples from a wide range of historical samples that have similar variation characteristics of meteorological elements in the target assimilation period and region. This method, referred to as the preferred historical forecast sample, can effectively increase the number of time-lagged ensemble samples at a low computational cost, thereby introducing flow-related features that more accurately reflect current weather conditions.

[0010] Although numerous studies have demonstrated that ensemble-variational hybrid assimilation methods outperform single variational or ensemble methods, a common limitation is that the climatological background error covariance matrix B and the dynamically changing forecast error covariance Be are typically combined using a linear weighting scheme, with the blending weight often being a fixed empirical parameter (usually less than 1). This approach has two significant drawbacks: first, it inefficiently utilizes the data samples used to estimate the error covariance; second, the selection of the blending weight lacks scientific basis and cannot flexibly adapt to changing weather conditions. Therefore, a reasonable improvement is to adapt the blending parameter to the inherent characteristics of the data samples providing the background error information and to adapt it to the spatiotemporal changes of the weather conditions. This is also one of the core issues of this study.

[0011] When it comes to data feature extraction, traditional geosciences typically integrate known mathematical and physical mechanisms, construct numerical simulation models based on idealized assumptions, and utilize supercomputers to simulate and predict relevant weather or climate phenomena. This approach is undoubtedly limited by human understanding and cognition of Earth system processes, and inevitably incorporates excessive human intervention. In contrast to this theory-driven paradigm, data-driven machine learning methods, particularly deep learning, have flourished in recent years. Deep learning, through multi-layered neural network structures and nonlinear transformations, effectively combines low-level data features into abstract and easily distinguishable high-level representations. This allows for a more comprehensive understanding of the distributed nature of the data and transfers these representations to various downstream tasks, such as prediction, classification, and point cloud registration. End-to-end models, such as large language models, leverage massive databases for training and optimization through supervised or self-supervised learning to achieve natural language understanding and generation. This successful paradigm has brought tremendous potential to the field of "AI for Science," spawning numerous large-scale deep learning-based meteorological models that are completely independent of traditional numerical forecasting methods and ushering in a new scientific research paradigm. Although these models have demonstrated great advantages in improving the accuracy of global medium-term forecasts and accelerating the forecast inference process, the end-to-end intelligent forecast results are usually relatively smooth and have poor effects on predicting convective-scale weather with turning characteristics.

[0012] In operational numerical forecast simulations, because the physical mechanisms are not yet fully understood, researchers often rely on a large number of empirical parameters to closely simulate the evolution of actual weather phenomena while ensuring forecast availability. These empirical parameters include key parameters of different weather processes, key settings in various parameterization schemes, and mixing parameters in hybrid assimilation. This approach relies heavily on researchers' professional experience with weather phenomena and lacks a more scientific theoretical basis.

[0013] Excitingly, reinforcement learning in the field of machine learning offers new insights into this challenge. Reinforcement learning is a learning method that transitions from environmental states to action selection, with the goal of maximizing the cumulative reward of the agent in its interactions with the environment and ultimately obtaining the optimal strategy. This method of selecting actions based on prior states can alleviate the "interpretability" problem caused by the "black box" nature of deep learning models. Although a large number of studies are still attempting to incorporate physical information into deep learning models, these methods are collectively referred to as physical information neural networks (PINNs). These methods customize network models by using different activation functions, gradient optimization techniques, neural network structures, and loss functions in order to improve the physical consistency of the model.

[0014] However, the limitations of PINN are also quite obvious. It can only effectively solve problems with known physical models. For research problems that are not yet clear (such as complex precipitation processes), mechanically adding physical information may introduce more errors and even negatively affect the performance of neural networks. Therefore, although PINN has made progress in some fields, there is still much room for improvement in unresolved or unknown theoretical problems. SUMMARY

[0015] The application provides a numerical weather prediction method and system based on deep reinforcement learning, which can improve the prediction accuracy of numerical weather prediction.

[0016] To achieve the above object, the application provides the following scheme:

[0017] A numerical weather prediction method based on deep reinforcement learning comprises:

[0018] Setting a custom environment for ensemble variational data assimilation;

[0019] Obtaining policy network parameters and value network parameters;

[0020] According to the custom environment for ensemble variational data assimilation, the policy network parameters and the value network parameters, the updated policy network parameters are obtained based on the deep reinforcement learning method;

[0021] According to the updated policy network parameters, a hybrid background-error covariance matrix is determined, which is a weighted average of a static covariance matrix and an ensemble covariance matrix;

[0022] According to the hybrid background-error covariance matrix, the adaptive hybrid parameter policy model training is updated to obtain a numerical weather prediction model;

[0023] According to the numerical weather prediction model, numerical weather prediction is performed.

[0024] To achieve the above object, the application also provides the following scheme:

[0025] A system for numerical weather prediction based on deep reinforcement learning comprises:

[0026] A custom environment setting module for setting a custom environment for ensemble variational data assimilation;

[0027] A parameter acquisition module for acquiring policy network parameters and value network parameters;

[0028] A policy network parameter updating module is used to obtain updated policy network parameters based on the custom environment, policy network parameters and value network parameters assimilated by the set variational data based on a deep reinforcement learning method;

[0029] A mixed background-error covariance matrix determination module is used to determine a mixed background-error covariance matrix according to the updated policy network parameters, wherein the mixed background-error covariance matrix is ​​a weighted average of a static covariance matrix and a collective covariance matrix;

[0030] A numerical weather forecast prediction model training module is used to update the adaptive hybrid parameter strategy model training according to the hybrid background-error covariance matrix to obtain a numerical weather forecast prediction model;

[0031] The numerical weather forecast prediction module is used to perform numerical weather forecast prediction based on the numerical weather forecast prediction model.

[0032] To achieve the above object, the present invention also provides the following solution:

[0033] An electronic device includes a memory and a processor, wherein the memory is used to store a computer program, and the processor runs the computer program to enable the electronic device to perform a numerical weather forecasting method based on deep reinforcement learning.

[0034] To achieve the above object, the present invention also provides the following solution:

[0035] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements a numerical weather forecasting method based on deep reinforcement learning.

[0036] According to the specific embodiments provided by the present invention, the present invention discloses the following technical effects:

[0037] The present invention provides a numerical weather forecasting method based on deep reinforcement learning. The method includes: setting a custom environment for ensemble variational data assimilation; obtaining policy network parameters and value network parameters; obtaining updated policy network parameters based on the custom environment, policy network parameters, and value network parameters for ensemble variational data assimilation using a deep reinforcement learning method; determining a mixed background-error covariance matrix based on the updated policy network parameters, wherein the mixed background-error covariance matrix is ​​a weighted average of a static covariance matrix and an ensemble covariance matrix; updating an adaptive mixed parameter policy model training based on the mixed background-error covariance matrix to obtain a numerical weather forecasting model; and performing numerical weather forecasting based on the numerical weather forecasting model. The present invention can improve the prediction accuracy of numerical weather forecasting. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0039] Figure 1 This is a flow chart of the numerical weather forecasting method based on deep reinforcement learning of the present invention;

[0040] Figure 2 This is the DRL-EnVar framework diagram;

[0041] Figure 3 This is an implementation example of C-CNN;

[0042] Figure 4 This is a system structure diagram of the numerical weather forecast based on deep reinforcement learning in the present invention. DETAILED DESCRIPTION

[0043] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0044] The present invention provides a numerical weather forecasting method and system based on deep reinforcement learning, which can improve the prediction accuracy of numerical weather forecasting.

[0045] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.

[0046] Example 1:

[0047] Figure 1 This is a flow chart of the numerical weather forecasting method based on deep reinforcement learning of the present invention, as shown in Figure 1 As shown, the present invention provides a numerical weather forecasting method based on deep reinforcement learning, the method comprising:

[0048] Step 101: Set up a custom environment for ensemble variational data assimilation.

[0049] The goal of the EnVar assimilation system is to use error information from two sources of information (B hand R), effectively integrating the background field and the observation field, providing the optimal analysis field at the target time (time i) for the numerical prediction model, so that the numerical prediction model uses the analysis field as the initial field (in order to more intuitively evaluate the quality of the analysis field, the present invention defaults to the assimilation system based on the Lorenz96 model and can ignore the initialization and directly perform subsequent model forecasts). After a period of numerical forecasting (time i), the numerical model has the best forecast performance, that is, the forecast result is close to the real field, that is: argmin||x f -x t || I , where x t The real field at the target moment; The analysis field at time i For numerical mode The initial field of , the forecast duration is (Ii), and the model forecast field at time I is obtained. Weight in It is expressed as a mixture of static background error covariances (subscripted as static, abbreviated as st) and flow-dependent background error information (subscripted as flow-dependent, abbreviated as fl), and the parameters satisfy α1+α2+...+α m +β1+β2+...+β n =1.

[0050] The research background of this invention is that under the premise of sparse observation (masked) and noise, the continuous assimilation forecast cycle is one year and there is a mutation period (one month) during the period. The variable value during the mutation period is significantly higher than the variable value of the climate state, simulating strong convective weather conditions such as heavy rainfall. In the above example, the two statistical results R and H based on the same and unchanged observation data are completely known and fixed, then B h The accuracy of the description of the background error information of the real-time state variable and the degree of timely response to the sudden change, that is, the degree of characterization of the background error flow dependency information, determines the accuracy of the analysis field. h It mainly depends on the static background error covariance obtained by statistics of historical data samples and the collective background error covariance with flow dependence calculated by collective data samples, as well as the appropriate selection of the fusion parameters α and β of multiple background error information sources (such as model Figure 2 (part (c)).

[0051] In order to restore the background error information as close to the real background error information as possible, the present application will make full use of all data samples that can provide effective background error information, including multiple historical climate data samples and real-time ensemble prediction samples, to statistically obtain background error information without significantly increasing the statistical calculation cost. Under this precondition, the present application mainly discusses the influence of intelligent selection of mixing parameters of different background error information sources on the system assimilation performance, and emphasizes the real-time and self-adaptability of the mixing parameters.

[0052] According to the above description, the task of selecting the mixing matrix parameters in real time in the hybrid assimilation prediction cycle is expressed as a decision-making task of mapping from the current state (analysis field) to action, that is

[0053] f(s i )=a i

[0054] a i =[α1,...,α m ,β1,...,β n ]

[0055] Where m represents the number of static B, and n represents the number of flow-dependent B.

[0056] Formalized as an MDP, as follows:

[0057] ·State is the vector given by the EnVar system at time step i.

[0058] ·Action a i is a set of different weighting factors, with the constraint that the sum is equal to 1.

[0059] ·State transition s i+1 is the numerical solution that minimizes the cost function in our task:

[0060]

[0061] ·Reward R(s i , a i , s i+1 is the immediate reward for taking action a i in state s i and reaching new state s i+1 .

[0062] ·The discount factor γ is set to 0.99 because our task involves long-term continuous planning, and the future rewards brought by continuous actions are crucial.

[0063] The goal is to trust and retain the complete framework of the EnVar assimilation system, and to more finely characterize B hEnabling, making it possible to effectively deliver mode evolution related information flow, improving the assimilation performance of EnVar assimilation system, especially in the mutation period, can respond in time and make effective adjustment, and the mixed B h , which is crucial for the subsequent assimilation prediction cycle to get a better analysis field. Specifically, given the initial field of the numerical model, based on the available known observations, the EnVar system is used for assimilation prediction cycle, which needs to find an optimal adaptive strategy for mixing weights of multiple background error information sources θ , so that the assimilation prediction cycle can run stably in the climate state stable period or mutation period of numerical model evolution, and the assimilation performance is as high as possible.

[0064] Deep reinforcement learning in the field of mixed data assimilation is relatively less than other fields such as robot control and automatic driving. The invention does not simply use the classic deep reinforcement learning neural network based on MLP (model parameters are large, and the training cost is large), but according to the characteristics of the numerical prediction model and the assimilation prediction cycle system verified by experiment, an innovative and intelligent algorithm is designed, called DRL-EnVar. The algorithm model needs to reduce the computational complexity as much as possible, reduce the neural network parameters to control the model training time, and still be able to extract effective information from limited available data. The information generated by the agent and the environment is interacted, the reasoning is completed, and the optimal strategy of adaptive selection of mixing parameters is learned, especially when there is a mutation period in the evolution of the numerical model, the strategy can still be used stably, ensuring high assimilation performance.

[0065] DRL-EnVar algorithm is an intelligent hybrid assimilation strategy algorithm proposed by the invention, aiming to solve the real-time adaptive mixing parameter selection problem in EnVar hybrid assimilation prediction cycle. In this cycle, the mixing proportion of static background error covariance matrix B and flow-dependent background error covariance matrix (flow-dependent B) needs to be dynamically adjusted, so that the algorithm can effectively respond to changing observations and prediction situations. This task can be regarded as a decision problem of continuously selecting the optimal action in a high-dimensional state space, which conforms to the classic decision problem type in reinforcement learning. First, define the EnVar self-defined environment.

[0066]

[0067]

[0068] The above algorithm is the definition process of EnVar self-defined environment. First, give the initial background field of Lorenz96 model , that is, X j = F, if j≠20; X j = 1.001F, if j=20.

[0069] The first state of the environment has not yet been assimilated and is directly determined by Assign to As the initial state of the environment. According to the above analysis, the present invention provides a total of 5 B matrices of background error information, so that the action space agent needs to select 5 mixing parameters to perform an action, and the sum of the parameters is always 1. The reward function consists of two parts. One part is to analyze the field x after each assimilation. a With real field x t The distance, that is, the root mean square error RMSE a To directly measure the quality of the analysis field; the other part is assimilated, and the assimilated analysis field is used as the initial field to make a 48h forecast. RMSE root mean square error with the real field f To evaluate the forecast performance of the numerical forecast model, and thus indirectly measure the quality of the current analysis field. The larger the two root mean square errors are, the greater the penalty for the action taken by the agent, that is, R = -RMSE a -RMSE f .

[0070] Regarding the state transition rules, the research context of this paper is based on a continuous assimilation forecast cycle of one year (assuming each month is 30 days, a year is 360 days). Furthermore, to maintain the stability of the Loren96 model and eliminate transient behavior, 90 days of initial Loren96 model startup data are discarded, resulting in a total of 450 days of continuous model integration. The numerical solution method uses the fourth-order Runge-Kutta format (RK4) with a time step of dt = 0.05 (conventionally 6 hours). This means that four assimilation cycles are performed per day, resulting in a total of 1800 iterations (line 5 of the pseudocode in Algorithm 3.1).

[0071] In order to study whether the B matrix can transmit the flow-related information in time during the mutation period, a switch is set in the state transition rule. If the pattern evolution mutation index is detected, the forcing term of the numerical model is set to 15.0, otherwise it remains at 8.0. In addition, due to the flow-dependent B 48 and B 24 A certain number of collective samples need to be accumulated to perform statistical calculations, so Figure 2 (c) It is necessary to complete 8 assimilations (num_DA=8) before performing a flow-dependent B 48 and B 24 , thereby updating B h .

[0072] If num_DA<8, based on repeated experimental analysis, directly use the climate state background error information that meets the experimental settings, that is, BF8S15 As B h The EnVar system assimilation performance is relatively the best. Based on the above definition, the EnVar custom environment finally returns the state s i , action a i , reward r i , and provided to the DRL-EnVar algorithm for subsequent training.

[0073] This step specifically includes:

[0074] Customize the reinforcement learning simulation environment based on the assimilation and forecasting cycle of the ensemble variational data assimilation system, design the state space and action space of the intelligent agent, and establish a decision-feedback reward mechanism to obtain the reward function at the current moment;

[0075] Get the current action and status;

[0076] A custom environment that assimilates the reward function, actions, and states as collective variational data.

[0077] The EnVar hybrid scheme of the ensemble variational data assimilation system is built on the basis of the existing 3DVar system and directly introduces the integrated information through the error covariance matrix.

[0078]

[0079] Among them, x represents the state variables, including model variables such as temperature, wind force, pressure and humidity; x b Indicates the background state, y o Represents observation data, all data are three-dimensional data. Assume that the nonlinear observation operator is linear and completely known, simplifying to R is the observation error covariance matrix, which can be obtained from the instrument's observation error to describe the observation value y o Correlation with the model grid points.

[0080] It is worth noting that the mixed background-error covariance matrix B h is defined as the static covariance matrix and the collective covariance matrix (B e and B s ), effectively replacing B in the original 3DVar system.

[0081] B h =(1-β)B s +βB e

[0082] Among them, β (0<β<1) is an adjustable factor that controls B e and B s The weight of .

[0083] Step 102: Obtain policy network parameters and value network parameters.

[0084] Step 103: Based on the custom environment, policy network parameters and value network parameters assimilated by the ensemble variational data, an updated policy network parameter is obtained based on a deep reinforcement learning method.

[0085] Deep learning, a subset of machine learning, is a general artificial intelligence approach that utilizes multi-layer neural networks to simulate complex patterns and representations in large datasets. Deep learning models are typically composed of multiple layers of nonlinear computational units. They automatically learn high-level, abstract feature representations from large amounts of training data by feeding the outputs of lower layers into higher layers. This allows them to fully perceive the distributed characteristics of the data, demonstrating superior performance in various tasks such as image and speech recognition and natural language processing. Among the numerous architectures that have emerged, fully connected neural networks (FCNNs), convolutional neural networks (CNNs), and the gated recurrent unit (GRU) in recurrent neural networks (RNNs) have attracted significant attention due to their effectiveness in processing specific types of data and tasks.

[0086] Fully connected neural networks (FCNNs), also known as dense neural networks, are a class of artificial neural networks with extensive connectivity. Fully connected neural networks represent the most basic form of deep learning, in which every neuron in one layer is connected to every neuron in the next layer. This dense connectivity enables FCNNs to model complex nonlinear relationships in data, making them suitable for a wide range of applications, including classification, regression, and pattern recognition. A special and widely used type of FCNN is the Multi-Layer Perceptron (MLP).

[0087] Convolutional neural networks (CNNs) are specifically designed to process grid-like data structures, such as images. CNNs are inspired by biological processes in the visual cortex, where neurons are arranged to cover the visual field, using convolutional layers to automatically and adaptively learn a spatial hierarchy of features from the input image. These layers apply convolution operations to capture local patterns, such as edges and textures, which are then aggregated in deeper layers to form more complex structures and objects. This hierarchical feature extraction has enabled CNNs to achieve remarkable performance in image recognition, object detection, and related tasks.

[0088] Recurrent Neural Networks (RNNs) are designed to process patterns in sequences of data, such as time series or natural language. The core component of an RNN is the recurrent unit, which processes input data sequentially, updating its internal state based on the current input and previous state. This feedback mechanism allows the network to capture temporal dependencies and long-term correlations in the data. However, standard RNNs face challenges such as vanishing and exploding gradients, which can hinder their ability to learn long-term dependencies. To address this issue, the gated recurrent unit (GRU) was introduced as a variant of the RNN architecture. As a simplified variant of the Long Short-Term Memory (LSTM) network, the GRU combines the functionality of the input gate and the forget gate into a single update gate, simplifying the architecture while maintaining similar performance. By using update and reset gates, the GRU effectively manages the information passing through the network, resulting in better performance in tasks such as time series prediction, language modeling, and speech recognition.

[0089] Reinforcement learning (RL) is a subfield of machine learning (ML) that focuses on training an agent to make sequential decisions by maximizing cumulative rewards through interactions with an environment. The framework is based on the Markov decision process (MDP), which is characterized by a tuple (S, A, P, R, γ), where S is the state space, A is the action space, P represents the state transition probability, R is the reward function, and γ is the discount factor.

[0090] RL problems are mainly solved by value-based and policy-based methods. Value-based methods, such as Q-learning and DQN, derive optimal policies by optimizing value functions and are applicable to discrete environments such as Go. Policy-based methods, such as policy gradients, gradually improve policies, making them applicable to continuous action scenarios such as robot control. A well-known method in RL is the Actor-Critic framework, which combines these two methods to solve problems in continuous action spaces and high-dimensional state spaces. The framework uses two networks: an actor-network, which is used to generate parameterized action policies; and a critic-network, which is used to evaluate the value of state-action pairs. This dual structure combines value function approximation with direct policy optimization, improving learning efficiency and stability.

[0091] Proximal Policy Optimization (PPO) is an advanced RL algorithm that improves the Actor-Critic framework by introducing a clipped proxy objective function to achieve stable policy updates. PPO solves the high variance and instability problems of traditional policy gradients, achieving a balance between simplicity and performance, making it the preferred method for complex RL tasks.

[0092] To better extract abstract representations of cyclical data, we also proposed a cyclical one-dimensional convolutional network (C-CNN) to remove symmetries in the data and improve learning efficiency. Considering the specificity of this research task, a softmax module was implemented to ensure that the sum of the mixing parameters is always 1. Cyclic-Convolutional Neural Network (C-CNN): C-CNN is a specialized convolutional architecture designed for Earth system data. The numerical forecast model used in the experimental verification is Lorenz96, which features J variables at equidistant grid points around the equator, resulting in a data structure that is connected end-to-end in a cyclical manner. Due to atmospheric circulation patterns, variables at adjacent longitudes within the same latitude circle influence each other. C-CNN is inspired by one-dimensional convolution. As shown in the figure, considering the physical spatial dependencies of the observations, conventional one-dimensional convolution cannot model cyclical physical spatial information, resulting in a decrease in modeling capability. Therefore, we use a cyclic convolution that is symmetric in physical space to replace one-dimensional convolution. After multiple rounds of circular convolution, the model can extract higher-level abstract features, and this feature will not be affected by the symmetry of the physical space (that is, when the data is initialized, no matter where the starting point of the ring is, the same representation can always be obtained in the end), so that the next step of processing can be done in a smaller representation space, so that the model can achieve better performance on less data. In addition, after each circular convolution, the RELU (Rectified Linear Unit) activation function is used to add nonlinear features to the neural network to help the network better learn data features. That is, the input features are transformed element by element nonlinearly. On the one hand, for negative inputs, the ReLU function outputs 0, which is equivalent to setting some neurons to an inactive state, thereby making the neural network have a certain sparsity and reducing the correlation between parameters; on the other hand, for positive intervals, the ReLU function maintains the gradient of 1, which can effectively alleviate the problem of gradient disappearance. In addition, combined with the BatchNorm operation, it helps to make model training more stable and accelerate convergence.

[0093] Figure 3As an example of the implementation of the C-CNN, the input channel number of the C-CNN is 1, the number of convolution kernels is 8, the size of each convolution kernel is 5, and the output channel is 8. Starting from any starting point of the ring data, each convolution kernel moves counterclockwise along the ring data, and the value of the corresponding position is multiplied and summed every time it moves to a position. For example, one convolution process is: the yellow dashed oval frame in the input channel data is multiplied by 8 convolution kernels and summed, and 8 values are output to form 8 output channel data (8 yellow dashed rectangular frames in the output channel data). Then the convolution kernel moves counterclockwise by one step to the blue dashed frame data position for convolution operation, then to the orange and green, and so on, until the beginning and end are connected, and the complete ring convolution operation is completed. The values of each convolution kernel are different, and are learned and updated together with the strategy network model parameters during model training.

[0094] Step 103 specifically comprises:

[0095] The self-defined environment, strategy network parameters and value network parameters of the ensemble variational data assimilation are based on a deep proximal policy optimization algorithm to obtain updated strategy network parameters.

[0096]

[0097]

[0098] The above algorithm describes the training process of the DRL-EnVar algorithm for adaptive hybrid parameter selection.

[0099] According to extensive research and comparison of deep reinforcement learning (DRL) methods, it is found that PPO (Proximal Policy Optimization) has superior performance and robustness in current DRL algorithms. Compared with other methods (such as DQN, TRPO, DDPG, etc.), PPO has the advantages of low parameter sensitivity, stable update, wide applicability, etc., especially suitable for optimization problems in continuous space, which makes it the preferred algorithm in the research and application of reinforcement learning. Therefore, PPO is used as the reinforcement learning algorithm in this study to ensure the training convergence and model stability in high-dimensional complex environments, and to improve the effect of adaptive hybrid parameter selection.

[0100] The core of the DRL-EnVar algorithm consists of two components: an actor and a critic, forming a typical actor-critic architecture. The actor network is responsible for generating actions, specifically selecting appropriate parameter ratios during the hybrid assimilation process; the critic network estimates a value function, used to evaluate the effectiveness of the current strategy. Both utilize an MLP (Multi-Layer Perceptron) architecture. Based on previous experience in data assimilation and forecasting tasks, the MLP has been proven to be a versatile approximator with excellent performance in the spatiotemporal series assimilation-forecasting cycle.

[0101] Specifically, the PPO algorithm introduces a clipping operation in actor updates. This restricts the range of the ratio of the old and new policies to control the policy gradient update amplitude, thereby avoiding training instability caused by excessive policy updates. The clipping range ε ​​is set to 0.2, an empirical value. The critic network is updated using temporal difference (TD) error, minimizing the TD error to improve the accuracy of the value function estimation. Furthermore, to fully extract data features, this algorithm combines a C-CNN and MLP architecture within a deep learning network. The C-CNN extracts spatiotemporal features, while the MLP estimates the policy and value. This architecture ensures that the model effectively captures data features and selects optimal mixing parameters in the complex EnVar assimilation-forecasting cycle. Ultimately, the trained DRL-EnVar model can dynamically adjust the mixing parameter ratio during real-time assimilation, achieving optimal assimilation results under different meteorological and observational scenarios.

[0102] Step 104: Determine a mixed background-error covariance matrix based on the updated policy network parameters, where the mixed background-error covariance matrix is ​​a weighted average of a static covariance matrix and a collective covariance matrix.

[0103] This step specifically includes:

[0104] According to the updated strategy network parameters, use formula B h =(1-β)B s +βB e , determine the mixed background-error covariance matrix;

[0105] Among them, B h is the mixed background-error covariance matrix, B e is the static covariance matrix, B s is the collective covariance matrix, β is the updated policy network parameter, which controls B e and B s The weight of , 0<β<1.

[0106] The static B and the B with flow dependence characteristics are both calculated by statistical methods, but the methods used are different. First, the static B is calculated by the most commonly used NMC method in business, that is, the difference between different lengths of forecast pairs is used, but each forecast pair is valid at the same time. The formula is:

[0107]

[0108] It can be seen that the key of the NMC method lies in the composition of the historical data sample (3 years) used for statistical calculation. Three historical data samples are used in this study, which respectively reflect the different climate characteristics of numerical models. The first data sample completely follows the research background of the present application, that is, a month of mutation period occurs year by year, and the Lorenz96 forcing term F = 15.0 in this period, and the remaining 11 months are the regular stable period F = 8.0; the second data sample assumes that the whole year is a mutation period, the purpose is to more prominently concentrate the error correlation between the state variables when F = 15.0, and it is expected that this data sample can provide significant correlation error information when the mutation period occurs in the research EnVar assimilation forecast cycle; the third data sample assumes that the whole year does not occur mutation, that is, F is always 8.0, because the mutation period is relatively low in frequency compared with the regular and stable change of the model, that is, more times in a year the model is relatively slowly evolving, and the background error information of this climate state is very important.

[0109] Regarding the real-time B that provides flow-dependent information, in order to avoid the unacceptable computational cost brought by ensemble prediction, and to ensure the effective statistics and transmission of flow-dependent information as much as possible, the present application uses two real-time ensemble data samples. Part of it is inspired by the acquisition of time-lag ensemble samples in variational assimilation, that is, the difference between the model forecast fields at the same time but different effective times in the historical forecast sample of the EnVar system, that is, the 24h forecast value and the 48h forecast value in the blue highlight and green highlight on the left in figure (c);

[0110] Another part is inspired by the acquisition of time-lag ensemble samples in ensemble assimilation, that is, considering that numerical model prediction may have a certain degree of temporal and spatial phase error, especially strong convective weather phenomena such as heavy rain, which usually have a time error of large value prediction in advance or delay and a spatial deviation of large value area, the spatiotemporal adjacent forecast values at a target time are all regarded as the forecast values at the current time, so as to alleviate the error rate and effectively increase the number of ensemble samples, and to alleviate the sampling error caused by insufficient ensemble samples and other problems, that is, the 21h and 27h forecasts in the blue highlight in figure (c) are regarded as the 24h forecast, and the 45h and 51h forecasts in the green highlight are regarded as the 48h forecast.

[0111] The two parts of samples together constitute two real-time aggregate data samples. It should be noted in advance that the displacement prediction in this unrepresented space will be realized by the subsequent circular convolutional neural network, which will be explained in detail later. Specifically, the present invention performs assimilation every 6 hours, that is, fuses the observation field and the background field, and performs a 48-hour forecast. The forecast results are saved every 3 hours, and then B is calculated. 24 The number of set members is 12. According to the calculation principle of VTSP (reference), B is calculated. 24 . Accordingly, calculate B 48 The number of set members is 24. Statistics B 48 and B 24 The forecast timeliness of the two samples is significantly different, and the accumulated errors are different, which provides different flow-dependent error-related information for the hybrid assimilation system.

[0112] Step 105: updating the adaptive hybrid parameter strategy model training according to the hybrid background-error covariance matrix to obtain a numerical weather forecast prediction model;

[0113] Step 106: Perform numerical weather forecast prediction according to the numerical weather forecast prediction model.

[0114] Figure 2 The overall process of the DRL-EnVar algorithm is shown, which can be divided into three main stages: the adaptive selection of mixing parameters, the calculation of the mixed B matrix, and the h ) and the EnVar hybrid system assimilation forecast cycle. The three stages are closely linked and influence each other to achieve an intelligent hybrid assimilation system.

[0115] In the DRL phase (a), we first customize the reinforcement learning simulation environment based on the EnVar hybrid assimilation system's assimilation prediction cycle. We design the agent's state and action spaces, and establish a long-term "decision-feedback" reward mechanism. Through the agent's continuous interaction and exploration with the environment, and using the PPO proximal policy optimization algorithm to train the model, we continuously update the policy network's parameters θ. The goal is to maximize the expected cumulative reward and obtain an optimal strategy for selecting hybrid parameters that adapt to the current state. During this phase, to better extract abstract representations of cyclic data, we also propose a cyclic one-dimensional convolutional network (C-CNN) to eliminate symmetries in the data and improve learning efficiency. Considering the specificity of this research task, we set a softmax module to ensure that the sum of the hybrid parameters is always 1.

[0116] Continue to calculation B h Stage (c), based on the action a given in the DRL stage i ( Figure 2Arrow ①), the mixing parameters are multiplied by the corresponding background error covariance matrix, among which there are three static B obtained by one statistical analysis of historical data samples (respectively B F8S15 、B F8 and B F15 ) and two Bs with flow-dependent characteristics calculated from real-time data samples (respectively B 48 and B 24 ).

[0117] In the EnVar stage, based on the B calculated in the previous stage h After completing an assimilation, the analysis field is passed to the DRL stage as the state part of the next agent collection information ( Figure 2 Arrow ③), conduct subsequent adaptive hybrid parameter strategy model training and update strategy model parameters. In addition, each assimilation is performed, a 48h forecast is performed, and the forecast results are saved every 3 hours. The iterative stacking is provided as the calculation flow dependent B 48 and B 24 Real-time data samples.

[0118] The numerical weather forecasting method based on deep reinforcement learning is an intelligent ensemble variational hybrid data assimilation method based on deep reinforcement learning (DRL-EnVar) to improve the performance of the assimilation system under the conditions of sparse observations and important turning points in weather evolution. This method is based on the traditional EnVar framework and directly uses the B estimated by the ensemble sample to e It is embedded in a variational cost function for iterative assimilation. To meet the needs of high-frequency assimilation, a time-lagged ensemble method is primarily employed, which is highly time-efficient and supplemented by other efficient ensemble expansion strategies. By leveraging the robust feature extraction capabilities of DL and the intelligent decision-making capabilities of RL, DRL-EnVar significantly improves assimilation performance, providing a physically consistent optimal initial state for the NWP model, thereby optimizing forecast accuracy.

[0119] This experiment uses the Lorenz96 chaotic model, widely used in assimilation research, as the numerical model. The model's initial fields consist of a set of state vectors with a given number of variables, N. The true value field is generated by evolving the model over an integration step, with the model forcing term set to F = 8.0 or F = 15.0. Short-term forecast fields are also generated by model integration, but to simulate the error characteristics of model forecasts, the Lorenz96 model's forcing term is adjusted to F = 8.4 or F = 15.75 (references). The assimilation framework is based on three-dimensional variational assimilation (3DVar), with simulated observations assimilated every six hours. Observations are generated by superimposing small Gaussian random perturbations on the model's true value field. The Gaussian noise has a mean of zero, and the covariance matrix is ​​the specified observation error covariance. The observation errors are assumed to be spatially independent and uncorrelated, meaning the observation error covariance matrix is ​​the identity matrix with a scaling factor (I), and the observation error standard deviation is set to 1.0 by default. The background error covariance matrix uses a variety of mixed background error covariance schemes. The specific settings are detailed in the comparative experimental scheme section. The validation period of the cyclic assimilation system experiment is set to one year, with an initial spin-up period of 90 days.

[0120] According to the experimental premise and verification objectives, two groups of comparative experiments were designed.

[0121] The first set of experiments aims to evaluate the performance of a hybrid assimilation scheme that uses reinforcement learning to intelligently select blending weights for the time-lagged ensemble background error covariance and the static background error covariance under conditions of varying observation sparsity. The experiments focus on verifying whether the ensemble background error covariance, derived from the statistics of the time-lagged ensemble members, can be used to select appropriate blending weights during the blending process, effectively improving the assimilation system's performance.

[0122] The second set of experiments examines transitional weather processes under sparse observation conditions. During a cyclic assimilation cycle, when the system experiences a sudden change, the experiments investigate whether, through intelligent weighting through reinforcement learning, multiple static background error covariances and multiple time-lagged ensemble background error covariances can effectively capture background error information about real-time state variables and rapidly respond to changes in the current meteorological situation, thereby improving assimilation performance during the sudden change period. Furthermore, the experiments focus on comparing the adaptability and responsiveness of the mixed background error covariance matrix across different methods under the sudden change period, verifying their comprehensive improvement in year-round assimilation performance.

[0123] The two sets of experiments compared six methods: 3DVar, a fixed-weight time-lagged hybrid assimilation method (CTL-EnVar), EnKF, a hybrid assimilation method (MLP-EnVar) that uses only a multi-layer perceptron (MLP) for feature extraction in a deep reinforcement learning model, a CCNN-based reinforcement learning method proposed in this paper (DRL-EnVar), and the classic hybrid covariance data assimilation method (HCDA). To mitigate sampling errors and long-range spurious correlations caused by a small number of ensemble members, all methods except pure 3DVar localized the ensemble error covariance matrix. This localization employed the Gaspari-Cohn function (a fifth-order piecewise rational function), whose weights decreased from 1 to 0 within a threshold radius using a Gaussian distribution approximation.

[0124] Experiment 1: Comparative experimental setup with sparse observations and no mutation

[0125] In this experiment, all observational data were sparsified using a uniform masking method, masking out one observation value for every n state variables. Since the total number of model state variables used in the experiment was 40, to ensure mask uniformity, n = 9, 4, 3, and 1 were chosen, corresponding to observation masking ratios of 10%, 20%, 25%, and 50%, respectively. Furthermore, this experiment did not involve transitional weather processes; that is, F = 8.0 was maintained throughout the cyclic assimilation cycle. The main objectives of the experiment were twofold: first, to verify whether, under the condition of sparse observations, the time-lagged ensemble members counted through the cyclic assimilation cycle can provide flow-dependent background error information and effectively propagate observational information from observed variables to unobserved variables; second, to test whether a deep reinforcement learning method, combined with deep learning to extract data features, can optimize the blending weights based on a long-term reward mechanism and the current environmental state, thereby improving the utilization efficiency of multiple background error information sources and significantly improving the average assimilation performance over the validation cycle. This experimental design aims to explore the potential of combining time-lagged ensemble background error covariance with deep reinforcement learning under the special conditions of sparse observations, and provide practical support for improving the performance of the cyclic assimilation system.

[0126] Experiment 2: Comparative Experimental Setup with Sparse Observations and Mutation

[0127] On the basis of observational uniform masking, this experiment increases the complexity and specificity of the evolution of state variables during the verification period. Specifically, a one-month mutation period (June 15th to July 15th) was added to the cyclic assimilation period. During this period, the values ​​of the model state variables were significantly greater than those in the stable state period, simulating the plum rain weather phenomenon\cite{yihui2005east}. During these 30 days, the forcing term of the mutation period was set to F = 15.0, and the rest of the time (including the spin up period) was kept at F = 8.0, as shown in the figure. Since it involves a one-month state variable mutation period, this experiment changes the climate state characteristics during the verification period. Therefore, compared with Experiment 1, this experiment counts the background error covariance information under two different climate states, namely B F8S15 and B F15 , and conduct a comprehensive analysis on them.

[0128] All experiments were performed on a uniformly configured workstation equipped with a single GeForce RTX 3090 GPU and an AMD EPYC 7H12 CPU.

[0129] The fixed mixing parameters of the three methods are reproduced through the enumeration method adopted by each method itself, aiming to find the optimal fixed parameters.

[0130] (1) Fixed mixing parameter time lag ensemble - three-dimensional variational mixing assimilation method CTL-EnVar, where the mixing background error covariance calculation formula is: B h =(1-α)B nmc +αB TL

[0131] In the method, the construction process of the time lag set member is: in the 3DVar assimilation process, the model prediction fields at the same time but different response times in the historical prediction sample are used to construct the time lag set member. The specific operation is to perform assimilation and 48-hour prediction every 3 hours, save the prediction results every 3 hours, construct 16 time lag set members, and obtain 120 difference fields. In order to ensure the fairness of the comparative experiment, the other methods in this experiment are set to perform assimilation every 6 hours. Therefore, the CTL-EnVar method is also changed to perform assimilation and 48-hour prediction every 6 hours. The static background error covariance matrix in the method is calculated by the NMC method, and different from the four mixing parameters 0.25, 0.5, 0.75 and 1.0 listed in the original text, the mixing parameter is regarded as the weight of the set background error covariance. In order to ensure that the method can achieve optimal performance, the enumeration range is expanded to 0 to 1, the preliminary verification step is 0.1, a total of 10 parameters, each mixing parameter is repeated 10 times, and the average value is taken as the final result. The experiment found that when the mixing parameter is 0.1, the performance is relatively optimal. In order to further confirm whether a smaller mixing parameter can bring better performance, the second enumeration range is 0.01 to 0.1, and the step is 0.01. It is found again that when 0.01, the performance is optimal. Then, the enumeration range is further refined to 0.001 to 0.01, the step is 0.001, and finally refined to 0.0001 to 0.001, the step is 0.0001. Finally, in two experimental verification scenarios, the optimal parameters of CTL-EnVar are concentrated in 0.0001 to 0.01, and the optimal mixing parameter is different under different observation sparse ratios and experimental settings.

[0132] (2) EnKF method: The purpose of setting EnKF as a comparative experiment is to verify whether the performance of the innovative method can approach or be equivalent to EnKF under different set member numbers, while the calculation cost is much lower than EnKF. It needs to be determined that the selection of set member number under different experimental settings. The enumeration range of set member number is 5 to 40, the step is 5, and the upper limit is set to 40, because the number of state variables of the model used in the application is 40. Compared with the business EnKF method, the ratio of the number of state variables to the number of set members is 1:1, which is almost impossible to achieve, and will bring extremely high calculation cost. Each enumeration experiment is repeated 10 times and the average value is taken.

[0133] (3) HCDA method: The purpose of setting up HCDA as a comparative experiment is to verify whether the method proposed in the present invention can be equivalent to or close to HCDA under different numbers of set members, while the computational cost is much lower than HCDA. The parameters that need to be determined include the selection of the number of set members and the weight parameter between the static background error covariance and the set prediction error covariance. The enumeration range of the number of set members is 5 to 40, with a step size of 5, and the enumeration range of the weight parameter is 0 to 1, with a step size of 0.1. The experiment found that the number of set members and the weight parameter jointly affect the performance of HCDA, which is manifested in that under the same number of set members, as the weight parameter increases, the performance first significantly improves and then gradually decreases. The optimal performance under different numbers of set members is similar, that is, the increase in the number of set members does not significantly improve the performance of HCDA. Therefore, the enumeration range of the number of set members is narrowed to 2 to 5, with a step size of 1, and the enumeration range of the weight parameter is still 0 to 1, with a step size of 0.1. Each enumeration experiment is repeated 10 times and the average value is taken.

[0134] The present invention uses two evaluation indicators to evaluate the assimilation performance of the assimilation method within the experimental verification cycle: one is a direct evaluation indicator - the distance between the analysis field and the true field, that is, the root mean square error (RMSEa); the other is an indirect evaluation indicator - the distance between the forecast field and the true field, RMSEf. The forecast duration is selected according to different verification requirements, ranging from 3 hours to 48 hours, and the average performance of the forecast every 3 hours. In addition, for the special settings of Experiment 2, since the generation and disappearance cycle of convective-scale weather in the simulated plum rain weather phenomenon is usually in the order of hours or shorter, the performance comparison of the 1-hour forecast results during the mutation period is added. To avoid experimental contingency, the experiment of each method was repeated 50 times, and the average value was taken as the final evaluation result.

[0135] Example 2:

[0136] This embodiment provides a numerical weather forecast system based on deep reinforcement learning. Figure 4 This is a system structure diagram of the numerical weather forecast based on deep reinforcement learning in the present invention. Figure 4 As shown, the system includes:

[0137] A custom environment setting module 201 is used to set a custom environment for ensemble variation data assimilation;

[0138] Parameter acquisition module 202, used to obtain policy network parameters and value network parameters;

[0139] A policy network parameter updating module 203 is configured to obtain updated policy network parameters based on a deep reinforcement learning method according to the customized environment, policy network parameters, and value network parameters assimilated by the ensemble variational data;

[0140] A mixed background-error covariance matrix determination module 204 is used to determine a mixed background-error covariance matrix based on the updated policy network parameters, where the mixed background-error covariance matrix is ​​a weighted average of a static covariance matrix and a collective covariance matrix;

[0141] A numerical weather forecast prediction model training module 205 is used to update the adaptive hybrid parameter strategy model training according to the hybrid background-error covariance matrix to obtain a numerical weather forecast prediction model;

[0142] The numerical weather forecast prediction module 206 is used to perform numerical weather forecast prediction according to the numerical weather forecast prediction model.

[0143] Example 3:

[0144] This embodiment provides an electronic device, including a memory and a processor, wherein the memory is used to store a computer program, and the processor runs the computer program to enable the electronic device to execute the numerical weather forecasting method based on deep reinforcement learning of embodiment 1.

[0145] Optionally, the above-mentioned electronic device may be a server.

[0146] In addition, an embodiment of the present invention also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the numerical weather forecasting method based on deep reinforcement learning of embodiment one.

[0147] Embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0148] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0149] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0150] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0151] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Reference can be made to the common and similar parts between the various embodiments. For the systems disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the method description.

[0152] The present invention uses specific examples to illustrate the principles and implementation methods of the present invention. The above examples are only intended to help understand the method and core concept of the present invention. At the same time, those skilled in the art will find that the specific implementation methods and application scopes may vary based on the concept of the present invention. In summary, the contents of this specification should not be construed as limiting the present invention.

Claims

1. A numerical weather forecasting method based on deep reinforcement learning, characterized in that: The method comprises: Set up a custom environment for ensemble variational data assimilation; Get policy network parameters and value network parameters; Obtaining updated policy network parameters based on a deep reinforcement learning method according to the customized environment, policy network parameters, and value network parameters assimilated from the ensemble variational data; Determining a mixed background-error covariance matrix based on the updated policy network parameters, wherein the mixed background-error covariance matrix is ​​a weighted average of a static covariance matrix and a collective covariance matrix; Updating the adaptive hybrid parameter strategy model training according to the hybrid background-error covariance matrix to obtain a numerical weather forecast prediction model; Performing numerical weather forecast prediction according to the numerical weather forecast prediction model; The custom environment for setting up ensemble variational data assimilation specifically includes: Customize the reinforcement learning simulation environment based on the assimilation and forecasting cycle of the ensemble variational data assimilation system, design the state space and action space of the intelligent agent, and establish a decision-feedback reward mechanism to obtain the reward function at the current moment; Get the current action and status; A custom environment that assimilates the reward function, actions, and states as collective variational data; The customized environment, policy network parameters, and value network parameters assimilated according to the set variational data are used to obtain updated policy network parameters based on a deep reinforcement learning method, specifically including: Obtaining updated policy network parameters based on a deep proximal policy optimization algorithm according to the customized environment, policy network parameters, and value network parameters assimilated by the ensemble variational data; Determining the mixed background-error covariance matrix according to the updated strategy network parameters specifically includes: According to the updated strategy network parameters, use formula B h =(1-β)B s +βB e , determine the mixed background-error covariance matrix; Among them, B h is the mixed background-error covariance matrix, B e is the static covariance matrix, B s is the collective covariance matrix, β is the updated policy network parameter, which controls B e and B s The weight of , 0<β<1.

2. The numerical weather forecasting method based on deep reinforcement learning according to claim 1, characterized in that: The deep proximal policy optimization algorithm specifically includes an actor network and a critic network. The actor network is responsible for generating actions and selecting appropriate parameter ratios in the hybrid assimilation process; the critic network estimates the value function to evaluate the effectiveness of the current policy.

3. A numerical weather forecasting system based on deep reinforcement learning, characterized in that: The system comprises: Custom environment setting module, used to set custom environment for ensemble variational data assimilation; Parameter acquisition module, used to obtain policy network parameters and value network parameters; A policy network parameter updating module is used to obtain updated policy network parameters based on the custom environment, policy network parameters and value network parameters assimilated by the set variational data based on a deep reinforcement learning method; A mixed background-error covariance matrix determination module is used to determine a mixed background-error covariance matrix according to the updated policy network parameters, wherein the mixed background-error covariance matrix is ​​a weighted average of a static covariance matrix and a collective covariance matrix; A numerical weather forecast prediction model training module is used to update the adaptive hybrid parameter strategy model training according to the hybrid background-error covariance matrix to obtain a numerical weather forecast prediction model; A numerical weather forecast prediction module, configured to perform numerical weather forecast prediction based on the numerical weather forecast prediction model; The custom environment setting module specifically includes: The reward function determination unit is used to customize the reinforcement learning simulation environment according to the ensemble variational data assimilation system assimilation prediction cycle process, design the state space and action space of the intelligent agent, and establish a decision-feedback reward mechanism to obtain the reward function at the current moment; An action and state determination unit, used to obtain the action and state at the current moment; a custom environment determination unit, configured to assimilate the reward function, action, and state as a custom environment for collective variational data; The policy network parameter updating module specifically includes: A policy network parameter updating unit, configured to obtain updated policy network parameters based on a deep proximal policy optimization algorithm according to the custom environment, policy network parameters, and value network parameters assimilated by the ensemble variational data; The mixed background-error covariance matrix determination module specifically includes: Mixed background-error covariance matrix determination unit, used to adopt formula B according to the updated strategy network parameters h =(1-β)B s +βB e , determine the mixed background-error covariance matrix; Among them, B h is the mixed background-error covariance matrix, B e is the static covariance matrix, B s is the collective covariance matrix, β is the updated policy network parameter, which controls B e and B s The weight of , 0<β<1.

4. An electronic device, characterized in that: The electronic device includes a memory and a processor, the memory is used to store a computer program, and the processor runs the computer program to enable the electronic device to execute the numerical weather forecasting method based on deep reinforcement learning as described in any one of claims 1-2.

5. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the numerical weather forecasting method based on deep reinforcement learning according to any one of claims 1-2.

Citation Information

Patent Citations

  • Shared bicycle demand prediction method based on multi-strategy improved GWOBP neural network

    CN112766533A

  • Method for scene modeling and change detection

    US20050286764A1

Cited By

  • Method and system for estimating numerical prediction error growth rate of complex terrain

    CN122712084A