Numerical weather forecasting method and system based on deep reinforcement learning

Through a method based on deep reinforcement learning, the policy network parameters are updated and the mixed background-error covariance matrix is ​​determined, which solves the problem of insufficient design of background error covariance matrix in weather forecasts, and improves prediction accuracy and assimilation system performance.

CN119940410AActive Publication Date: 2025-05-06NAT UNIV OF DEFENSE TECH
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510007362.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-03
Publication Date
2025-05-06
Estimated Expiration
2045-01-03

AI Technical Summary

Technical Problem

The prior art is difficult to accurately count the distribution information of real background errors in weather forecasts, resulting in insufficient design and optimization of background error covariance matrix, especially when observation data is scarce and noise is present.

Method used

The numerical weather forecast prediction method based on deep reinforcement learning is adopted. By setting a custom environment for assimilation of set variational data, the policy network parameters and value network parameters are obtained, the policy network parameters are updated, the mixed background-error covariance matrix is ​​determined, and the adaptive mixed parameter strategy model is updated to obtain the numerical weather forecast prediction model.

Benefits of technology

The prediction accuracy of numerical weather forecasts is improved, especially in the case of observation of scarcity and weather evolution mutation periods, the background error information can be more accurately reflected and the performance of the assimilation system can be improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119940410A_ABST
    Figure CN119940410A_ABST
Patent Text Reader

Abstract

The invention discloses a numerical weather forecast prediction method and system based on deep reinforcement learning. The method comprises the following steps: setting a self-defined environment for assimilation of the set variation data; acquiring a policy network parameter and a value network parameter; obtaining updated strategy network parameters based on a deep reinforcement learning method according to the self-defined environment, the strategy network parameters and the value network parameters assimilated by the set variation data; according to the updated strategy network parameters, a mixed background-error covariance matrix is determined, and the mixed background-error covariance matrix is a weighted average value of a static covariance matrix and a set covariance matrix; updating self-adaptive mixed parameter strategy model training according to the mixed background-error covariance matrix to obtain a numerical weather forecasting model; and performing numerical weather forecast prediction according to the numerical weather forecast prediction model. According to the invention, the prediction precision of numerical weather prediction can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of weather forecast technology, and in particular to a numerical weather forecast method and system based on deep reinforcement learning. Background Art

[0002] Data assimilation was first proposed to solve the need for initial values ​​in numerical model integration. It aims to solve the inverse problem in numerical forecast models. Its core idea is to define an atmospheric motion state as accurately as possible by making full use of all existing information. In the data assimilation system, the short-term forecast results of the numerical model are usually used as background field information. However, the numerical model is essentially an approximate estimate of the real weather system through a large amount of prior knowledge, so there are systematic errors. This requires the introduction of real observation data in the data assimilation analysis cycle to correct the background field, so as to provide the optimal initial field for the model forecast at the next moment.

[0003] Although in recent years, with the increasing demand for global weather forecasts, countries have established relatively complete comprehensive meteorological observation networks including ground-based, air-based and space-based, however, in the actual operational data assimilation process, due to the constraints of various objective and subjective factors, the amount of global multi-source observation data available at a specific moment is far less than the dimension of the numerical forecast model state, resulting in the dominant position of the model background field in the assimilation process. ECMWF has conducted a quantitative assessment of the contribution of various types of information in the analysis field. The results show that of the analysis field information obtained through the data assimilation system, about 15% comes from the observation data during the assimilation period, and the remaining 85% comes from the model background field.

[0004] Nevertheless, the key role of the background error covariance matrix B in data assimilation, especially in the characterization of background field uncertainty, can still ensure that a small amount of observation data is effectively propagated through the matrix. Specifically, the background error covariance matrix can not only effectively propagate observation information, but also smooth observation increments, introduce equilibrium properties, and construct flow structure, thereby ensuring the smooth progress of the assimilation process. The role of this mechanism is particularly significant, especially when observation data are scarce and noisy. The design and optimization of the background error covariance matrix B determines the overall performance of the data assimilation system to a certain extent.

[0005] As a statistic, the background error covariance matrix B describes the probability distribution function of the model prediction error (i.e., the background field error in the data assimilation system). However, in the actual modeling process, the core problem faced is that the "real" state cannot be obtained, so it is impossible to accurately count the distribution information of the real background error. In order to make up for this deficiency, the background error can usually only be simulated under specific assumptions. The common simulation method in business is the NMC (National Meteorological Center) method proposed by Parrish and Derber, which simulates the prediction error by comparing the integral difference between the model forecast values ​​at the same time and different prediction times. Because the NMC method is concise in theory and simple in business implementation, it is not limited by the resolution of the observation network, and can count the effective background error covariance matrix B containing the balance and constraint relationship between different variables. Therefore, it is widely used in the global meteorological assimilation system and has been adopted by most numerical prediction centers.

[0006] However, the statistical samples of background errors in the NMC method are usually calculated based on the model prediction differences over a considerable period of time (e.g., one year). In this method, the error changes are mainly due to the natural variability of the model, so the resulting background error covariance matrix B is climatological, static, and isotropic. However, the real background error covariance matrix B usually has complex properties such as significant anisotropy, flow dependence, and baroclinicity, which are not fully reflected in the NMC method. To this end, the researchers proposed an ensemble-variational hybrid assimilation method, which aims to generate short-term forecast ensemble samples in real time by combining the Monte Carlo method (ensemble method) in the data assimilation cycle, and obtain the forecast error covariance Be with flow-related characteristics through these sample statistics. This method mixes the real-time generated background error covariance matrix Be with the static background error covariance matrix Bs obtained by the NMC method through linear combination, thereby more accurately reflecting the error characteristics that change with the weather situation.

[0007] There are many ways to obtain prediction ensemble samples, the most common of which is the Ensemble Kalman Filter (EnKF) method. In this method, a set is generated simultaneously in a certain number of data assimilation cycles through reasonable perturbations (including initial value perturbations, observation perturbations, and parameterized perturbations). The resulting set of prediction data at the same time can be used to estimate the prediction error covariance. However, as the number of ensemble members increases, the computational cost increases exponentially. At present, the number of ensemble members that the operational numerical forecast center can bear is about 30 to 100, which is much smaller than the dimension of the numerical model state variables, resulting in the estimated forecast error covariance Be is usually non-full rank, has significant sampling errors, and is prone to false long-distance correlations and other problems.

[0008] An effective and low-cost approach is to increase the number of EnKF ensemble members through the effective time-shifted ensemble members (VTS). VTS expands the ensemble size by adding subset members before and after the central analysis time, including VTSM for sampling time and / or phase errors, and VTSP for eliminating spurious covariance through time smoothing. Another efficient alternative is the time-lagged ensemble forecast method, which directly uses the forecast results of the same time at different initial times in the historical forecast field of the variational assimilation system as ensemble samples, greatly reducing the computational cost and storage cost. The uncertainty of the time-lagged ensemble mainly comes from the differences in the initial field, lateral boundaries and observational data at different times, so it can effectively reflect the covariance information of the flow-dependent forecast error that evolves over time.

[0009] However, since the credibility of the time-lagged ensemble members is closely related to the forecast timeliness, the forecast results with longer timeliness often deviate from the true value and cannot be used as valid ensemble members. Therefore, the number of samples in the time-lagged ensemble is usually limited. To supplement the number of time-lagged ensemble samples, the empirical orthogonal decomposition (EOF) technique can be used to select prediction samples with similar change characteristics of meteorological elements in the target assimilation period and region from a wide range of historical samples, which is called the preferred historical forecast sample. This method can effectively increase the number of time-lagged ensemble samples at a lower computational cost, thereby introducing flow-related features that more accurately reflect the current weather conditions.

[0010] Although a large number of studies have shown that the assimilation performance of the ensemble-variation hybrid assimilation method is better than that of a single variational or ensemble method, a common shortcoming is that the combination of the climatological background error covariance matrix B and the dynamically changing forecast error covariance Be is usually performed in a linear weighted manner, and the hybrid weight is often a fixed empirical parameter (usually less than 1). This processing method has two significant defects: first, the utilization efficiency of the data samples used to estimate the error covariance is low, and second, the selection of the hybrid weight lacks a scientific basis and cannot flexibly adapt to the changes in different weather conditions. Therefore, a reasonable improvement idea is that the hybrid parameters should be adaptively adjusted according to the inherent characteristics of the data samples that provide background error information as the weather situation changes in space and time, which is also one of the core issues of this study.

[0011] In terms of data feature extraction, the traditional field of earth science usually integrates known mathematical and physical mechanisms, builds numerical simulation models based on ideal assumptions, and uses supercomputers to simulate and predict relevant weather or climate phenomena. This method is undoubtedly limited by the limitations of human understanding and cognition of earth system processes, and therefore inevitably incorporates too much human intervention. Unlike this theory-driven paradigm, data-driven machine learning methods, especially deep learning, have flourished in recent years. Deep learning can effectively combine low-level features of data through multi-layer neural network structures and nonlinear transformations to form abstract and easy-to-distinguish high-level representations, thereby more comprehensively perceiving the distributed features of data and migrating them to various downstream tasks, such as prediction, classification, and point cloud registration. End-to-end models represented by large language models rely on massive databases for training and optimization through supervised or self-supervised learning to achieve natural language understanding and generation. This successful paradigm has brought huge potential to the field of "AI for Science" and has spawned a number of large meteorological models based on deep learning, which are completely independent of traditional numerical forecasting models and have created a new scientific research paradigm. Although these models have shown great advantages in improving the accuracy of global medium-term forecasts and accelerating the prediction reasoning process, the end-to-end intelligent forecasting results are usually smooth and have poor prediction effects on convective-scale weather with turning characteristics.

[0012] In operational numerical forecast simulation, since the physical mechanism is not yet fully understood, researchers often have to rely on a large number of empirical parameters to simulate the evolution of actual weather phenomena as much as possible while ensuring the availability of forecasts. These empirical parameters include key parameters in different weather processes, key settings in various parameterization schemes, and hybrid parameters in hybrid assimilation. This approach relies more on researchers' professional experience with weather phenomena, and lacks a more scientific theoretical basis.

[0013] Excitingly, reinforcement learning in the field of machine learning provides new ideas for this challenge. Reinforcement learning is a learning method from environmental state to action selection, whose goal is to enable the agent to obtain the maximum cumulative reward in the interaction with the environment and ultimately obtain the optimal strategy. This method of selecting actions based on prior states can alleviate the "interpretability" problem caused by the "black box" characteristics of deep learning models. Although a large number of studies are still trying to incorporate physical information into deep learning models, these methods are collectively referred to as physical information neural networks (PINN). Such methods customize network models by adopting different activation functions, gradient optimization techniques, neural network structures, and loss functions in order to improve the physical consistency of the model.

[0014] However, the limitations of PINN are also very obvious. It can only effectively solve problems with known physical models. For unclear research problems (such as complex precipitation processes), mechanically adding physical information may introduce more errors and even have a negative impact on the performance of the neural network. Therefore, although PINN has made progress in some areas, there is still a lot of room for improvement in unresolved or unknown theoretical problems. Summary of the invention

[0015] The present invention provides a numerical weather forecast prediction method and system based on deep reinforcement learning, which can improve the prediction accuracy of numerical weather forecast.

[0016] To achieve the above object, the present invention provides the following solutions:

[0017] A numerical weather forecasting method based on deep reinforcement learning includes:

[0018] Set up a custom environment for ensemble variational data assimilation;

[0019] Get strategy network parameters and value network parameters;

[0020] Obtain updated policy network parameters based on the custom environment, policy network parameters, and value network parameters assimilated by the ensemble variational data based on a deep reinforcement learning method;

[0021] Determine a mixed background-error covariance matrix according to the updated policy network parameters, wherein the mixed background-error covariance matrix is ​​a weighted average of a static covariance matrix and a set covariance matrix;

[0022] Updating the adaptive hybrid parameter strategy model training according to the hybrid background-error covariance matrix to obtain a numerical weather forecast prediction model;

[0023] Numerical weather forecast prediction is performed according to the numerical weather forecast prediction model.

[0024] To achieve the above object, the present invention also provides the following solution:

[0025] A numerical weather forecasting system based on deep reinforcement learning includes:

[0026] Custom environment setting module, used to set custom environment for ensemble variational data assimilation;

[0027] Parameter acquisition module, used to obtain policy network parameters and value network parameters;

[0028] A policy network parameter updating module, used to obtain updated policy network parameters based on a deep reinforcement learning method according to the custom environment, policy network parameters and value network parameters assimilated by the set variational data;

[0029] A mixed background-error covariance matrix determination module, used to determine a mixed background-error covariance matrix according to the updated policy network parameters, wherein the mixed background-error covariance matrix is ​​a weighted average of a static covariance matrix and a set covariance matrix;

[0030] A numerical weather forecast prediction model training module is used to update the adaptive hybrid parameter strategy model training according to the hybrid background-error covariance matrix to obtain a numerical weather forecast prediction model;

[0031] The numerical weather forecast prediction module is used to perform numerical weather forecast prediction according to the numerical weather forecast prediction model.

[0032] To achieve the above object, the present invention also provides the following solution:

[0033] An electronic device includes a memory and a processor, wherein the memory is used to store a computer program, and the processor runs the computer program to enable the electronic device to perform a numerical weather forecasting method based on deep reinforcement learning.

[0034] To achieve the above object, the present invention also provides the following solution:

[0035] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements a numerical weather forecasting method based on deep reinforcement learning.

[0036] According to the specific embodiments provided by the present invention, the present invention discloses the following technical effects:

[0037] The present invention provides a numerical weather forecast prediction method based on deep reinforcement learning, the method comprising: setting a custom environment for ensemble variational data assimilation; obtaining policy network parameters and value network parameters; obtaining updated policy network parameters based on the custom environment, policy network parameters and value network parameters for ensemble variational data assimilation based on a deep reinforcement learning method; determining a mixed background-error covariance matrix based on the updated policy network parameters, the mixed background-error covariance matrix being a weighted average of a static covariance matrix and a ensemble covariance matrix; updating adaptive mixed parameter policy model training based on the mixed background-error covariance matrix to obtain a numerical weather forecast prediction model; and performing numerical weather forecast prediction based on the numerical weather forecast prediction model. The present invention can improve the prediction accuracy of numerical weather forecasts. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.

[0039] Figure 1 This is a flow chart of the numerical weather forecasting method based on deep reinforcement learning of the present invention;

[0040] Figure 2 This is the DRL-EnVar framework diagram;

[0041] Figure 3 This is an implementation example of C-CNN;

[0042] Figure 4 This is a system structure diagram of the numerical weather forecast based on deep reinforcement learning in the present invention. DETAILED DESCRIPTION

[0043] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0044] The present invention provides a numerical weather forecast prediction method and system based on deep reinforcement learning, which can improve the prediction accuracy of numerical weather forecast.

[0045] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.

[0046] Embodiment 1:

[0047] Figure 1 This is a flow chart of the numerical weather forecasting method based on deep reinforcement learning of the present invention, such as Figure 1 As shown, the present invention provides a numerical weather forecasting method based on deep reinforcement learning, the method comprising:

[0048] Step 101: Set up a custom environment for ensemble variational data assimilation.

[0049] The goal of the EnVar assimilation system is to use error information from two sources of information (B hand R), effectively integrating the background field and the observation field, providing the optimal analysis field at the target time (time i) for the numerical prediction model, so that the numerical prediction model uses the analysis field as the initial field (in order to more intuitively evaluate the pros and cons of the analysis field, the present invention defaults to the assimilation system based on the Lorenz96 model and can directly perform subsequent model forecasts after ignoring the initialization). After a period of numerical forecasting (time i), the numerical model has the best forecasting performance, that is, the forecast result is close to the real field, that is: argmin||x f -x t || I , where x t The real field at the target moment; is the analysis field at time i For numerical mode The initial field of the forecast time is (Ii), and the model forecast field at the Ith moment is obtained. Among them, Weight in It is represented by a variety of static background error covariances (subscripted as static, abbreviated as st) and mixed background error covariances of flow-dependent background error information (subscripted as flow-dependent, abbreviated as fl), and the parameters satisfy α1+α2+...+α m +β1+β2+...+β n =1.

[0050] The research background of this invention is that under the premise of sparse observation (masked) and noise, the continuous assimilation forecast cycle is one year and there is a mutation period (one month) during the period. The variable value during the mutation period is significantly higher than the variable value of the climate state, simulating strong convective weather conditions such as heavy precipitation. In the above example, two statistical results R and H based on the same and unchanged observation data are completely known and fixed, then B h The accuracy of the description of the background error information of the real-time state variable and the degree of timely response to the sudden change, that is, the degree of characterization of the background error flow dependency information, determines the accuracy of the analysis field. h It mainly depends on the static background error covariance obtained by statistics of historical data samples and the collective background error covariance with flow-dependent properties calculated by collective data samples, as well as the appropriate selection of fusion parameters α and β of multiple background error information sources (such as model Figure 2 (part (c)).

[0051] In order to restore the background error information closer to the real one as much as possible, the present invention will make full use of all data samples that can provide effective background error information, including multiple historical climate data samples and real-time ensemble forecast samples, to calculate the background error information without significantly increasing the statistical calculation cost. Under this premise, the present invention mainly explores the impact of the intelligent selection of mixing parameters of different background error information sources on the assimilation performance of the system, emphasizing the real-time and adaptability of the mixing parameters.

[0052] According to the above description, the task of selecting the mixing matrix parameters in real time in the hybrid assimilation forecast cycle can be expressed as a decision-making task that maps from the current state (analysis field) to the action, namely:

[0053] f(s i )=a i

[0054] a i =[α1,...,α m ,β1,...,β n ]

[0055] Among them, m represents the number of static Bs, and n represents the number of flow-dependent Bs.

[0056] It is formalized as an MDP as follows:

[0057] ·state is the vector given by the EnVar system at time step i.

[0058] Action a i is a set of different weighting factors with the constraint that their sum is equal to 1.

[0059] State transfers i+1 is the numerical solution that minimizes the cost function in our task:

[0060]

[0061] Reward R(s i , a i ,s i+1 is in state s i Take action i and arrive at a new state s i+1 direct reward.

[0062] The discount factor γ is set to 0.99 because our task involves long-term sequential planning and the future rewards from sequential actions are crucial.

[0063] The goal is to trust and retain the complete framework of the EnVar assimilation system, and to more finely characterize B through deep reinforcement learning. hEmpowerment enables it to effectively transmit the flow-related information of the model evolution, improve the assimilation performance of the EnVar assimilation system, especially during the mutation period, to respond in time and make effective adjustments, and to obtain a B that can better reflect the background error information of the mutation period in real time. h , which is crucial for obtaining better analysis fields in subsequent assimilation forecast cycles. Specifically, given the initial field of the numerical model, based on available known observations, using the EnVar system for assimilation forecast cycles, it is necessary to find an optimal strategy π for adaptively giving the mixed weights of multiple background error information sources. θ , so that the assimilation forecast cycle can operate stably regardless of the stable or sudden change period of the climate state of the numerical model evolution, and make the assimilation performance as high as possible.

[0064] The attempts and researches of deep reinforcement learning in the field of mixed data assimilation are relatively few compared to other fields such as robot control and autonomous driving. The present invention does not simply use the classic MLP-based deep reinforcement learning neural network (many model parameters and high training cost), but designs an innovative and intelligent algorithm called DRL-EnVar based on the experimentally verified numerical forecast model and the characteristics of the assimilation forecast cycle system. The algorithm model needs to reduce the computational complexity as much as possible, reduce the neural network parameters to control the model training time, and still extract effective information from the limited available data. The intelligent agent interacts with the information generated by the environment, completes reasoning, and learns to obtain the optimal strategy for adaptively selecting mixed parameters. Especially when there is a mutation period in the evolution of the numerical model, the strategy can still be used stably to ensure high assimilation performance.

[0065] The DRL-EnVar algorithm is an intelligent hybrid assimilation strategy algorithm proposed in the present invention, which aims to solve the real-time adaptive hybrid parameter selection problem in the EnVar hybrid assimilation forecast cycle. In this cycle, the mixing ratio of the static background error covariance matrix B and the flow-dependent background error covariance matrix (flow-dependent B) needs to be dynamically adjusted so that the algorithm can effectively respond to changing observation and forecasting scenarios. This task can be regarded as a decision-making problem of continuously selecting the optimal action in a high-dimensional state space, which conforms to the classic decision-making problem type in reinforcement learning. First, define the EnVar custom environment.

[0066]

[0067]

[0068] The above algorithm is the definition process of EnVar custom environment. First, the initial background field of Lorenz96 model is given That is X j =F, if j≠20; X j =1.001F,if j=20.

[0069] The first state of the environment has not yet been assimilated and is directly determined by Assign to As the initial state of the environment. According to the above analysis, the present invention provides a total of 5 B matrices of background error information, so that the action space agent needs to select 5 mixed parameters to perform an action, and the sum of the parameters is always 1. The reward function consists of two parts. One part is to analyze the field x after each assimilation. a With real field x t The distance, that is, the root mean square error RMSE a The other part is to directly measure the quality of the analysis field after assimilation, and use the assimilated analysis field as the initial field to make a 48h forecast. RMSE with the true field f To evaluate the forecasting performance of the numerical forecasting model, and thus indirectly measure the quality of the current analysis field. The larger the two root mean square errors are, the greater the penalty for the actions taken by the agent, that is, R = -RMSE a -RMSE f .

[0070] Regarding the setting of state transition rules, the research background of the present invention is set as a continuous assimilation forecast cycle of one year (assuming that each month is 30 days, then one year is 360 days). In addition, in order to keep the Loren96 model stable and eliminate transient behavior, it is set to discard 90 days of Loren96 model initial startup data, so a total of 450 days of model continuous integration is required. The numerical solution method uses the fourth-order Runge-Kutta format (RK4), with a time step of dt=0.05 (agreed to be 6 hours), that is, 4 assimilations need to be performed in 1 day, and the total number of iterations is 1800 times (the pseudo code line number of Algorithm 3.1 is 5).

[0071] In order to study whether the B matrix can transmit flow-related information in time during the mutation period, a switch is set in the state transition rule. If the mode evolution mutation index is detected, the forcing term of the numerical model is set to 15.0, otherwise it remains at 8.0. In addition, due to the flow-dependent B 48 and B 24 A certain number of collective samples need to be accumulated to perform statistical calculations, so Figure 2 (c) It takes 8 assimilations (num_DA = 8) before a flow-dependent B is performed. 48 and B 24 , thereby updating B h .

[0072] If num_DA<8, based on repeated experimental analysis, directly use the climate state background error information that meets the experimental settings, that is, BF8S15 As B h The EnVar system assimilation performance is relatively the best. Based on the above definition, the EnVar custom environment finally returns the state s i , action a i , reward r i , and provided to the DRL-EnVar algorithm for subsequent training.

[0073] This step specifically includes:

[0074] According to the assimilation forecast cycle process of the ensemble variational data assimilation system, the reinforcement learning simulation environment is customized, the state space and action space of the intelligent agent are designed, and a decision-feedback reward mechanism is established to obtain the reward function at the current moment;

[0075] Get the current action and status;

[0076] A custom environment that assimilates the reward function, actions, and states as collective variational data.

[0077] The EnVar hybrid scheme of the ensemble variational data assimilation system is built on the basis of the existing 3DVar system and directly introduces the integrated information through the error covariance matrix.

[0078]

[0079] Among them, x represents the state variables, including model variables such as temperature, wind force, pressure and humidity; x b Indicates the background state, y o Represents observation data, all data are three-dimensional data. Assume that the nonlinear observation operator is linear and completely known, simplifying to R is the observation error covariance matrix, which can be obtained from the instrument observation error to describe the observation value y o Correlation with the model grid points.

[0080] It is worth noting that the mixed background-error covariance matrix B h is defined as the static covariance matrix and the aggregate covariance matrix (B e and B s ), effectively replacing B in the original 3DVar system.

[0081] B h =(1-β)B s +βB e

[0082] Among them, β (0<β<1) is an adjustable factor that controls B e and B s The weight of .

[0083] Step 102: Obtain strategy network parameters and value network parameters.

[0084] Step 103: Obtain updated policy network parameters based on the custom environment, policy network parameters, and value network parameters assimilated by the ensemble variational data based on a deep reinforcement learning method.

[0085] Deep learning is a subset of machine learning and a general artificial intelligence method that uses multi-layer neural networks to simulate complex patterns and representations of large-scale data sets. Deep learning models are usually composed of multiple layers of nonlinear operation units. It uses the output of the lower layer as the input of the higher layer. In this way, it automatically learns high-level abstract feature representations from a large amount of training data to fully perceive the distributed characteristics of the data, and has shown excellent performance in various tasks such as image and speech recognition and natural language processing. Among the many architectures that have emerged, the gated recurrent unit (GRU) in the fully connected neural network (FCNN), convolutional neural network (CNN) and recurrent neural network (RNN) has attracted great attention due to its effectiveness in processing specific types of data and tasks.

[0086] Fully Connected Neural Networks (FCNNs), also known as dense neural networks, are a class of artificial neural networks with extensive connectivity. Fully connected neural networks represent the most basic form of deep learning, where every neuron in one layer is connected to every neuron in the next layer. This dense connectivity enables FCNNs to model complex nonlinear relationships in data, making them suitable for a wide range of applications, including classification, regression, and pattern recognition. Multi-Layer Perceptron (MLP) is a special and widely used type of FCNN.

[0087] Convolutional Neural Networks (CNNs) are specifically designed to process grid-like data structures, such as images. CNNs are inspired by biological processes in the visual cortex, where neurons are arranged in a way that covers the visual field, using convolutional layers to automatically and adaptively learn a spatial hierarchy of features from the input image. These layers apply convolution operations to capture local patterns, such as edges and textures, which are then aggregated in deeper layers to form more complex structures and objects. This hierarchical feature extraction has enabled CNNs to achieve remarkable performance in image recognition, object detection, and related tasks.

[0088] Recurrent Neural Networks (RNNs) are designed to process patterns in sequences of data, such as time series or natural language. The core component of RNNs is the recurrent unit, which processes input data sequentially, updating its internal state based on the current input and previous state. This feedback mechanism allows the network to capture temporal dependencies and long-term correlations in the data. However, standard RNNs face challenges such as vanishing and exploding gradients, which can hinder their ability to learn long-term dependencies. To address this issue, the Gated Recurrent Unit (GRU) was introduced as a variant of the RNN architecture. As a simplified variant of the Long Short-Term Memory (LSTM) network, the GRU combines the functionality of the input gate and the forget gate into a single update gate, simplifying the architecture while maintaining similar performance. By using update and reset gates, the GRU effectively manages the information passing through the network, resulting in better performance in tasks such as time series prediction, language modeling, and speech recognition.

[0089] Reinforcement learning (RL) is a subfield of machine learning (ML) that focuses on training agents to make sequential decisions by interacting with the environment to maximize cumulative rewards. The framework is based on Markov decision processes (MDPs), which are characterized by a tuple (S, A, P, R, γ), where S is the state space, A is the action space, P represents the state transition probability, R is the reward function, and γ is the discount factor.

[0090] RL problems are mainly solved by value-based and policy-based methods. Value-based methods, such as Q-learning and DQN, derive optimal policies by optimizing value functions, and are applicable to discrete environments such as Go. Policy-based methods, such as policy gradients, gradually improve policies, making them applicable to continuous action scenarios such as robot control. A well-known method in RL is the Actor-Critic framework, which combines these two methods to solve problems in continuous action spaces and high-dimensional state spaces. The framework uses two networks: actor-network, which is used to generate parameterized action policies; critic-network, which is used to evaluate the value of state-action pairs. This dual structure combines value function approximation with direct policy optimization, improving learning efficiency and stability.

[0091] Proximal Policy Optimization (PPO) is an advanced RL algorithm that improves the Actor-Critic framework by introducing a clipped proxy objective function to achieve stable policy updates. PPO solves the high variance and instability problems of traditional policy gradients and achieves a balance between simplicity and performance, making it the preferred method for complex RL tasks.

[0092] During this period, in order to better extract the abstract representation of the cyclic data, we also proposed a cyclic one-dimensional convolution module (C-CNN) to remove the symmetry in the data and improve the learning efficiency. Considering the particularity of this research task, a softmax module is set so that the sum of the mixed parameters is always 1. Cyclic-Convolutional neural network (C-CNN): C-CNN is a special convolution structure designed for the data of the earth system. The numerical prediction model used in the experimental verification is Lorenz96, which is characterized by J variables on equidistant grid points around the equator, that is, its data structure is a ring structure connected head to tail. According to the circulation situation of the atmosphere, the variables of adjacent longitudes in the same latitude circle affect each other. C-CNN is inspired by one-dimensional convolution. As shown in the figure, considering the physical space dependency of the observations, if ordinary one-dimensional convolution calculation is used, the physical space information of the ring cannot be modeled, resulting in a decrease in the modeling ability of the model. Therefore, we use a symmetrical cyclic convolution in the physical space to replace the one-dimensional convolution. After multiple rounds of circular convolution, the model can extract higher-level abstract features, and this feature will not be affected by the symmetry of the physical space (that is, when the data is initialized, no matter where the starting point of the ring is, the same representation can always be obtained in the end), so that the next step of processing can be done in a smaller representation space, so that the model can achieve better performance on less data. In addition, after each circular convolution, the nonlinear features of the neural network are added using the RELU (Rectified Linear Unit) activation function to help the network better learn data features. That is, the input features are transformed element by element nonlinearly. On the one hand, for negative inputs, the ReLU function outputs 0, which is equivalent to setting some neurons to an inactive state, so that the neural network has a certain sparsity and reduces the correlation between parameters; on the other hand, for positive intervals, the ReLU function keeps the gradient at 1, which can effectively alleviate the problem of gradient disappearance. In addition, combined with the BatchNorm operation, it helps to make the model training more stable and accelerate convergence.

[0093] Figure 3This is an implementation example of C-CNN. The number of input channels of C-CNN is 1, the number of convolution kernels is 8, the size of each convolution kernel is 5, and the number of output channels is 8. Starting from any starting point of the ring data, each convolution kernel moves counterclockwise along the ring data. Every time it moves to a position, the values ​​of the corresponding position are multiplied and summed. For example, a convolution process is: the yellow dotted ellipse in the input channel data is multiplied by 8 convolution kernels and then summed, and 8 values ​​are output to form 8 output channel data (the corresponding 8 yellow dotted rectangular boxes in the output channel data). Then the convolution kernel moves counterclockwise 1 step to the blue dotted box data position for convolution operation, then to orange and then to green, and so on, until the end is connected to complete the complete ring convolution operation. Among them, the value of each convolution kernel is different, and it is learned and updated together with the policy network model parameters during model training.

[0094] Step 103 specifically includes:

[0095] According to the customized environment, policy network parameters and value network parameters of the ensemble variational data assimilation, based on a deep proximal policy optimization algorithm, updated policy network parameters are obtained.

[0096]

[0097]

[0098] The above algorithm describes the training process of the DRL-EnVar algorithm that implements adaptive mixing parameter selection.

[0099] Based on extensive research and comparison of deep reinforcement learning (DRL) methods, it is found that PPO (Proximal Policy Optimization) shows superior performance and robustness in current DRL algorithms. Compared with other methods (such as DQN, TRPO, DDPG, etc.), PPO has the advantages of low parameter sensitivity, stable update, and wide applicability, especially for optimization problems in continuous space, which makes it the preferred algorithm in the research and application of reinforcement learning. Therefore, this study uses PPO as a reinforcement learning algorithm to ensure training convergence and model stability in high-dimensional complex environments, thereby improving the effect of adaptive hybrid parameter selection.

[0100] The core of the DRL-EnVar algorithm consists of two parts: Actor and Critic, forming a typical Actor-Critic structure. The Actor network is responsible for generating actions, that is, selecting the appropriate parameter ratio in the hybrid assimilation process; the Critic network estimates the value function to evaluate the effectiveness of the current strategy. Both use the MLP (Multi-layer Perceptron) structure. Based on previous experience in data assimilation and forecasting tasks, MLP has been verified as a universal approximator and has performed well in the spatiotemporal series assimilation-forecasting cycle.

[0101] Specifically, the PPO algorithm introduces the Clip operation in the update of the Actor, and controls the update amplitude of the policy gradient by limiting the range of the ratio of the new and old policies, thereby avoiding training instability caused by excessive policy updates. The clipping range ε ​​is taken as 0.2, which is an empirical value. The Critic network is updated by the temporal difference (TD) error, and the estimation accuracy of the value function is improved by minimizing the TD error. In addition, in order to fully extract data features, this algorithm combines C-CNN and MLP structures in the deep learning network. C-CNN is used to process the extraction of spatiotemporal features, and MLP is used for the estimation of strategies and values. This structure ensures that in the complex EnVar assimilation-forecasting cycle, the model can effectively capture data features and select the optimal mixing parameters. Finally, the trained DRL-EnVar model can dynamically adjust the proportion of mixing parameters during the real-time assimilation process, thereby achieving the best assimilation effect in different meteorological and observation scenarios.

[0102] Step 104: Determine a mixed background-error covariance matrix according to the updated policy network parameters, where the mixed background-error covariance matrix is ​​a weighted average of a static covariance matrix and a set covariance matrix.

[0103] This step specifically includes:

[0104] According to the updated strategy network parameters, formula B is used h =(1-β)B s +βB e , determine the mixed background-error covariance matrix;

[0105] Among them, B h is the mixed background-error covariance matrix, B e is the static covariance matrix, B s is the set covariance matrix, β is the updated policy network parameter, which controls B e and B s The weight of , 0<β<1.

[0106] Both static B and B with flow-dependent characteristics are calculated by statistical methods, but the methods used are different. First, static B is calculated by the most commonly used NMC method in business, that is, using the difference between forecast pairs of different lengths, but each forecast pair is valid at the same time. The formula is:

[0107]

[0108] It can be seen that the key to the NMC method lies in the composition of the historical data samples (3 years) used for statistical calculations. This study uses three historical data samples, which respectively reflect the different climate state characteristics of the numerical model. The first data sample fully follows the research background of the present invention, that is, a mutation period of one month occurs every year, during which the forcing term of Lorenz96 F = 15.0, and the remaining 11 months are the conventional stable period F = 8.0; the second data sample assumes that the whole year is a mutation period, the purpose is to more prominently focus on the error correlation between state variables when F = 15.0, and it is expected that when studying the mutation during the EnVar assimilation forecast cycle, this data sample can provide significant relevant error information; the third data sample assumes that no mutation occurs throughout the year, that is, F is always 8.0, because the mutation period is lower in frequency than the conventional stable change of the model, that is, the model is relatively slowly evolving, that is, stable, more often in a year, and the background error information of this climate state is quite important.

[0109] Regarding the real-time B that provides flow-dependent information, in order to avoid the unacceptable computational cost brought by the ensemble forecast and to ensure the effective statistics and transmission of the flow-dependent information as much as possible, the present invention uses two real-time ensemble data samples. Inspired in part by the acquisition of time-delay ensemble samples in variational assimilation, that is, the difference field of the model forecast field of different time periods at the same time in the historical forecast samples of the EnVar system, that is, the 24h forecast value and the 48h forecast value in the blue and green highlights on the left side of Figure (c);

[0110] The other part is inspired by the acquisition of time-shifted ensemble samples in ensemble assimilation, that is, considering that there will be a certain degree of spatiotemporal phase error in numerical model forecasts, especially severe convective weather phenomena with strong mutation such as heavy rainfall, which usually have time errors of large-value forecast advance or postponement and spatial deviation in large-value areas. In this way, the forecast values ​​close to a target moment in time and space are regarded as the forecast values ​​at the current moment, so as to alleviate the false alarm rate and effectively increase the number of ensemble samples, alleviate the sampling error caused by insufficient ensemble samples, that is, the 21h and 27h forecasts in the blue highlight of Figure (c) are regarded as 24h forecasts, and the 45h and 51h forecasts in the green highlight are regarded as 48h forecasts.

[0111] The two parts of samples together constitute two real-time set data samples. It should be noted in advance that the displacement prediction in this unrepresented space will be realized by the subsequent ring convolutional neural network, which will be described in detail later. Specifically, the present invention performs assimilation every 6 hours, that is, fusion of the observation field and the background field, and performs a 48-hour forecast. The forecast result is saved every 3 hours, and then B is calculated. 24 The number of members of the set is 12. According to the calculation principle of VTSP (reference), B is calculated. 24 . Accordingly, calculate B 48 The number of members in the set is 24. Statistics B 48 and B 24 The forecast timeliness of the two samples is significantly different, and the accumulated errors are different, which provides different flow-dependent error-related information for the hybrid assimilation system.

[0112] Step 105: updating the adaptive hybrid parameter strategy model training according to the hybrid background-error covariance matrix to obtain a numerical weather forecast prediction model;

[0113] Step 106: Perform numerical weather forecast prediction according to the numerical weather forecast prediction model.

[0114] Figure 2 The overall process of the DRL-EnVar algorithm is shown, which can be divided into three main stages: the adaptive selection of mixing parameters, the calculation of the mixed B matrix, and the h ) and EnVar hybrid system assimilation forecast cycle stage. The three stages are closely linked and influence each other to jointly realize the intelligent hybrid assimilation system.

[0115] In the DRL stage (a), we first customize the reinforcement learning simulation environment according to the assimilation forecast cycle process of the EnVar hybrid assimilation system, design the state space and action space of the agent, and establish a long-term "decision-feedback" reward mechanism. Through the agent's continuous interaction and exploration with the environment, and using the PPO proximal policy optimization algorithm to train the model, the parameters θ of the policy network are continuously updated. The goal is to maximize the expected cumulative reward and obtain a corresponding optimal strategy for selecting hybrid parameters that are adaptive to the current state. During this period, in order to better extract the abstract representation of the ring data, we also proposed a ring-shaped one-dimensional convolution module (C-CNN) to eliminate the symmetry in the data and improve learning efficiency. Considering the particularity of this research task, a softmax module is set so that the sum of the hybrid parameters is always 1.

[0116] Continue to Calculation B h Stage (c), based on the action a given in the DRL stage i ( Figure 2Arrow ①), the mixing parameters are multiplied by the corresponding background error covariance matrix, among which there are three static B (respectively B F8S15 , B F8 and B F15 ) and two B with flow-dependent characteristics calculated from real-time data samples (respectively B 48 and B 24 ).

[0117] In the EnVar stage, based on the B calculated in the previous stage h After completing an assimilation, the analysis field is passed to the DRL stage as the state part of the next agent collection information ( Figure 2 Arrow ③), conduct subsequent adaptive hybrid parameter strategy model training and update strategy model parameters. In addition, each assimilation is performed for 48h forecast, and the forecast result is saved every 3h. The iterative stacking is provided as the B dependent on the calculation flow. 48 and B 24 Real-time data samples.

[0118] The numerical weather forecast prediction method based on deep reinforcement learning is an intelligent ensemble variational hybrid data assimilation method based on deep reinforcement learning (DRL-EnVar) to improve the performance of the assimilation system under the conditions of sparse observations and important turning points in weather evolution. This method is based on the traditional EnVar framework and directly converts the B estimated by the ensemble sample into e Embedded into the variational cost function for iterative assimilation. In order to meet the needs of high-frequency assimilation, the time-lagged ensemble method is mainly used, which is highly time-efficient and supplemented by other efficient ensemble expansion strategies. By leveraging the robust feature extraction capability of DL and the intelligent decision-making capability of RL, DRL-EnVar significantly improves the assimilation performance and provides a physically consistent optimal initial state for the NWP model, thereby optimizing the forecast accuracy.

[0119] This experiment uses the Lorenz96 chaos model, which is widely used in the study of assimilation methods, as the numerical model. The initial field of the model is a set of state vectors with a given number of variables N. The true value field is generated by the evolution of the model under the integration step, and the forcing term of the model is set to F = 8. or F = 15.0. The short-term forecast field is also generated by the model integration, but in order to simulate the error characteristics of the model forecast, the forcing term of the Lorenz96 model is adjusted to F = 8.4 or F = 15.75 (reference) during the forecast. The assimilation framework is based on three-dimensional variational assimilation (3DVar), and the simulated observations are assimilated every 6 hours. The observations are generated by superimposing a small Gaussian random perturbation on the model true value field. The mean of the Gaussian noise is 0, and the covariance matrix is ​​the specified observation error covariance. The observation errors are assumed to be spatially independent and uncorrelated, that is, the observation error covariance matrix is ​​the proportional coefficient (I) of the unit matrix, and the observation error standard deviation is set to 1.0 by default. The background error covariance matrix uses a variety of mixed background error covariance schemes, and its specific settings are detailed in the comparative experimental scheme section. The verification period of the cyclic assimilation system experiment is set to one year, with an initial spin-up period of 90 days.

[0120] According to the experimental premise and verification objectives, two groups of comparative experiments were designed.

[0121] The first set of experiments aims to evaluate the effect of using reinforcement learning to intelligently select the mixing weights of the mixed assimilation scheme of the time-lagged ensemble background error covariance and the static background error covariance under different observation sparsity conditions, and to improve the average performance of the annual cycle assimilation. The experiment focuses on verifying whether the ensemble background error covariance obtained by the statistics of the time-lagged ensemble members can select appropriate integration weights through reinforcement learning during the mixing process, thereby effectively improving the performance of the assimilation system.

[0122] The second group of experiments verifies the transitional weather process under sparse observation conditions. During the cyclic assimilation period, when the system experiences a mutation period, the experiment explores whether multiple static background error covariances and multiple time-lagged set background error covariances can effectively capture the background error information of real-time state variables after mixing weights through reinforcement learning intelligent decision-making, and quickly respond to changes in the current meteorological situation, thereby improving the assimilation performance during the mutation period. At the same time, the experiment focuses on comparing the adaptability and response efficiency of the mixed background error covariance matrix of different methods under the mutation period, verifying its comprehensive improvement effect on the assimilation performance throughout the year.

[0123] Six comparison methods were set up in the two groups of experiments: 3DVar, a hybrid assimilation method with fixed weights and time lag (CTL-EnVar), EnKF, a hybrid assimilation method (MLP-EnVar) that only uses a multi-layer perceptron (MLP) in the feature extraction part of the deep reinforcement learning model, a reinforcement learning method based on CCNN (DRL-EnVar) proposed in this invention, and a classic hybrid covariance data assimilation method (HCDA). In order to alleviate the problems of sampling error and long-distance false correlation caused by the small number of set members, all methods except pure 3DVar localized the set error covariance matrix. The localization uses the Gaspari-Cohn function (fifth-order piecewise rational function), and its weight value decreases from 1 to 0 in the form of a Gaussian distribution approximation within the threshold radius.

[0124] Experiment 1: Comparative experimental setup with sparse observations and no mutations

[0125] In this experiment, all observation data are sparsely processed by uniform masking, that is, one observation value is masked every n state variables. Since the total number of model state variables used in the experiment is 40, in order to ensure the uniformity of the mask, n=9, 4, 3, 1 are selected respectively, and the corresponding observation masking ratios are 10%, 20%, 25% and 50%. In addition, this experiment does not involve transitional weather processes, that is, F=8.0 is maintained throughout the entire cycle of cyclic assimilation in the experimental setting. The main objectives of the experiment include the following two aspects: on the one hand, verify whether the time-lagged set members statistically analyzed by the cyclic assimilation cycle can provide flow-dependent background error information under the condition of sparse observations, and whether the observation information can be effectively propagated from observed variables to unobserved variables; on the other hand, verify whether the deep reinforcement learning method can make an optimal decision on the mixing weight based on the long-term reward mechanism and the current environmental state on the basis of extracting data features in combination with deep learning, thereby improving the utilization efficiency of multiple background error information sources and achieving a significant improvement in the average assimilation performance within the verification period. This experimental design aims to explore the potential of combining time-lagged ensemble background error covariance with deep reinforcement learning under the special conditions of sparse observations, and to provide practical support for improving the performance of the cyclic assimilation system.

[0126] Experiment 2: Comparative experimental setup with sparse observations and mutations

[0127] On the basis of observational uniform masking, this experiment increases the complexity and particularity of the evolution of state variables during the verification period. Specifically, a one-month mutation period (June 15 to July 15) was added to the cyclic assimilation period, during which the values ​​of the model state variables were significantly greater than those in the stable state period, simulating the plum rain weather phenomenon\cite{yihui2005east}. During these 30 days, the forcing term in the mutation period was set to F=15.0, and the rest of the time (including the spin up period) was kept at F=8.0, as shown in the figure. Since it involves a one-month state variable mutation period, this experiment changes the climate state characteristics during the verification period. Therefore, compared with Experiment 1, this experiment counts the background error covariance information under two different climate states, namely, B F8S15 and B F15 , and conduct a comprehensive analysis on them.

[0128] All experiments were performed on a uniformly configured workstation equipped with a single Geforce RTX 3090 GPU and an AMD EPYC 7H12 CPU.

[0129] The fixed mixing parameters of the three methods are reproduced through the enumeration method adopted by each method itself, aiming to find the optimal fixed parameters.

[0130] (1) Fixed mixing parameter time lag ensemble - three-dimensional variational mixing assimilation method CTL-EnVar, where the mixed background error covariance calculation formula is: B h =(1-α)B nmc +αB TL

[0131] In this method, the construction process of the time-lag set members is as follows: in the 3DVar assimilation process, the model forecast fields of different time periods at the same time in the historical forecast samples are used to form the time-lag set members. The specific operation is to perform assimilation and 48-hour forecast every 3 hours, save the forecast results every 3 hours, form 16 time-lag set members, and obtain 120 difference fields. To ensure the fairness of the comparative experiment, other methods in this experiment are set to perform assimilation every 6 hours. Therefore, the CTL-EnVar method is also changed to perform assimilation and 48-hour forecast every 6 hours. The static background error covariance matrix in this method is calculated by the NMC method. Different from the four mixing parameters of 0.25, 0.5, 0.75 and 1.0 listed in the original text, the mixing parameters are regarded as the weights of the set background error covariance. In order to ensure that the method can achieve optimal performance, this experiment expanded the enumeration range to 0 to 1, with a preliminary verification step of 0.1, a total of 10 parameters, and each mixed parameter was verified 10 times, and the average value was taken as the final result. The experiment found that when the mixed parameter was 0.1, the performance was relatively optimal. In order to further confirm whether a smaller mixed parameter can bring better performance, the second enumeration range was 0.01 to 0.1, with a step of 0.01. It was found again that when 0.01, the performance was optimal. After that, the enumeration range was further refined to 0.001 to 0.01, with a step of 0.001, and finally refined to 0.0001 to 0.001, with a step of 0.0001. Finally, in the two experimental verification scenarios, the optimal parameters of CTL-EnVar were concentrated between 0.0001 and 0.01, and the optimal mixed parameters were different under different observation sparsity ratios and experimental settings.

[0132] (2) EnKF method: The purpose of setting EnKF as a comparative experiment is to verify whether the performance of the proposed innovative method can be close to or equivalent to EnKF under different numbers of set members, while the computational cost is much lower than EnKF. What needs to be determined is the choice of the number of set members under different experimental settings. The enumeration range of the number of set members is 5 to 40, the step size is 5, and the upper limit is set to 40, because the pattern state variables used in the present invention are 40. Compared with the business EnKF method, the ratio of the number of pattern state variables to the number of set members of 1:1 is almost impossible to achieve with 40 set members, and it will bring extremely high computational costs. Each enumeration experiment is repeated 10 times and the average value is taken.

[0133] (3) HCDA method: The purpose of setting HCDA as a comparative experiment is to verify whether the method proposed in the present invention is equivalent to or close to HCDA under different numbers of set members, and the computational cost is much lower than HCDA. The parameters that need to be determined include the selection of the number of set members and the weight parameter between the static background error covariance and the set prediction error covariance. The enumeration range of the number of set members is 5 to 40, with a step size of 5, and the enumeration range of the weight parameter is 0 to 1, with a step size of 0.1. The experiment found that the number of set members and the weight parameter jointly affect the performance of HCDA, which is manifested in that under the same number of set members, as the weight parameter increases, the performance first significantly improves and then gradually decreases. The optimal performance under different numbers of set members is similar, that is, the increase in the number of set members does not significantly improve the performance of HCDA. Therefore, the enumeration range of the number of set members is reduced to 2 to 5, with a step size of 1, and the enumeration range of the weight parameter is still 0 to 1, with a step size of 0.1. Each enumeration experiment is repeated 10 times and the average value is taken.

[0134] The present invention uses two evaluation indicators to evaluate the assimilation performance of the assimilation method within the experimental verification cycle: one is a direct evaluation indicator - the distance between the analysis field and the real field, which is the root mean square error RMSEa; the other is an indirect evaluation indicator - the distance between the forecast field and the real field, RMSEf. The forecast duration is selected according to different verification requirements, from 3 hours to 48 hours, and the average performance of the forecast every 3 hours. In addition, for the special settings of Experiment 2, since the generation and disappearance cycle of convective-scale weather in the simulated plum rain weather phenomenon is usually hours or shorter, the performance comparison of the 1-hour forecast results in the mutation period is added. In order to avoid the contingency of the experiment, the experiment of each method was repeated 50 times, and the average value was taken as the final evaluation result.

[0135] Embodiment 2:

[0136] This embodiment provides a system for numerical weather forecasting based on deep reinforcement learning. Figure 4 This is a system structure diagram of the numerical weather forecast based on deep reinforcement learning in the present invention. Figure 4 As shown, the system includes:

[0137] A custom environment setting module 201 is used to set a custom environment for ensemble variation data assimilation;

[0138] Parameter acquisition module 202, used to acquire policy network parameters and value network parameters;

[0139] A policy network parameter updating module 203 is used to obtain updated policy network parameters based on the custom environment, policy network parameters and value network parameters of the set variational data assimilation based on a deep reinforcement learning method;

[0140] A mixed background-error covariance matrix determination module 204 is used to determine a mixed background-error covariance matrix according to the updated policy network parameters, wherein the mixed background-error covariance matrix is ​​a weighted average of a static covariance matrix and a set covariance matrix;

[0141] A numerical weather forecast prediction model training module 205 is used to update the adaptive hybrid parameter strategy model training according to the hybrid background-error covariance matrix to obtain a numerical weather forecast prediction model;

[0142] The numerical weather forecast prediction module 206 is used to perform numerical weather forecast prediction according to the numerical weather forecast prediction model.

[0143] Embodiment three:

[0144] This embodiment provides an electronic device, including a memory and a processor, the memory is used to store a computer program, and the processor runs the computer program to enable the electronic device to execute the numerical weather forecast prediction method based on deep reinforcement learning of embodiment one.

[0145] Optionally, the above-mentioned electronic device may be a server.

[0146] In addition, an embodiment of the present invention further provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the numerical weather forecast prediction method based on deep reinforcement learning of embodiment one.

[0147] Embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0148] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0149] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0150] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0151] In this specification, each embodiment is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the system disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part.

[0152] The present invention uses specific examples to illustrate the principles and implementation methods of the present invention. The above examples are only used to help understand the method and core ideas of the present invention. At the same time, for those skilled in the art, according to the ideas of the present invention, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting the present invention.

Claims

1. A numerical weather forecasting method based on deep reinforcement learning, characterized in that: The method comprises: Set up a custom environment for ensemble variational data assimilation; Get strategy network parameters and value network parameters; Obtain updated policy network parameters based on the custom environment, policy network parameters, and value network parameters assimilated by the ensemble variational data based on a deep reinforcement learning method; Determine a mixed background-error covariance matrix according to the updated policy network parameters, wherein the mixed background-error covariance matrix is ​​a weighted average of a static covariance matrix and a set covariance matrix; Updating the adaptive hybrid parameter strategy model training according to the hybrid background-error covariance matrix to obtain a numerical weather forecast prediction model; Numerical weather forecast prediction is performed according to the numerical weather forecast prediction model.

2. The numerical weather forecasting method based on deep reinforcement learning according to claim 1, characterized in that: The custom environment for setting up ensemble variational data assimilation specifically includes: According to the assimilation forecast cycle process of the ensemble variational data assimilation system, the reinforcement learning simulation environment is customized, the state space and action space of the intelligent agent are designed, and a decision-feedback reward mechanism is established to obtain the reward function at the current moment; Get the current action and status; A custom environment that assimilates the reward function, actions, and states as collective variational data.

3. The numerical weather forecasting method based on deep reinforcement learning according to claim 2, characterized in that: The customized environment, policy network parameters and value network parameters assimilated according to the set variational data are obtained based on a deep reinforcement learning method after updating the policy network parameters, specifically including: According to the customized environment, policy network parameters and value network parameters of the ensemble variational data assimilation, based on a deep proximal policy optimization algorithm, updated policy network parameters are obtained.

4. The numerical weather forecasting method based on deep reinforcement learning according to claim 3, characterized in that: The deep proximal policy optimization algorithm specifically includes an Actor network and a Critic network. The Actor network is responsible for generating actions and selecting appropriate parameter ratios in the hybrid assimilation process; the Critic network estimates the value function to evaluate the effectiveness of the current strategy.

5. The numerical weather forecasting method based on deep reinforcement learning according to claim 3, characterized in that: Determining the mixed background-error covariance matrix according to the updated policy network parameters specifically includes: According to the updated strategy network parameters, formula B is used h =(1-β)B S +βB e , determine the mixed background-error covariance matrix; Among them, B h is the mixed background-error covariance matrix, B e is the static covariance matrix, B s is the set covariance matrix, β is the updated policy network parameter, which controls B e and B s The weight of , 0<β<1.

6. A numerical weather forecast system based on deep reinforcement learning, characterized in that: The system comprises: Custom environment setting module, used to set custom environment for ensemble variational data assimilation; Parameter acquisition module, used to obtain policy network parameters and value network parameters; A policy network parameter updating module, used to obtain updated policy network parameters based on a deep reinforcement learning method according to the custom environment, policy network parameters and value network parameters assimilated by the set variational data; A mixed background-error covariance matrix determination module, used to determine a mixed background-error covariance matrix according to the updated policy network parameters, wherein the mixed background-error covariance matrix is ​​a weighted average of a static covariance matrix and a set covariance matrix; A numerical weather forecast prediction model training module is used to update the adaptive hybrid parameter strategy model training according to the hybrid background-error covariance matrix to obtain a numerical weather forecast prediction model; The numerical weather forecast prediction module is used to perform numerical weather forecast prediction according to the numerical weather forecast prediction model.

7. An electronic device, characterized in that: The electronic device includes a memory and a processor, the memory is used to store a computer program, and the processor runs the computer program to enable the electronic device to execute the numerical weather forecasting method based on deep reinforcement learning as described in any one of claims 1-5.

8. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the numerical weather forecasting method based on deep reinforcement learning described in any one of claims 1 to 5 is implemented.

Citation Information

Patent Citations

  • Rapid updating mixing assimilation method based on time lag set

    CN105447593A

  • Shared bicycle demand prediction method based on multi-strategy improved GWOBP neural network

    CN112766533A

  • Method for scene modeling and change detection

    US20050286764A1