Quantitative rainfall estimation method and system, storage medium and electronic equipment
By using a non-parametric Bayesian machine learning algorithm in quantitative precipitation estimation, combining multiple factors and inference methods, the factor screening problem in quantitative precipitation estimation is solved, and the accuracy and application performance of the estimation are improved.
Patent Information
- Application Number
- CN202510094673.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-01-21
AI Technical Summary
Existing quantitative precipitation estimation techniques are difficult to accurately screen the factors that affect precipitation estimation, resulting in inaccurate estimation results.
The quantitative precipitation estimation algorithm based on non-parametric Bayesian machine learning is used, combined with nine elements such as radar reflectivity factor, automatic station information, radar inversion wind field, and weather type, and the use of a priori distribution model, variational reasoning and Jason's inequality, the lower bound of precipitation estimation is determined and maximized.
It improves the accuracy of quantitative precipitation estimation and business application performance, and can more effectively screen key factors affecting precipitation estimation to obtain more accurate precipitation estimation results.
Smart Images

Figure CN120013004A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of meteorological forecasting, and in particular to a quantitative precipitation estimation method, system, storage medium and electronic equipment. Background Art
[0002] At present, Quantitative Precipitation Estimation (QPE) is the basis of quantitative precipitation forecast (QPF), heavy precipitation near-warning and other services, and is an important part of short-term forecasting. It has always been the focus and difficulty in forecasting services. Automatic weather stations are currently the most direct way to observe precipitation. Due to the uneven spatial distribution of automatic weather stations, it is easy for the observation data to not fully reflect the distribution characteristics of precipitation.
[0003] Radar reflectivity is the main factor affecting precipitation, but precipitation is affected by many factors. How to screen out the factors that affect the accuracy of quantitative precipitation estimation and obtain the most optimized precipitation estimation is a difficult problem in quantitative precipitation estimation.
[0004] Therefore, a solution is urgently needed. Summary of the invention
[0005] One of the purposes of the present invention is to provide a quantitative precipitation estimation method, which is based on long-term massive observation data and adopts a quantitative precipitation estimation algorithm based on non-parametric Bayesian machine learning. The algorithm combines 9 elements such as radar reflectivity factor, automatic station information, radar inversion wind field, weather type, etc. to constrain the radar quantitative precipitation estimation algorithm. First, a radar data set of the quantitative precipitation estimation area and an automatic station real-time observation information data set are constructed, and a prior distribution model is constructed based on the data set. Based on variational reasoning, the approximate distribution of the posterior distribution of precipitation estimation is defined according to the prior distribution model. Based on Jason's inequality, the lower bound of precipitation estimation is determined according to the approximate distribution, and the lower bound of precipitation estimation is maximized to obtain a quantitative precipitation estimation result, thereby carrying out quantitative precipitation estimation in order to improve its business application performance.
[0006] An embodiment of the present invention provides a quantitative precipitation estimation method, comprising:
[0007] Get data set M;
[0008] Based on the data set M, construct a prior distribution model p(M|D); where D is a family of models;
[0009] Based on variational inference, according to the prior distribution model p(M|D), the approximate distribution q(Θ) of the posterior distribution of precipitation estimation is defined; where Θ is the probability model parameter;
[0010] Based on Jason's inequality, the lower bound E of precipitation estimation is determined according to the approximate distribution q(Θ)q [log(p(Θ,M))]-E q [log(q(Θ))]; where E q is the distribution range of precipitation estimation data;
[0011] The lower bound E for precipitation estimation q [log(p(Θ,M))]-E q [log(q(Θ))] is maximized to obtain quantitative precipitation estimation results
[0012] Optionally, the data set M includes at least: a radar data set of the quantitative precipitation estimation area and an automatic station information data set, wherein the radar data set is 1Km*1Km radar reflectivity factor grid data within the quantitative precipitation estimation area, and the automatic station information data set contains the latitude and longitude position coordinates and actual precipitation observation data of each automatic station in the quantitative precipitation estimation area.
[0013] Optionally, constructing a prior distribution model p(M|D) based on the data set M includes:
[0014] Dirichlet nonparametric Bayesian model G 0 is the basic probability distribution on the probability measure space Ω, with the concentration parameter α 0 >0, if the probability distribution G on the probability measure space Ω obeys the basic probability distribution G 0 ,but:
[0015] G~DP(α 0 ,G 0 )
[0016] Where: Basic probability distribution G 0 Determine the distribution of basic components in the prior distribution model p(M|D); DP is a Dirichlet process;
[0017] Based on the Dirichlet process mixture model DPM, a generation probability is added to each data point in the given precipitation estimation area as the prior distribution of the data:
[0018] m i ~p(m|θ i )
[0019] Where: parameter θ i It obeys the probability distribution G, i∈[N] is a set with values ranging from 1 to the total number of data points N, each data point generates a probability parameter, N is the total number of data points; p is the conditional probability density function; m is the precipitation estimate of the data point;
[0020] By comparing different families of prior distribution probability models m i~p(m|θ i ) is used to select the optimal model as the prior distribution model p(M|D):
[0021] p(M|D)=∫ Θ (M|Θ)p(Θ|D)
[0022] Where: D represents a family of models.
[0023] Optionally, the selecting the optimal model as the prior distribution model p(M|D) by comparing the likelihood functions of different families of models includes:
[0024] Assume that the data set M = {m 1 ,m 2 ,m 3 …m n} data is independent, read the actual precipitation observation data of n automatic stations in the quantitative precipitation estimation area, and randomly arrange the actual precipitation observation data of n automatic stations to obtain {F(i)}, where i = 1, 2, ... n;
[0025] set up Ω=Ω t-1 , for each automatic station real-time precipitation observation data i∈{F(1),F(2)…F(n)}, the indicator factor β of the real-time precipitation observation data of the automatic station composed of n random permutations formed based on the function {F(i)} i Sampling is performed; among them, is the concentrated parameter at time t-1; Ω t-1 is the probability measure space at time t-1;
[0026] Calculate the likelihood estimate f of the observed data based on the current K clusters k (m i ),
[0027] f k (m i )=p(m i |β i =k,M \i ,ζ)
[0028]
[0029] In the formula, β i For a new category, M \i From the data set M = {m 1 ,m 2 ,m 3 …m n}, the corresponding subscript data is removed; ζ is the distribution parameter; k is the observed value; m i is a random variable; Similar to k, it represents another observation value other than k;
[0030] β i To take a sample:
[0031]
[0032]
[0033] In the formula, is the amount of data already in the kth class; E i is the preset observation data sample; K is the number of observation data samples; f k represents the probability density function for the kth class; δ is the Kronecker delta function, when β i = k, δ(β i , k) = 1, otherwise 0; To represent the probability density function except for the kth class;
[0034] if Then increase the number of clusters K by 1, K = K + 1;
[0035] Check the amount of observation data for calculating the likelihood function of each type of clustering. If the total number of observation data for one type is 0, delete the corresponding type and the number of clusters K will be reduced by 1, K=K-1.
[0036] An embodiment of the present invention provides a quantitative precipitation estimation system, comprising:
[0037] An acquisition module, used to acquire a data set M;
[0038] A construction module is used to construct a prior distribution model p(M|D) based on a data set M, where D is a family of models;
[0039] A definition module is used to define the approximate distribution q(Θ) of the posterior distribution of precipitation estimation based on the prior distribution model p(M|D) based on variational inference; wherein Θ is a probability model parameter;
[0040] Determination module, used to determine the lower bound E of precipitation estimation based on the approximate distribution q(Θ) based on Jason inequality q [log(p(Θ,M))]-E q [log(q(Θ))]; where E q is the interval of the estimated distribution of precipitation;
[0041] Maximization module, used to estimate the lower bound E of precipitation q [log(p(Θ,M))]-E q [log(q(Θ))] is maximized to obtain quantitative precipitation estimation results
[0042] Optionally, the data set M includes at least: a radar data set of the quantitative precipitation estimation area and an automatic station information data set, wherein the radar data set is 1Km*1Km radar reflectivity factor grid data within the quantitative precipitation estimation area, and the automatic station information data set contains the latitude and longitude position coordinates and actual precipitation observation data of each automatic station in the quantitative precipitation estimation area.
[0043] Optionally, constructing a prior distribution model p(M|D) based on the data set M includes:
[0044] Dirichlet nonparametric Bayesian model G 0 is the basic probability distribution on the probability measure space Ω, with the concentration parameter α 0 >0, if the probability distribution G on the probability measure space Ω obeys the basic probability distribution G 0 ,but:
[0045] G~DP(α 0 ,G 0 )
[0046] Where: Basic probability distribution G 0 Determine the distribution of basic components in the prior distribution model p(M|D); DP is a Dirichlet process;
[0047] Based on the Dirichlet process mixture model DPM, a generation probability is added to each data point in the given precipitation estimation area as the prior distribution of the data:
[0048] m i ~p(m|θ i )
[0049] Where: parameter θ i It obeys the probability distribution G, i∈[N] is a set with values ranging from 1 to the total number of data points N, each data point generates a probability parameter, N is the total number of data points; p is the conditional probability density function; m is the precipitation estimate of the data point;
[0050] By comparing different families of prior distribution probability models m i ~p(m|θ i ) is used to select the optimal model as the prior distribution model p(M|D):
[0051] p(M|D)=∫ Θ (M|Θ)p(Θ|D)
[0052] Where: D represents a family of models.
[0053] Optionally, the selecting the optimal model as the prior distribution model p(M|D) by comparing the likelihood functions of different families of models includes:
[0054] Assume that the data set M = {m 1 ,m 2 ,m 3 …m n} data is independent, read the actual precipitation observation data of n automatic stations in the quantitative precipitation estimation area, and randomly arrange the actual precipitation observation data of n automatic stations to obtain {F(i)}, where i = 1, 2, ... n;
[0055] set up Ω=Ω t-1 , for each automatic station real-time precipitation observation data i∈{F(1),F(2)…F(n)}, the indicator factor β of the real-time precipitation observation data of the automatic station composed of n random permutations formed based on the function {F(i)} i Sampling is performed; among them, is the concentrated parameter at time t-1; Ω t-1 is the probability measure space at time t-1;
[0056] Calculate the likelihood estimate f of the observed data based on the current K clusters k (m i ),
[0057] f k (m i )=p(m i |β i =k,M \i ,ζ)
[0058]
[0059] In the formula, β i For a new category, M \i From the data set M = {m 1 ,m 2 ,m 3 …m n}, the corresponding subscript data is removed; ζ is the distribution parameter; k is the observed value; m i is a random variable; Similar to k, it represents another observation value other than k;
[0060] β i To take a sample:
[0061]
[0062] In the formula, is the amount of data already in the kth class; Ei is the preset observation data sample; K is the number of observation data samples; f k represents the probability density function for the kth class; δ is the Kronecker delta function, when β i = k, δ(β i , k) = 1, otherwise 0; To represent the probability density function except for the kth class;
[0063] if Then increase the number of clusters K by 1, K = K + 1;
[0064] Check the amount of observation data for calculating the likelihood function of each type of clustering. If the total number of observation data for one type is 0, delete the corresponding type and the number of clusters K will be reduced by 1, K=K-1.
[0065] An embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored, and a processor executes the computer program to implement any of the above methods.
[0066] An embodiment of the present invention provides an electronic device, which includes a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement any one of the methods described above.
[0067] Other features and advantages of the present invention will be described in the following description, and partly become apparent from the description, or understood by practicing the present invention. The purpose and other advantages of the present invention can be realized and obtained by the structures particularly pointed out in the written description and the accompanying drawings.
[0068] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0069] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings:
[0070] Figure 1 is a schematic diagram of a quantitative precipitation estimation method in an embodiment of the present invention;
[0071] Figure 2 Schematic diagram of a quantitative precipitation estimation system in an embodiment of the present invention. DETAILED DESCRIPTION
[0072] The preferred embodiments of the present invention are described below in conjunction with the accompanying drawings. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present invention, and are not used to limit the present invention.
[0073] The embodiment of the present invention provides a quantitative precipitation estimation method, such as Figure 1 As shown, including:
[0074] S1. Obtain data set M;
[0075] S2. Based on the data set M, construct a prior distribution model p(M|D); where D is a family of models;
[0076] S3. Based on variational inference, according to the prior distribution model p(M|D), the approximate distribution q(Θ) of the posterior distribution of precipitation estimation is defined; where Θ is the probability model parameter;
[0077] S4. Based on Jason's inequality, the lower bound E of precipitation estimation is determined according to the approximate distribution q(Θ) q [log(p(Θ,M))]-E q [log(q(Θ))]; where E q is the distribution range of precipitation estimation data;
[0078] S5. Lower bound E for precipitation estimation q [log(p(Θ,M))]-E q [log(q(Θ))] is maximized to obtain quantitative precipitation estimation results
[0079] The data set M at least includes: a radar data set of the quantitative precipitation estimation area and an automatic station information data set, wherein the radar data set is 1Km*1Km radar reflectivity factor grid data in the quantitative precipitation estimation area, and the automatic station information data set contains the latitude and longitude position coordinates and actual precipitation observation data of each automatic station in the quantitative precipitation estimation area.
[0080] The prior distribution model p(M|D) is constructed based on the data set M, including:
[0081] Dirichlet nonparametric Bayesian model G 0 is the basic probability distribution on the probability measure space Ω, with the concentration parameter α 0 >0, if the probability distribution G on the probability measure space Ω obeys the basic probability distribution G 0 ,but:
[0082] G~DP(α 0 ,G 0 )
[0083] Where: Basic probability distribution G 0Determine the distribution of the basic components in the prior distribution model p(M|D); DP is the Dirichlet process; "Dirichlet Process" is a random process widely used in Bayesian nonparametric statistics. In this method, it is used to model the probability of precipitation distribution when the specific form of precipitation distribution is unknown but there is some prior knowledge;
[0084] Based on the Dirichlet process mixture model DPM, a generation probability is added to each data point in the given precipitation estimation area as the prior distribution of the data:
[0085] m i ~p(m|θ i )
[0086] Where: parameter θ i It obeys the probability distribution G, i∈[N] is a set with values ranging from 1 to the total number of data points N, each data point generates a probability parameter, N is the total number of data points; p is the conditional probability density function (PDF); m is the precipitation estimate of the data point; this expression represents the probability distribution of precipitation estimate data under the condition of given parameters;
[0087] By comparing different families of prior distribution probability models m i ~p(m|θ i ) is used to select the optimal model as the prior distribution model p(M|D):
[0088] p(M|D)=∫ Θ (M|Θ)p(Θ|D)
[0089] Where: D represents a family of models.
[0090] The method of selecting the optimal model as the prior distribution model p(M|D) by comparing the likelihood functions of different family models includes:
[0091] Assume that the data set M = {m 1 ,m 2 ,m 3 …m n} data is independent, read the actual precipitation observation data of n automatic stations in the quantitative precipitation estimation area, randomly arrange the actual precipitation observation data of n automatic stations, and obtain {F(i)}, where i = 1, 2, ... n; randomly arrange the obtained precipitation observation data to obtain observation data with different arrangements;
[0092] set up Ω=Ω t-1, for each automatic station real-time precipitation observation data i∈{F(1),F(2)…F(n)}, the indicator factor β of the real-time precipitation observation data of the automatic station composed of n random permutations formed based on the function {F(i)} i Sampling is performed; among them, is the concentrated parameter at time t-1; Ω t-1 is the probability measurement space at time t-1; for each automatic station's actual precipitation observation data, a random arrangement is generated based on a specific function, and sampling is performed to form the corresponding indicator factor, which represents the centralized parameters (such as meteorological parameters, precipitation values, etc.) at the current moment, including the probability measurement space at the moment.
[0093] Calculate the likelihood estimate f of the observed data based on the current K clusters k (m i ),
[0094] f k (m i )=p(m i |β i =k,M \i ,ζ)
[0095]
[0096] In the formula, β i For a new category, M \i From the data set M = {m 1 ,m 2 ,m 3 …m n} removes the data corresponding to the subscript (the unique index corresponding to each data point in the data set) (that is, k in the formula); ζ is the distribution parameter; k is the observed value; m i is a random variable; Similar to k, it represents another observation value other than k; for each current cluster, calculate the likelihood estimate of its observed data. That is, based on the current data set and clustering model, evaluate the degree of fit of the data. By setting the number of clusters and sample data, gradually fit different models and calculate the likelihood function of each cluster. The relevant parameters in the calculation formula, such as the clustering parameters, observation values, and random variables of the data, are finally obtained to obtain the likelihood function value of each cluster. By comparing the likelihood functions of different family models, the model with the highest likelihood value is selected as the optimal model and used as the prior distribution model.
[0097] β i To take a sample:
[0098]
[0099]
[0100] In the formula, is the amount of data already in the kth class; E i is the preset observation data sample; K is the number of observation data samples; f k represents the probability density function for the kth class; δ is the Kronecker delta function, when β i = k, δ(β i , k) = 1, otherwise 0; To represent the probability density function except for the kth class;
[0101] if Then increase the number of clusters K by 1, K = K + 1;
[0102] Check the amount of observation data for calculating the likelihood function of each type of clustering. If the total number of observation data for one type is 0, delete the corresponding type and the number of clusters K will be reduced by 1, K=K-1.
[0103] During the clustering process, sampling is performed to determine whether the number of clusters needs to be adjusted. If the amount of data in a cluster is 0, the cluster is deleted and the number of clusters is reduced. Adjusting the number of clusters can improve the estimation accuracy and ensure that each cluster can effectively represent the distribution of the data. When the number of clusters reaches the preset condition, the number of clusters increases by 1 to further refine the classification. All clusters are calculated and checked to ensure that the total number of observed data for the likelihood function of each cluster is appropriate. During the data sampling process, the Kronecker function is used to represent the relationship between different data categories. If the number of clusters still does not meet the conditions after adjustment (such as a category is empty), the category needs to be deleted and the number of clusters needs to be reduced. By dynamically adjusting the number of clusters, it can be ensured that the clustering results of precipitation data are more consistent with the distribution of actual data. After multiple iterations, when the number of clusters and the likelihood function reach the preset accuracy, the optimal precipitation estimation result is finally output.
[0104] The working principle and beneficial effects of the above technical solution are:
[0105] Bayesian machine learning is an important branch of machine learning. The Bayesian method was first proposed by British mathematician Thomas Bayes. After more than 200 years of development, it has become an important part of machine learning and has been widely used in many statistical machine learning fields such as multivariate structured output prediction. The basic concept of Bayes' theorem is that given the prior distribution and likelihood function of a model, the posterior probability distribution p(Θ|M) of the model can be obtained by the Bayesian formula:
[0106]
[0107] In formula (1), Θ is the probability model parameter, M is the data set, which includes radar data set and automatic station information data set in this application, and p 0 (Θ) is the prior distribution function of the model, p(M|Θ) is the likelihood function, and p(M) is a constant.
[0108] The choice of model is the basic problem of Bayesian method. Dirichlet non-parametric Bayesian model can obtain the number of cluster centers through machine automatic learning based on the characteristics of the data itself.
[0109] The Dirichlet nonparametric Bayesian model assumes that G 0 is a random probability distribution on the probability measure space Ω, with the concentration parameter α 0 >0, the probability distribution G on the space Ω obeys the basic probability distribution G 0 ,but:
[0110] G~DP(α 0 ,G 0 )(2)
[0111] Where: Basic distribution G 0 Determines the distribution of basic components in the model.
[0112] However, the probability distribution obtained by the Dirichlet process is discrete. In order to cluster different groups of data with certain similarities, the Dirichlet process mixture model DPM (Dirichlet Process Mixture) (Antoniak.1974) is introduced to add a generation probability to each data point as the prior distribution of the data:
[0113] m i ~p(m|θ i )(3)
[0114] In formula (3): parameter θ i It obeys the G distribution, i∈[N] is the probability of generating each data point, and N is the number of data points. When G obeys the Dirichlet process distribution, the model is called the Dirichlet process mixture model. The Bayesian method selects the optimal model by comparing the likelihood functions of different family models:
[0115] p(M|D)=∫ Θ (M|Θ)p(Θ|D) (4)
[0116] In formula (4), D represents a family of models. In the absence of an obvious prior function, assuming that p(Θ|D) is uniformly distributed, overfitting of the model can be avoided by integrating formula (4).
[0117] In this study, it is assumed that the observation set M = {m 1 ,m 2 ,m 3 …m n} data are independent. In order to obtain the indicator factor for each observation, in the non-parametric Bayesian model using the Dirichlet process as the prior distribution, Gibbs sampling is used to obtain the likelihood function of different family models to select the optimal model. The steps are as follows:
[0118] (1) Data initialization: The system reads data and randomly arranges n observation data to obtain {F(i)}, where i = 1, 2, … n.
[0119] (2) Category clustering: Ω=Ω t-1 , for each observation data i∈{F(1),F(2)…F(n)}, the indicator factor β for each data i Take samples.
[0120] Calculate the likelihood estimate of the observed data based on the existing K clusters:
[0121] f k (m i )=p(m i |β i =k,M \i ,ζ) (5)
[0122]
[0123] In formula (5) and formula (6), β i It is a new category, M \i Represents removing the corresponding subscripted data from the corresponding observed data set; ζ is the distribution parameter.
[0124] β i To take a sample:
[0125]
[0126] In formula (7) and formula (8), is the amount of data already in the kth class. If Then increase the number of clusters by one, K=K+1.
[0127] (3) Cluster update: Check the number of observations in each category. If the total number of observations in a category is 0, delete the category and reduce the number of clusters by one, K = K-1.
[0128] The reasoning method of the Bayesian model is an important part of Bayesian learning. Given the prior distribution, the posterior distribution of the Bayesian model is usually unsolvable and requires an effective reasoning method. In this application, variational reasoning is used for prediction.
[0129] In quantitative estimation of precipitation forecast, given a data set M and a prior distribution model p(M|D), a variational method is used to define the approximate distribution q(Θ) of the posterior distribution of precipitation estimation. Using Jason's inequality, a lower bound for precipitation estimation can be obtained:
[0130] logp(M)≥E q [log(p(Θ,M))]-E q [log(q(Θ))] (9)
[0131] By maximizing this lower bound estimate:
[0132]
[0133] The quantitative precipitation estimation can be solved.
[0134] To make quantitative precipitation estimates, we first need to collect some actual data related to precipitation. The dataset consists of a radar dataset and an automatic station information dataset. The radar dataset contains radar reflectivity factor data in the precipitation area. The spatial resolution of the radar data is 1km*1km. These radar reflectivity data can be used to estimate the distribution of precipitation, because there is a certain relationship between radar signals and precipitation. The automatic station information dataset contains the automatic weather station data in the quantitative precipitation estimation area, which records the latitude and longitude information of each station and the precipitation observation data at the corresponding time point. These real-time observation data serve as a reference for actual precipitation and play a role in verification and auxiliary estimation.
[0135] In order to be able to infer accurate precipitation estimates from the collected data, it is necessary to construct a prior distribution model, the purpose of which is to set the prior information of precipitation based on the existing data and background knowledge. The Dirichlet process (DP) is a non-parametric Bayesian model that can be used to model distributions with unknown number of populations. In quantitative precipitation estimation, the Dirichlet process is used to model different types of precipitation processes. Through the Dirichlet process, a mixed model can be constructed to divide the data into several categories and assign a probability to each category. In this application, DPM is used to establish a prior distribution model for precipitation. Specifically, the distribution of data points can be regarded as generated from a mixed model, and the parameters of each mixed component are unknown. Therefore, the distribution is defined by the Dirichlet process, thereby providing prior distribution information for the estimation of precipitation. In DPM, a prior distribution is used to describe the distribution pattern of precipitation. By constructing a prior distribution, the model can be updated when new observation data is available, thereby gradually approaching the true precipitation distribution.
[0136] After establishing the prior distribution model, posterior inference is performed. Since it is often impossible to directly calculate the posterior distribution in actual situations, it is necessary to use variational inference to approximate the calculation. Variational inference is an optimization method used to find approximate solutions to complex probability distributions. In the precipitation estimation problem, the purpose of variational inference is to approximate the complex posterior distribution by constructing a simple approximate distribution. Specifically, this application uses variational inference to infer the posterior distribution of precipitation from the prior distribution. During the reasoning process, the goal of variational inference is to minimize the KL divergence of the approximate posterior distribution by optimizing a set of parameters, thereby approximating the true posterior distribution as accurately as possible.
[0137] In order to further optimize the estimation results and ensure the effectiveness of the model under given data, this application introduces Jensen's inequality to calculate the lower bound. Jensen's inequality is a mathematical tool in Bayesian reasoning, which provides a method to approximate the objective function by optimizing the lower bound. In this application, Jensen's inequality is used to establish the lower bound of precipitation estimation. By calculating the lower bound, it can be ensured that the estimate of the posterior distribution will not deviate too far from the true value, and the maximization of this lower bound also makes it possible to find the optimal precipitation estimation result in the multiple solution space.
[0138] After the lower bound is calculated, the lower bound is maximized. By maximizing the lower bound, a quantitative precipitation estimate that is as accurate as possible can be obtained. Specifically, for a given data set, based on variational inference and Jason's inequality, an objective function is obtained, which represents the size of the lower bound. By optimizing the objective function and maximizing the lower bound, the optimal precipitation estimate is finally obtained. It is often necessary to estimate the likelihood function of the data points and to select the optimal model for precipitation estimation. The likelihood function describes the probability of the observed data when the model parameters are given. In precipitation estimation, the likelihood function is used to calculate the difference between the precipitation data and the model prediction. By maximizing the likelihood function, the precipitation can be estimated more accurately. In actual data, there are often many different types of precipitation. Therefore, clustering methods are used to group the data to better understand the distribution characteristics of different types of precipitation and obtain more accurate estimates through sampling calculations.
[0139] Through the above steps, the result of quantitative precipitation estimation can be finally obtained. The whole process combines methods such as Bayesian reasoning, variational reasoning, and Jason inequality to ensure the accuracy and reliability of the estimation results. The model continuously optimizes and approximates the actual precipitation through prior distribution models, variational reasoning, and lower bound maximization, and finally achieves high-precision precipitation estimation. This application has broad application prospects in the field of meteorology, especially in extreme weather warning and climate research.
[0140] The quantitative precipitation estimation method and system of the present invention are based on long-term massive observation data and adopt a quantitative precipitation estimation algorithm based on non-parametric Bayesian machine learning. The algorithm combines 9 elements such as radar reflectivity factor, automatic station information, radar inversion wind field, weather type, etc. to constrain the radar quantitative precipitation estimation algorithm. First, a radar data set of the quantitative precipitation estimation area and an automatic station real-time observation information data set are constructed, and a prior distribution model is constructed based on the data set. Based on variational reasoning, the approximate distribution of the posterior distribution of precipitation estimation is defined according to the prior distribution model. Based on Jason's inequality, the lower bound of precipitation estimation is determined according to the approximate distribution, and the lower bound of precipitation estimation is maximized to obtain a quantitative precipitation estimation result, thereby carrying out quantitative precipitation estimation in order to improve its business application performance.
[0141] The embodiment of the present invention provides a quantitative precipitation estimation system, such as Figure 2 As shown, including:
[0142] Acquisition module 1, used to acquire data set M;
[0143] Construction module 2 is used to construct a prior distribution model p(M|D) based on the data set M, where D is a family of models;
[0144] A definition module 3 is used to define an approximate distribution q(Θ) of the posterior distribution of precipitation estimation based on variational reasoning according to a prior distribution model p(M|D); wherein Θ is a probability model parameter;
[0145] Determination module 4 is used to determine the lower bound E of precipitation estimation based on Jason inequality and the approximate distribution q(Θ) q [log(p(Θ,M))]-E q [log(q(Θ))]; where E q is the interval of the estimated distribution of precipitation;
[0146] Maximization module 5 is used to estimate the lower bound E of precipitation q [log(p(Θ,M))]-E q [log(q(Θ))] is maximized to obtain quantitative precipitation estimation results
[0147] The data set M at least includes: a radar data set of the quantitative precipitation estimation area and an automatic station information data set, wherein the radar data set is 1Km*1Km radar reflectivity factor grid data in the quantitative precipitation estimation area, and the automatic station information data set contains the latitude and longitude position coordinates and actual precipitation observation data of each automatic station in the quantitative precipitation estimation area.
[0148] The prior distribution model p(M|D) is constructed based on the data set M, including:
[0149] Dirichlet nonparametric Bayesian model G 0 is the basic probability distribution on the probability measure space Ω, with the concentration parameter α 0 >0, if the probability distribution G on the probability measure space Ω obeys the basic probability distribution G 0 ,but:
[0150] G~DP(α 0 ,G 0 )
[0151] Where: Basic probability distribution G 0 Determine the distribution of basic components in the prior distribution model p(M|D); DP is a Dirichlet process;
[0152] Based on the Dirichlet process mixture model DPM, a generation probability is added to each data point in the given precipitation estimation area as the prior distribution of the data:
[0153] m i ~p(m|θ i )
[0154] Where: parameter θ iIt obeys the probability distribution G, i∈[N] is a set with values ranging from 1 to the total number of data points N, each data point generates a probability parameter, N is the total number of data points; p is the conditional probability density function; m is the precipitation estimate of the data point;
[0155] By comparing different families of prior distribution probability models m i ~p(m|θ i ) is used to select the optimal model as the prior distribution model p(M|D):
[0156] p(M|D)=∫ Θ (M|Θ)p(Θ|D)
[0157] Where: D represents a family of models.
[0158] The method of selecting the optimal model as the prior distribution model p(M|D) by comparing the likelihood functions of different family models includes:
[0159] Assume that the data set M = {m 1 ,m 2 ,m 3 …m n} data is independent, read the actual precipitation observation data of n automatic stations in the quantitative precipitation estimation area, and randomly arrange the actual precipitation observation data of n automatic stations to obtain {F(i)}, where i = 1, 2, ... n;
[0160] set up Ω=Ω t-1 , for each automatic station real-time precipitation observation data i∈{F(1),F(2)…F(n)}, the indicator factor β of the real-time precipitation observation data of the automatic station composed of n random permutations formed based on the function {F(i)} i Sampling is performed; among them, is the concentrated parameter at time t-1; Ω t-1 is the probability measure space at time t-1;
[0161] Calculate the likelihood estimate f of the observed data based on the current K clusters k (m i ),
[0162] f k (m i )=p(m i |β i =k,M \i ,ζ)
[0163]
[0164] In the formula, β i For a new category, M\i From the data set M = {m 1 ,m 2 ,m 3 …m n}, the corresponding subscript data is removed; ζ is the distribution parameter; k is the observed value; m i is a random variable; Similar to k, it represents another observation value other than k;
[0165] β i To take a sample:
[0166]
[0167] In the formula, is the amount of data already in the kth class; E i is the preset observation data sample; K is the number of observation data samples; f k represents the probability density function for the kth class; δ is the Kronecker delta function, when β i = k, δ(β i , k) = 1, otherwise 0; To represent the probability density function except for the kth class;
[0168] if Then increase the number of clusters K by 1, K = K + 1;
[0169] Check the amount of observation data for calculating the likelihood function of each type of clustering. If the total number of observation data for one type is 0, delete the corresponding type and the number of clusters K will be reduced by 1, K=K-1.
[0170] An embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored, and a processor executes the computer program to implement any of the above methods.
[0171] An embodiment of the present invention provides an electronic device, which includes a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement any one of the methods described above.
[0172] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalents, the present invention is also intended to include these modifications and variations.
Claims
1. A quantitative precipitation estimation method, characterized in that: include: Get data set M; Based on the data set M, construct a prior distribution model p(M|D); where D is a family of models; Based on variational inference, according to the prior distribution model p(M|D), the approximate distribution q(Θ) of the posterior distribution of precipitation estimation is defined; where Θ is the probability model parameter; Based on Jason's inequality, the lower bound E of precipitation estimation is determined according to the approximate distribution q(Θ) q [log(p(Θ,M))]-E q [log(q(Θ))]; where E q is the distribution range of precipitation estimation data; The lower bound E for precipitation estimation q [log(p(Θ,M))]-E q [log(q(Θ))] is maximized to obtain quantitative precipitation estimation results 2. The quantitative precipitation estimation method according to claim 1, characterized in that: The data set M at least includes: a radar data set of the quantitative precipitation estimation area and an automatic station information data set, wherein the radar data set is 1Km*1Km radar reflectivity factor grid data in the quantitative precipitation estimation area, and the automatic station information data set contains the latitude and longitude position coordinates and actual precipitation observation data of each automatic station in the quantitative precipitation estimation area.
3. The quantitative precipitation estimation method according to claim 1, characterized in that: The prior distribution model p(M|D) is constructed based on the data set M, including: In the Dirichlet nonparametric Bayesian model, G0 is the basic probability distribution on the probability measure space Ω, and the concentration parameter α0>0. If the probability distribution G on the probability measure space Ω obeys the basic probability distribution G0, then: G~DP(α0,G0) Where: the basic probability distribution G0 determines the distribution of the basic components in the prior distribution model p(M|D); DP is the Dirichlet process; Based on the Dirichlet process mixture model DPM, a generation probability is added to each data point in the given precipitation estimation area as the prior distribution of the data: m i ~p(m|θ i ) Where: parameter θ i It obeys the probability distribution G, i∈[N] is a set with values ranging from 1 to the total number of data points N, each data point generates a probability parameter, N is the total number of data points; p is the conditional probability density function; m is the precipitation estimate of the data point; By comparing different families of prior distribution probability models m i ~p(m|θ i ) is used to select the optimal model as the prior distribution model p(M|D): p(M|D)=∫ Θ (M|Θ)p(Θ|D) Where: D represents a family of models.
4. The quantitative precipitation estimation method according to claim 1, characterized in that: The method of selecting the optimal model as the prior distribution model p(M|D) by comparing the likelihood functions of different family models includes: Assume that the data set M = {m1,m2,m3…m n } data is independent, read the actual precipitation observation data of n automatic stations in the quantitative precipitation estimation area, and randomly arrange the actual precipitation observation data of n automatic stations to obtain {F(i)}, where i = 1, 2, ... n; set up Ω=Ω t-1 , for each automatic station real-time precipitation observation data i∈{F(1),F(2)…F(n)}, the indicator factor β of the real-time precipitation observation data of the automatic station composed of n random permutations formed based on the function {F(i)} i Sampling is performed; among them, is the concentrated parameter at time t-1; Ω t-1 is the probability measure space at time t-1; Calculate the likelihood estimate f of the observed data based on the current K clusters k (m i ), f k (m i )=p(m i |b i =k,M \i ,g) In the formula, β i For a new category, M \i From the data set M = {m1,m2,m3…m n }, the corresponding subscript data is removed; ζ is the distribution parameter; k is the observed value; m i is a random variable; Similar to k, it represents another observation value other than k; β i To sample: In the formula, is the amount of data already in the kth class; E i is the preset observation data sample; K is the number of observation data samples; f k represents the probability density function for the kth class; δ is the Kronecker delta function, when β i =k, δ(βi, k) = 1, otherwise 0; To represent the probability density function except for the kth class; if Then increase the number of clusters K by 1, K = K + 1; Check the amount of observation data for calculating the likelihood function of each type of clustering. If the total number of observation data for one type is 0, delete the corresponding type and the number of clusters K will be reduced by 1, K=K-1.
5. A quantitative precipitation estimation system, characterized in that: include: An acquisition module, used to acquire a data set M; A construction module is used to construct a prior distribution model p(M|D) based on a data set M, where D is a family of models; A definition module is used to define the approximate distribution q(Θ) of the posterior distribution of precipitation estimation based on the prior distribution model p(M|D) based on variational inference; wherein Θ is a probability model parameter; Determination module, used to determine the lower bound E of precipitation estimation based on the approximate distribution q(Θ) based on Jason inequality q [log(p(Θ,M))]-E q [log(q(Θ))]; where E q is the interval of the estimated distribution of precipitation; Maximization module, used to estimate the lower bound E of precipitation q [log(p(Θ,M))]-E q [log(q(Θ))] is maximized to obtain quantitative precipitation estimation results 6. The quantitative precipitation estimation system according to claim 5, characterized in that: The data set M at least includes: a radar data set of the quantitative precipitation estimation area and an automatic station information data set, wherein the radar data set is 1Km*1Km radar reflectivity factor grid data in the quantitative precipitation estimation area, and the automatic station information data set contains the latitude and longitude position coordinates and actual precipitation observation data of each automatic station in the quantitative precipitation estimation area.
7. The quantitative precipitation estimation system according to claim 5, characterized in that: The prior distribution model p(M|D) is constructed based on the data set M, including: In the Dirichlet nonparametric Bayesian model, G0 is the basic probability distribution on the probability measure space Ω, and the concentration parameter α0>0. If the probability distribution G on the probability measure space Ω obeys the basic probability distribution G0, then: G~DP(α0,G0) Where: the basic probability distribution G0 determines the distribution of the basic components in the prior distribution model p(M|D); DP is the Dirichlet process; Based on the Dirichlet process mixture model DPM, a generation probability is added to each data point in the given precipitation estimation area as the prior distribution of the data: m i ~p(m|θ i ) Where: parameter θ i It obeys the probability distribution G, i∈[N] is a set with values ranging from 1 to the total number of data points N, each data point generates a probability parameter, N is the total number of data points; p is the conditional probability density function; m is the precipitation estimate of the data point; By comparing different families of prior distribution probability models m i ~p(m|θ i ) is used to select the optimal model as the prior distribution model p(M|D): p(M|D)=∫ Θ (M|Θ)p(Θ|D) Where: D represents a family of models.
8. The quantitative precipitation estimation system according to claim 5, characterized in that: The method of selecting the optimal model as the prior distribution model p(M|D) by comparing the likelihood functions of different family models includes: Assume that the data set M = {m1,m2,m3…m n } data is independent, read the actual precipitation observation data of n automatic stations in the quantitative precipitation estimation area, and randomly arrange the actual precipitation observation data of n automatic stations to obtain {F(i)}, where i = 1, 2, ... n; set up Ω=Ω t-1 , for each automatic station real-time precipitation observation data i∈{F(1),F(2)…F(n)}, the indicator factor β of the real-time precipitation observation data of the automatic station composed of n random permutations formed based on the function {F(i)} i Sampling is performed; among them, is the concentrated parameter at time t-1; Ω t-1 is the probability measure space at time t-1; Calculate the likelihood estimate f of the observed data based on the current K clusters k (m i ), f k (m i )=p(m i |b i =k,M \i ,g) In the formula, β i For a new category, M \i From the data set M = {m1,m2,m3…m n }, the corresponding subscript data is removed; ζ is the distribution parameter; k is the observed value; m i is a random variable; Similar to k, it represents another observation value other than k; β i To take a sample: In the formula, is the amount of data already in the kth class; E i is the preset observation data sample; K is the number of observation data samples; f k represents the probability density function for the kth class; δ is the Kronecker delta function, when β i = k, δ(β i , k) = 1, otherwise 0; To represent the probability density function except for the kth class; if Then increase the number of clusters K by 1, K = K + 1; Check the amount of observation data for calculating the likelihood function of each type of clustering. If the total number of observation data for one type is 0, delete the corresponding type and the number of clusters K will be reduced by 1, K=K-1.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and the processor executes the computer program to implement the method according to any one of claims 1 to 4.
10. An electronic device, characterized in that: The electronic device comprises a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the method according to any one of claims 1 to 4.
Citation Information
Patent Citations
Runoff probability forecasting method
CN109344999A
Multi-source rainfall data fusion algorithm and device based on Bayesian regression
CN112612995A
High-precision wind speed soft measurement method for wind power prediction of wind power plant
CN116565840A
Generating probabilistic estimates of rainfall rates from radar reflectivity measurements
US20170075034A1