A quantitative precipitation estimation method, system, storage medium and electronic device

By using a quantitative precipitation estimation algorithm based on nonparametric Bayesian machine learning, a prior distribution model is constructed by combining multiple factors. Variational inference and Jason's inequality are then used to solve the problem of inaccurate precipitation estimation caused by uneven distribution of automatic weather stations, thus achieving higher accuracy in quantitative precipitation estimation.

CN120013004BActive Publication Date: 2025-10-24METEOROLOGICAL BUREAU OF SHENZHEN MUNICIPALITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510094673.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-10-24
Estimated Expiration
2045-01-21

AI Technical Summary

Technical Problem

In existing technologies, the uneven spatial distribution of automatic weather stations means that the observation data cannot fully reflect the precipitation distribution characteristics, affecting the accuracy of quantitative precipitation estimation. How to screen out the key factors affecting quantitative precipitation estimation in order to improve the estimation accuracy is a difficult problem.

Method used

A quantitative precipitation estimation algorithm based on nonparametric Bayesian machine learning is adopted, which combines nine elements including radar reflectivity factor, automatic weather station information, radar inversion wind field and weather type. By constructing a prior distribution model, variational inference and Jason's inequality, the lower bound of precipitation estimation is determined, and the lower bound is maximized to obtain accurate quantitative precipitation estimation results.

Benefits of technology

It improves the operational performance of quantitative precipitation estimation, ensuring the accuracy and reliability of the estimation results, and has broad application prospects, especially in extreme weather early warning and climate research.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120013004B_ABST
    Figure CN120013004B_ABST
Patent Text Reader

Abstract

The application provides a quantitative precipitation estimation method, system, storage medium and electronic equipment, wherein the method comprises the following steps: obtaining a data set; constructing a prior distribution model based on the data set; defining an approximate distribution of a precipitation estimation posterior distribution according to the prior distribution model based on variational inference; determining a lower bound of the precipitation estimation according to the approximate distribution based on the Jensen inequality; and maximizing the lower bound of the precipitation estimation to obtain a quantitative precipitation estimation result. First, the radar data set of the quantitative precipitation estimation area and the automatic station real-time observation information data set are constructed, the prior distribution model is constructed based on the data set, the approximate distribution of the precipitation estimation posterior distribution is defined according to the prior distribution model based on the variational inference, the lower bound of the precipitation estimation is determined according to the approximate distribution based on the Jensen inequality, the lower bound of the precipitation estimation is maximized, and the quantitative precipitation estimation result is obtained, so that the quantitative precipitation estimation is carried out, and the business application performance is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of weather prediction, in particular to a quantitative precipitation estimation method and system, a storage medium and an electronic device. BACKGROUND

[0002] At present, quantitative precipitation estimation (QPE) is the basis for quantitative precipitation forecast (QPF), near-term warning of heavy precipitation and other businesses, and is an important part of short-term near-term forecast, and has always been the focus and difficulty in forecast business. The automatic weather station is the most direct way to observe precipitation at present. Due to the uneven spatial distribution of automatic weather stations, the observation data cannot fully reflect the distribution characteristics of precipitation.

[0003] Radar reflectivity factor is the main factor affecting precipitation, but precipitation is affected by many factors. How to select the factors affecting the accuracy of quantitative precipitation estimation and obtain the most optimized precipitation estimation is a difficult problem in current quantitative precipitation estimation.

[0004] Therefore, a solution is urgently needed. SUMMARY

[0005] One of the purposes of the present application is to provide a quantitative precipitation estimation method based on long-term mass observation data, using a quantitative precipitation estimation algorithm based on non-parametric Bayesian machine learning. The algorithm combines radar reflectivity factor, automatic station information, radar inversion wind field, weather type and other 9 elements to constrain the radar quantitative precipitation estimation algorithm. First, the radar data set of the quantitative precipitation estimation area and the automatic station real-time observation information data set are constructed, the prior distribution model is constructed based on the data set, the approximate distribution of the posterior distribution of the precipitation estimation is defined based on the prior distribution model and the variational inference, the lower bound of the precipitation estimation is determined based on the approximate distribution according to the Jensen inequality, the lower bound of the precipitation estimation is maximized, and the quantitative precipitation estimation result is obtained, so as to carry out quantitative precipitation estimation and improve the business application performance.

[0006] The quantitative precipitation estimation method provided by the present application comprises:

[0007] Obtaining a data set M;

[0008] Based on the data set M, a prior distribution model p(M|D) is constructed; wherein D is a family of models;

[0009] Based on variational inference, the approximate distribution q(Θ) of the posterior distribution of the precipitation estimation is defined according to the prior distribution model p(M|D); wherein Θ is a probability model parameter;

[0010] Based on the Jensen inequality, the lower bound E of the precipitation estimation is determined according to the approximate distribution q(Θ)q [log(p(Θ,M))]-E q [log(q(Θ))]; where E q is the distribution range of precipitation estimation data;

[0011] The lower bound E for precipitation estimation q [log(p(Θ,M))]-E q [log(q(Θ))] is maximized to obtain quantitative precipitation estimation results

[0012] Optionally, the data set M includes at least: a radar data set of the quantitative precipitation estimation area and an automatic station information data set, wherein the radar data set is 1Km*1Km radar reflectivity factor grid data within the quantitative precipitation estimation area, and the automatic station information data set contains the latitude and longitude position coordinates and actual precipitation observation data of each automatic station in the quantitative precipitation estimation area.

[0013] Optionally, constructing a prior distribution model p(M|D) based on the data set M includes:

[0014] In the Dirichlet nonparametric Bayesian model, G0 is the basic probability distribution on the probability measure space Ω, and the concentration parameter α0>0. If the probability distribution G on the probability measure space Ω obeys the basic probability distribution G0, then:

[0015] G~DP(α0,G0)

[0016] Where: the basic probability distribution G0 determines the distribution of the basic components in the prior distribution model p(M|D); DP is the Dirichlet process;

[0017] Based on the Dirichlet process mixture model DPM, a generation probability is added to each data point in the given precipitation estimation area as the prior distribution of the data:

[0018] m i ~p(m|θ i )

[0019] Where: parameter θ i Obeying the probability distribution G, i∈[N] is a set with values ​​ranging from 1 to the total number of data points N, each data point generates a probability parameter, N is the total number of data points; p is the conditional probability density function; m is the precipitation estimate of the data point;

[0020] By comparing different families of prior distribution probability models m i ~p(m|θ i ) is used to select the optimal model as the prior distribution model p(M|D):

[0021] p(M|D) = ∫ Θ (M|Θ)p(Θ|D)

[0022] In the formula, D represents a family of models.

[0023] Optionally, the selecting the optimal model as the prior distribution model p(M|D) by comparing the likelihood functions of different families of models comprises:

[0024] Suppose that the data in the data set M = {m1, m2, m3…m n} are independent, read n automatic station real-time precipitation observation data in the quantitative precipitation estimation area, and the n automatic station real-time precipitation observation data are randomly arranged to obtain {F(i)}, wherein i = 1, 2, … n.

[0025] Suppose Ω = Ω t-1 , for each automatic station real-time precipitation observation data i ∈ {F(1), F(2)…F(n)}, an indicator factor β i of the automatic station real-time precipitation observation data formed based on the n random arrangement groups of the function {F(i)} is calculated. is the centralized parameter at t-1; Ω t-1 is the probability measure space at t-1;

[0026] Based on the current K clusters, the likelihood estimate f k (m i ) of the observation data is calculated.

[0027] f k (m i ) = p(m i |β i = k, M \i , ζ)

[0028]

[0029] In the formula, β i is a new class, M \i is the data corresponding to the index removed from the data set M = {m1, m2, m3…m n}; ζ is a distribution parameter; k is an observation value; m i is a random variable; , similar to k, represents another observation value other than k;

[0030] Sample β i :

[0031]

[0032]

[0033] wherein, is the amount of data already in the kth class; E i is a preset observation data sample; K is the number of observation data samples; f k represents a probability density function for the kth class; δ is a Kronecker delta function, δ(β i =k, δ(β i , k) = 1, otherwise 0; is a probability density function for a class other than the kth class;

[0034] If the number of clusters K is increased by 1, K = K + 1;

[0035] The amount of observation data for each cluster is checked, and if the total number of observation data for a class is 0, the corresponding class is deleted, and the number of clusters K is reduced by 1, K = K - 1.

[0036] The quantitative precipitation estimation system provided by the embodiment of the present application comprises:

[0037] An acquisition module is configured to acquire a data set M;

[0038] A construction module is configured to construct a prior distribution model p(M|D) based on the data set M; wherein D is a model family;

[0039] A definition module is configured to define an approximate distribution q(Θ) of a posterior distribution of precipitation estimation based on the prior distribution model p(M|D) according to variational inference; wherein Θ is a probability model parameter;

[0040] A determination module is configured to determine a lower bound E q [log(p(Θ,M))]-E q [log(q(Θ))] of the precipitation estimation based on the approximate distribution q(Θ) according to the Jensen inequality; wherein E q is an interval of the precipitation estimation distribution;

[0041] A maximization module is configured to maximize the lower bound E q [log(p(Θ,M))]-E q [log(q(Θ))] of the precipitation estimation, and obtain a quantitative precipitation estimation result

[0042] Optionally, the data set M at least comprises a radar data set for quantitatively estimating rainfall and an automatic station information data set, wherein the radar data set is 1Km*1Km radar reflectivity factor grid data in the quantitatively estimated rainfall area, and the automatic station information data set contains longitude and latitude position coordinates of each automatic station in the quantitatively estimated rainfall area and rainfall live observation data.

[0043] Optionally, the constructing the prior distribution model p(M|D) based on the data set M comprises:

[0044] In the Dirichlet non-parametric Bayesian model, G0 is a basic probability distribution on a probability measure space Ω, and the centralized parameter α0>0, and the probability distribution G on the probability measure space Ω is subject to the basic probability distribution G0, and:

[0045] G ~ DP(α0, G0)

[0046] In the formula, the basic probability distribution G0 determines the distribution of the basic component element in the prior distribution model p(M|D); and the DP is a Dirichlet process;

[0047] Based on the Dirichlet process mixture model DPM, a generation probability is added to each data point in the quantitatively estimated rainfall area as a prior distribution of the data:

[0048] m i ~ p(m|θ i )

[0049] In the formula, the parameter θ i is subject to the probability distribution G, i∈[N] is a set of values from 1 to the total number of data points N, a probability parameter is generated for each data point, N is the total number of data points; p is a conditional probability density function; and m is a rainfall estimation value of the data point.

[0050] The optimal model is selected as the prior distribution model p(M|D) by comparing likelihood functions of different families of prior distribution probability models m i ~ p(m|θ i ):

[0051] p(M|D) = ∫ Θ (M|Θ) p(Θ|D)

[0052] In the formula, D represents a family of models.

[0053] Optionally, the selecting the optimal model as the prior distribution model p(M|D) by comparing likelihood functions of different families of models comprises:

[0054] Supposing that the data set M = {m1, m2, m3…m nThe data independent, read quantitative precipitation estimation area within the n automatic station real-time precipitation observation data, n automatic station real-time precipitation observation data random arrangement, get {F(i)}, where i = 1, 2,..., n;

[0055] Let Ω = Ω t-1 , for each automatic station real-time precipitation observation data i ∈ {F(1), F(2)…F(n)}, based on the function {F(i)} formed by n random arrangement of automatic station real-time precipitation observation data indicator factor β i Sampling; wherein, is the centralized parameter at t-1; Ω t-1 is the probability measure space at t-1;

[0056] Based on the current K cluster calculation observation data likelihood estimate f k (m i ),

[0057] f k (m i ) = p(m i |β i = k, M \i , ζ)

[0058]

[0059] In the formula, β i is a new class, M \i is the data corresponding to the index removed from the data set M = {m1, m2, m3…m n}; ζ is the distribution parameter; k is the observation value; m i is a random variable; Similar to k, represents another observation value other than k;

[0060] Sample β i :

[0061]

[0062] In the formula, is the data amount already in the kth class; E i is the preset observation data sample; K is the number of observation data samples; f k represents the probability density function for the kth class; δ is the Kronecker delta function, δ(β i , k) = 1 when β i = k, otherwise 0; is the probability density function for the kth class; δ is the Kronecker delta function, δ(β i , k) = 1 when β i = k, otherwise 0; is the probability density function for the kth class; δ is the Kronecker delta function, δ(β i , k) = 1 when β i = k, otherwise 0;

[0063] If The number of clusters K is increased by 1, K = K + 1.

[0064] The observation data amount of each cluster is checked, and if the total number of observation data of a class is 0, the corresponding class is deleted, and the number of clusters K is reduced by 1, K = K - 1.

[0065] The computer readable storage medium provided by the embodiment of the application, the computer program is stored on the computer readable storage medium, and the processor executes the computer program to realize the method of any one of the above.

[0066] The electronic device provided by the embodiment of the application includes a memory and a processor, the memory stores a computer program, and the processor executes the computer program to realize the method of any one of the above.

[0067] Other features and advantages of the present application will be further described in the following specification, and some will become apparent from the specification, or will be understood by those skilled in the art. The purpose and other advantages of the present application can be achieved and obtained by the structure specifically pointed out in the written specification and the drawings.

[0068] The technical solutions of the present application will be further described in detail below with the help of the drawings and examples. BRIEF DESCRIPTION OF DRAWINGS

[0069] The accompanying drawings are used to provide a further understanding of the present application, and constitute a part of the specification, and are used to explain the present application together with the embodiments of the present application, and do not constitute a limitation on the present application. In the drawings:

[0070] Figure 1 It is a schematic diagram of a quantitative precipitation estimation method in the embodiment of the present application;

[0071] Figure 2 It is a schematic diagram of a quantitative precipitation estimation system in the embodiment of the present application. DETAILED DESCRIPTION

[0072] The preferred embodiments of the present application will be described below in conjunction with the drawings, and it should be understood that the preferred embodiments described herein are only used to illustrate and explain the present application, and do not limit the present application.

[0073] The embodiment of the present application provides a quantitative precipitation estimation method, as shown in the figure, which comprises: Figure 1

[0074] S1, obtaining a data set M;

[0075] S2, constructing a prior distribution model p(M|D) based on the data set M; wherein D is a family of models;​

[0076] S3, determining an approximate distribution q(Θ) of the precipitation estimation posterior distribution based on the variational inference according to the prior distribution model p(M|D); wherein, Θ is a probability model parameter;

[0077] S4, determining a lower bound E q [log(p(Θ,M))]-E q [log(q(Θ))] of the precipitation estimation based on the Jensen inequality according to the approximate distribution q(Θ); wherein, E q is a data distribution interval of the precipitation estimation;

[0078] S5, maximizing the lower bound E q [log(p(Θ,M))]-E q [log(q(Θ))] of the precipitation estimation to obtain a quantitative precipitation estimation result

[0079] The data set M at least includes: a radar data set of a quantitative precipitation estimation area and an automatic station information data set, wherein the radar data set is 1Km*1Km radar reflectivity factor grid data in the quantitative precipitation estimation area, and the automatic station information data set contains longitude and latitude position coordinates and precipitation real-time observation data of each automatic station in the quantitative precipitation estimation area.

[0080] The prior distribution model p(M|D) is constructed based on the data set M, and the construction includes:

[0081] In the Dirichlet non-parametric Bayesian model, G0 is a basic probability distribution on a probability measure space Ω, and the centralized parameter α0>0. If a probability distribution G on the probability measure space Ω is subject to the basic probability distribution G0, then:

[0082] G~DP(α0,G0)

[0083] In the formula, the basic probability distribution G0 determines the distribution of the basic component element in the prior distribution model p(M|D); DP is a Dirichlet process; wherein, the Dirichlet process is a random process widely used in Bayesian non-parametric statistics, and is used for modeling the precipitation distribution probability in the case that the specific form of the precipitation distribution is unknown but some prior knowledge is known in the method;

[0084] Based on the Dirichlet process mixture model DPM, a generation probability is added to each data point in the quantitative precipitation estimation area as a prior distribution of the data:

[0085] m i ~p(m|θ i )

[0086] Where: parameter θ i Obeying the probability distribution G, i∈[N] is a set of values ​​ranging from 1 to the total number of data points N, each data point generates a probability parameter, and N is the total number of data points; p is the conditional probability density function (PDF); m is the precipitation estimate at the data point; this expression represents the probability distribution of the precipitation estimate data under the given parameters;

[0087] By comparing different families of prior distribution probability models m i ~p(m|θ i ) is used to select the optimal model as the prior distribution model p(M|D):

[0088] p(M|D)=∫ Θ (M|Θ)p(Θ|D)

[0089] Where: D represents a family of models.

[0090] The method of selecting the optimal model as the prior distribution model p(M|D) by comparing the likelihood functions of different families of models includes:

[0091] Assume that the data set M = {m1,m2,m3…m n} data is independent, read the actual precipitation observation data of n automatic stations in the quantitative precipitation estimation area, randomly arrange the actual precipitation observation data of n automatic stations to obtain {F(i)}, where i = 1, 2, ... n; randomly arrange the obtained precipitation observation data to obtain observation data with different arrangement methods;

[0092] set up Ω=Ω t-1 , for each automatic station real-time precipitation observation data i∈{F(1),F(2)…F(n)}, the indicator factor β of the automatic station real-time precipitation observation data composed of n random permutations based on the function {F(i)} i Sampling is performed; among them, is the concentrated parameter at time t-1; Ω t-1 is the probability measurement space at time t-1; for the actual precipitation observation data of each automatic station, a random arrangement is generated based on a specific function, and the corresponding indicator factor is sampled to form the corresponding indicator factor. The indicator factor represents the centralized parameter (such as meteorological parameters, precipitation value, etc.) at the current moment, which includes the probability measurement space at the moment.

[0093] Calculate the likelihood estimate f of the observed data based on the current K clusters k (m i ),

[0094] fk (m i )=p(m i |β i =k,M \i ,ζ)

[0095]

[0096] Where, β i For a new category, M \i For the dataset M={m1,m2,m3…m n} removes the data corresponding to the subscript (the unique index corresponding to each data point in the data set) (that is, k in the formula); ζ is the distribution parameter; k is the observation value; m i is a random variable; Similar to k, it represents another observation value other than k; for each current cluster, the likelihood estimate of its observed data is calculated. That is, based on the current dataset and clustering model, the degree of fit of the data is evaluated. By setting the number of clusters and sample data, different models are gradually fitted to calculate the likelihood function of each cluster. The relevant parameters in the calculation formula, such as the clustering parameters, observation values, and random variables of the data, are finally obtained to obtain the likelihood function value of each cluster. By comparing the likelihood functions of different families of models, the model with the highest likelihood value is selected as the optimal model and used as the prior distribution model.

[0097] β i To take a sample:

[0098]

[0099]

[0100] Where, is the amount of data already in the kth class; E i is the preset observation data sample; K is the number of observation data samples; f k represents the probability density function for the kth class; δ is the Kronecker δ function, when β i =k, δ(β i , k) = 1, otherwise 0; To represent the probability density function except for the kth class;

[0101] if Then increase the number of clusters K by 1, K = K + 1;

[0102] Check the amount of observation data for calculating the likelihood function of each type of clustering. If the total number of observation data for a category is 0, delete the corresponding category and reduce the number of clusters K by 1, K = K-1.

[0103] In the clustering process, sampling is performed to determine whether the number of clusters needs to be adjusted. If the data volume in a certain cluster is 0, the cluster is deleted and the number of clusters is reduced. The adjustment of the number of clusters can improve the estimation accuracy and ensure that each cluster can effectively represent the distribution of the data. When the number of clusters reaches the preset condition, the number of clusters is increased by 1 to further refine the classification. The calculation and inspection are performed on all clusters to ensure that the total number of observation data of the likelihood function of each cluster is appropriate. In the data sampling process, the Kronecker function is used to represent the relationship between different data categories. If the number of clusters still does not meet the condition after adjustment (such as a certain category being empty), the category needs to be deleted and the number of clusters is reduced. By dynamically adjusting the number of clusters, it can be ensured that the clustering result of the precipitation data is more consistent with the distribution of the actual data. After multiple iterations, when the number of clusters and the likelihood function reach the preset accuracy, the optimal precipitation estimation result is finally output.

[0104] The working principle and beneficial effects of the above technical solutions are:

[0105] Bayesian machine learning is an important branch of machine learning. The Bayesian method was first proposed by British mathematician Thomas Bayes. After more than two hundred years of development, it has become an important part of machine learning and is widely used in many statistical machine learning fields such as multivariate structured output prediction. The basic concept of Bayes' theorem is to give the prior distribution and likelihood function of the model, and the posterior probability distribution p(Θ|M) of the model can be obtained by Bayes formula:

[0106]

[0107] In formula (1), Θ is the probability model parameter, M is the data set, which in this application includes the radar data set and the automatic station information data set, p0(Θ) is the prior distribution function of the model, p(M|Θ) is the likelihood function, and p(M) is a constant.

[0108] The selection of the model is a basic problem of the Bayesian method. The Dirichlet non-parametric Bayesian model can automatically learn the number of cluster centers by the characteristics of the data itself.

[0109] The Dirichlet non-parametric Bayesian model assumes that G0 is a random probability distribution on the probability measure space Ω, and the concentration parameter α0>0. If the probability distribution G on the space Ω is subject to the basic probability distribution G0, then:

[0110] G ~ DP(α0, G0) (2)

[0111] In the formula, the basic distribution G0 determines the distribution of the basic elements in the model.

[0112] But the probability distribution obtained by Dirichlet process is discrete, in order to cluster different groups of data with certain similarity, Dirichlet process mixture model (DPM) is introduced (Antoniak. 1974), by adding a generating probability to each data point as the prior distribution of data:

[0113] m i ~p(m|θ i )(3)

[0114] In formula (3), parameter θ i obeys G distribution, i∈[N] is a generating probability added to each data point, N is the number of data points. When G obeys Dirichlet process distribution, the model is called Dirichlet process mixture model. Bayesian method selects the optimal model by comparing the likelihood functions of different families of models:

[0115] p(M|D)=∫ Θ (M|Θ)p(Θ|D) (4)

[0116] In formula (4), D represents a family of models, in the case of no obvious prior function, it is assumed that p(Θ|D) is a uniform distribution, by integral of formula (4), model over-fitting can be avoided.

[0117] In the research, it is assumed that the data of observation set M={m1,m2,m3…m n} are independent, in order to obtain the indicator factor of each observation, in the non-parametric Bayesian model using Dirichlet process as prior distribution, Gibbs sampling is used to obtain the likelihood function of different families of models to select the optimal model, the steps are as follows:

[0118] (1) Data initialization: the system reads data, n observation data are randomly arranged, and {F(i)} is obtained, wherein i=1,2,…n.

[0119] (2) Class clustering: set Ω=Ω t-1 , for each observation data i∈{F(1),F(2)…F(n)}, the indicator factor β i of each data is sampled.

[0120] The likelihood estimate of the observation data based on the existing K clusters is calculated:

[0121] f k (m i )=p(m i |β i =k,M \i, ζ) (5)

[0122]

[0123] β i is a new class, M \i represents removing the data with the corresponding subscript from the corresponding observation dataset; ζ is a distribution parameter.

[0124] β i is sampled:

[0125]

[0126] β is the amount of data already in the kth class. If , the number of clusters is increased by one, K = K + 1.

[0127] (3) Cluster update: check the observation data amount of each class. If the total number of observation data of a certain class is 0, delete the class, and the number of clusters is reduced by one, K = K - 1.

[0128] The inference method of the Bayesian model is an important content in Bayesian learning. Given the prior distribution, the posterior distribution of the Bayesian model is usually unsolvable, and an effective inference method is needed. In this application, variational inference is used for prediction.

[0129] In quantitative estimation of precipitation prediction, given the data set M, the prior distribution model p(M|D), the variational method is used to define the approximate distribution q(Θ) of the precipitation estimation posterior distribution. Using the Jensen inequality, a lower bound of the precipitation estimation can be obtained:

[0130] logp(M)≥E q [log(p(Θ,M))]-E q [log(q(Θ))] (9)

[0131] By maximizing the lower bound of the estimation value:

[0132]

[0133] The solution of quantitative precipitation estimation can be completed.

[0134] To perform quantitative precipitation estimation, first, some actual data related to precipitation amount need to be collected. The dataset consists of a radar dataset and an automatic station information dataset. The radar dataset contains radar reflectivity factor data within the precipitation area. The spatial resolution of the radar data is 1 km*1 km. These radar reflectivity data can be used to estimate the distribution of precipitation, because there is a certain relationship between radar signals and precipitation amount. The automatic station information dataset contains automatic weather station data within the quantitative precipitation estimation area, recording the latitude and longitude information of each station and the precipitation observation data at the corresponding time point. These live observation data serve as a reference for the actual precipitation amount, playing a verification and auxiliary estimation role.

[0135] In order to be able to infer accurate precipitation estimation values from the collected data, a prior distribution model needs to be constructed, the purpose of which is to set the prior information of precipitation according to existing data and background knowledge. Dirichlet process (DP) is a non-parametric Bayesian model that can be used to model the distribution of an unknown number of populations. In quantitative precipitation estimation, Dirichlet process is used to model different types of precipitation processes. Through Dirichlet process, a mixture model can be constructed to divide the data into several categories and assign a probability to each category. In this application, DPM is used to establish a prior distribution model of precipitation. Specifically, the distribution of data points can be generated from a mixture model, and the parameters of each mixture component are unknown. Therefore, the Dirichlet process is used to define the distribution, thereby providing prior distribution information for precipitation estimation. In DPM, a prior distribution is used to describe the distribution form of precipitation amount. By constructing a prior distribution, the model can be updated when new observation data is available, thereby gradually approaching the true precipitation distribution.

[0136] After establishing the prior distribution model, posterior inference is performed. Since in actual situations, the posterior distribution cannot be directly calculated, it needs to be approximated by the method of variational inference. Variational inference is an optimization method used to solve approximate solutions of complex probability distributions. In the problem of precipitation estimation, the purpose of variational inference is to approximate the complex posterior distribution by constructing a simple approximate distribution. Specifically, this application uses variational inference to infer the posterior distribution of precipitation from the prior distribution. During the inference process, the goal of variational inference is to optimize a set of parameters to minimize the KL divergence of the approximate posterior distribution, so as to approximate the true posterior distribution as accurately as possible.

[0137] In order to further optimize the estimation result and ensure the effectiveness of the model under the given data, the present application introduces Jensen's Inequality to calculate the lower bound. Jensen's Inequality is a mathematical tool in Bayesian inference, which provides a method to approximate the objective function by optimizing the lower bound. In the present application, Jensen's Inequality is used to establish the lower bound of precipitation estimation. By calculating the lower bound, it can be ensured that the estimation of the posterior distribution will not deviate too far from the true value, and the maximization of the lower bound also enables the optimal precipitation estimation result to be found in the multiple solution space.

[0138] After calculating the lower bound, the process of maximizing the lower bound is carried out. By maximizing the lower bound, a quantitative precipitation estimation value as accurate as possible can be obtained. Specifically, for a given data set, based on variational inference and Jensen's Inequality, an objective function is obtained, which represents the size of the lower bound. By optimizing the objective function, the lower bound is maximized, and finally the optimal precipitation estimation result is obtained. It is often necessary to estimate the likelihood function of the data points and select the optimal model for precipitation estimation. The likelihood function describes the probability of observing the data given the model parameters. In precipitation estimation, the likelihood function is used to calculate the difference between the precipitation data and the model prediction. By maximizing the likelihood function, the precipitation can be more accurately estimated. In actual data, there are often many different types of precipitation, therefore, using clustering methods to group the data can better understand the distribution characteristics of different types of precipitation, and more accurate estimation values can be obtained through sampling calculation.

[0139] Through the above steps, the quantitative precipitation estimation result can be finally obtained. The whole process combines Bayesian inference, variational inference, Jensen's Inequality and other methods to ensure the accuracy and reliability of the estimation result. The model continuously optimizes and approximates the true precipitation through prior distribution model, variational inference and lower bound maximization, and finally realizes high-precision precipitation estimation. The present application has a wide application prospect in the field of meteorology, especially in the aspects of extreme weather warning and climate research.

[0140] The quantitative precipitation estimation method and system of the present application are based on long-term massive observation data, and use a quantitative precipitation estimation algorithm based on non-parametric Bayesian machine learning. The algorithm combines radar reflectivity factor, automatic station information, radar inversion wind field, weather type and other 9 elements to constrain the radar quantitative precipitation estimation algorithm. First, a radar data set of the quantitative precipitation estimation area and an automatic station real-time observation information data set are constructed, a prior distribution model is constructed based on the data set, an approximate distribution of the posterior distribution of the precipitation estimation is defined based on the prior distribution model according to the variational inference, the lower bound of the precipitation estimation is determined based on the approximate distribution according to Jensen's Inequality, the lower bound of the precipitation estimation is maximized, and the quantitative precipitation estimation result is obtained, so as to carry out quantitative precipitation estimation and improve the performance of its business application.

[0141] The embodiment of the present invention provides a quantitative precipitation estimation system, such as Figure 2 Shown, including:

[0142] Acquisition module 1, used to obtain data set M;

[0143] Construction module 2 is used to construct a prior distribution model p(M|D) based on the data set M; where D is a family of models;

[0144] A definition module 3 is used to define an approximate distribution q(Θ) of the posterior distribution of precipitation estimation based on the prior distribution model p(M|D) based on variational inference; where Θ is a probability model parameter;

[0145] Determination module 4 is used to determine the lower bound E of precipitation estimation based on the approximate distribution q(Θ) based on Jason inequality q [log(p(Θ,M))]-E q [log(q(Θ))]; where E q is the interval of the estimated precipitation distribution;

[0146] Maximize module 5, which is used to estimate the lower bound E of precipitation q [log(p(Θ,M))]-E q [log(q(Θ))] is maximized to obtain quantitative precipitation estimation results

[0147] The data set M includes at least: a radar data set of the quantitative precipitation estimation area and an automatic station information data set, wherein the radar data set is 1km*1km radar reflectivity factor grid data within the quantitative precipitation estimation area, and the automatic station information data set includes the latitude and longitude position coordinates and actual precipitation observation data of each automatic station within the quantitative precipitation estimation area.

[0148] The prior distribution model p(M|D) is constructed based on the data set M, including:

[0149] In the Dirichlet nonparametric Bayesian model, G0 is the basic probability distribution on the probability measure space Ω, and the concentration parameter α0>0. If the probability distribution G on the probability measure space Ω obeys the basic probability distribution G0, then:

[0150] G~DP(α0,G0)

[0151] Where: the basic probability distribution G0 determines the distribution of the basic components in the prior distribution model p(M|D); DP is the Dirichlet process;

[0152] Based on a Dirichlet process mixture model DPM, a generation probability is added to each data point in a quantitative precipitation estimation area as a prior distribution of data:

[0153] m i ~p(m|θ i )

[0154] In the formula, the parameter θ i obeys a probability distribution G, i∈[N] is a set of values from 1 to the total number of data points N, a probability parameter is generated for each data point, N is the total number of data points; p is a conditional probability density function; m is a precipitation estimation value of the data point;

[0155] By comparing the likelihood functions of different families of prior distribution probability models m i ~p(m|θ i ), the optimal model is selected as the prior distribution model p(M|D):

[0156] p(M|D)=∫ Θ (M|Θ)p(Θ|D)

[0157] In the formula, D represents a family of models.

[0158] The optimal model is selected as the prior distribution model p(M|D) by comparing the likelihood functions of different families of models, comprising:

[0159] Suppose that the data set M={m1, m2, m3…m n} is independent, n automatic station real-time precipitation observation data in a quantitative precipitation estimation area are read, the n automatic station real-time precipitation observation data are randomly arranged, and {F(i)} is obtained, wherein i=1, 2,…n;

[0160] Suppose Ω=Ω t-1 , for each automatic station real-time precipitation observation data i∈{F(1), F(2)…F(n)}, an indicator factor β i of the automatic station real-time precipitation observation data formed based on the n random arrangement groups of the function {F(i)} is calculated; is a centralized parameter at t-1; Ω t-1 is a probability measure space at t-1;

[0161] Based on the current K clusters, the likelihood estimation f k (m i ) of the observation data is calculated,

[0162] f k (m i )=p(m i |βi = k, M \i , ζ)

[0163]

[0164] where β i is a new class, M \i is the data removed from the dataset M = {m1, m2, m3…m n} with the corresponding index; ζ is the distribution parameter; k is the observation; m i is a random variable; Similar to k, represents another observation other than k;

[0165] Sampling β i :

[0166]

[0167] where is the amount of data already in the kth class; E i is the preset observation data sample; K is the number of observation data samples; f k represents the probability density function for the kth class; δ is the Kronecker delta function, δ(β i , k) = 1 when β i = k, and 0 otherwise; is the probability density function for the classes other than the kth class;

[0168] If , the number of clusters K is increased by 1, K = K + 1;

[0169] The amount of observation data for which the cluster likelihood function is calculated is checked, and if the total amount of observation data for a class is 0, the corresponding class is deleted, and the number of clusters K is reduced by 1, K = K - 1.

[0170] The computer readable storage medium provided by the embodiment of the application has a computer program stored thereon, and a processor executes the computer program to implement the method of any one of the above.

[0171] The electronic device provided by the embodiment of the application includes a memory and a processor, the memory has a computer program stored therein, and the processor executes the computer program to implement the method of any one of the above.

[0172] Obviously, those skilled in the art can make various modifications and variations to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and equivalent technologies thereof, the present application is also intended to include these modifications and variations.

Claims

1. A quantitative precipitation estimation method, characterized by, Comprising: Acquiring a dataset ; Based on a dataset , constructing a prior distribution model ; wherein, is a family of models; based on variational inference, according to a prior distribution model , defining an approximate distribution of a posterior distribution of precipitation estimates ; wherein, are parameters of a probabilistic model; Based on the Jensen inequality, a lower bound of the precipitation estimate is determined from the approximate distribution ; wherein is the interval of the precipitation estimate data distribution;​ Lower bound on precipitation estimates Maximization is performed to obtain quantitative precipitation estimates ; The data set-based , constructing a prior distribution model , comprising: Dirichlet nonparametric Bayesian model is the probability measure space The basic probability distribution on , the centralized parameter , probability measure space The probability distribution on If it obeys the basic probability distribution ,but: where: base probability distribution determining a prior distribution model distribution of the basic constituent elements; is a Dirichlet process; Based on Dirichlet process mixture model DPM, a generating probability is added to each data point in the estimation area of given amount precipitation as the prior distribution of data: where: parameters subject to a probability distribution , is a set of values from 1 to the total number of data points , each data point generating a probability parameter, is the total number of data points; is a conditional probability density function; is the precipitation estimate for the data point; By comparing likelihood functions of different family prior distribution probability models to select an optimal model as a prior distribution model : wherein: denotes a family of models; The optimal model is selected as the prior distribution model by comparing likelihood functions of different family prior distribution probability models , comprising:​ Set data independent of the data, read the quantitative precipitation estimation area of automatic station real-time precipitation observation data, automatic station real-time precipitation observation data randomly arranged, obtained , wherein ; Set , , for each automatic station live precipitation observation data , based on the function formed randomly arranged group of automatic station live precipitation observation data indicating factor sampling; wherein, is concentration parameters at time; is probability measure space at time; Based on the current likelihood estimate of the observation data , : where is a new class, is the data set with the corresponding index removed; is the distribution parameter; is the observation; is a random variable; and similarly, denotes another observation not of. right To take a sample: wherein is the number of classes is the amount of data already in the class is a predetermined observation data sample is the number of observation data samples denotes the probability density function for the class is the Kronecker function, when = 1, otherwise 0 = 1, otherwise 0 = 1, otherwise 0 = 1, otherwise 0 denotes the probability density function for the classes other than the class is the number of classes If = 0 , then the number of clusters is increased by 1, ; Check the observation data amount of each type of clustering likelihood function, if the total number of observation data of a type is 0, delete the corresponding category, the number of clusters is reduced by 1, .

2. The quantitative precipitation estimation method of claim 1, wherein, The dataset It at least includes: a radar data set of the quantitative precipitation estimation area and an automatic station information data set, wherein the radar data set is the 1km*1km radar reflectivity factor grid data within the quantitative precipitation estimation area, and the automatic station information data set contains the latitude and longitude position coordinates and actual precipitation observation data of each automatic station in the quantitative precipitation estimation area.

3. A quantitative precipitation estimation system, comprising: Comprising: An acquisition module is configured to acquire a data set ; a construction module for constructing a model based on a data set , a prior distribution model ; wherein, is a family of models; a bounding module configured to bound, based on variational inference, an approximate distribution of a posterior distribution of the precipitation estimate according to a prior distribution model ; wherein are parameters of the probabilistic model​ determining module configured to determine, based on Jensen's inequality, a lower bound of the precipitation estimate from the approximate distribution wherein is an interval of the precipitation estimate distribution;​ maximizing module for maximizing a lower bound of a precipitation estimate maximizing to obtain a quantitative precipitation estimate ; Based on the dataset , constructing a prior distribution model , comprising: In a Dirichlet nonparametric Bayesian model is a base probability distribution over a probability measure space , a parametric probability distribution over a probability measure space if it is drawn from a base probability distribution then:​​ where: base probability distribution determining a prior distribution model distribution of the basic constituent elements; is a Dirichlet process; Based on Dirichlet process mixture model DPM, a generating probability is added to each data point in the estimation area of given amount precipitation as the prior distribution of data: where: parameters subject to a probability distribution , is a set of values from 1 to the total number of data points , each data point generating a probability parameter, is the total number of data points; is a conditional probability density function; is the precipitation estimate for the data point; By comparing likelihood functions of different family prior distribution probability models to select an optimal model as a prior distribution model : wherein: denotes a family of models; selecting an optimal model as a prior distribution model by comparing likelihood functions of different family models comprising: Set data independent of data, read quantitative precipitation estimation area of automatic station real-time precipitation observation data, automatic station real-time precipitation observation data randomly arranged, get , wherein ; Set , , for each automatic station live precipitation observation data , based on the function formed randomly arranged into an automatic station live precipitation observation data indicating factor sampling; wherein, is concentration parameters at the moment; is probability measure space at the moment; Based on the current likelihood estimate of the observation data , : where is a new class, is the data set with the data corresponding to the index removed; is the distribution parameter; is the observation; is a random variable; and denotes another observation not equal to right To take a sample: wherein is the number of classes; is the amount of data already in the class; is a predetermined observation data sample; is the number of observation data samples; denotes the probability density function for the class ; is the Kronecker function, when = , and = 1, otherwise 0; denotes the probability density function for all classes except the class ; If = 0 , then the number of clusters is increased by 1, ; Check the observation data amount of each type of clustering likelihood function, if the total number of observation data of a type is 0, delete the corresponding category, the number of clusters is reduced by 1, .

4. The quantitative precipitation estimation system of claim 3, wherein, The data set At least includes: a radar data set of quantitative precipitation estimation area and an automatic station information data set, wherein the radar data set is 1Km*1Km radar reflectivity factor grid data in the quantitative precipitation estimation area, and the automatic station information data set contains longitude and latitude position coordinates and precipitation live observation data of each automatic station in the quantitative precipitation estimation area.

5. A computer readable storage medium, characterized in that, The computer readable storage medium stores a computer program, and the processor executes the computer program to realize the method in any one of claims 1-2.

6. An electronic device, comprising: The electronic device includes a memory and a processor, the memory stores a computer program, and the processor executes the computer program to realize the method in any one of claims 1-2.

Citation Information

Patent Citations

  • Runoff probability forecasting method

    CN109344999A

  • Multi-source rainfall data fusion algorithm and device based on Bayesian regression

    CN112612995A