Station network optimization method based on high-dimensional Copula entropy and Kriging
By constructing the C-Vine Copula tree structure and maximum likelihood estimation method, combining multivariate mutual information and sliding window method, the problem of probability density function characterization in high-dimensional rainfall site network optimization is solved, and the optimization of total information and estimation error is achieved, and the dynamic site network optimization is adapted to climate change is achieved.
Patent Information
- Application Number
- CN202210040869.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-14
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2042-01-14
AI Technical Summary
The traditional rainfall site optimization method cannot effectively characterize the probability density function in high-dimensional situations, and cannot evaluate the dynamic characteristics of the site optimization results caused by climate change. The existing Copula entropy method lacks the accuracy of the high-dimensional redundant information portraying amount, so it cannot adapt to the nonlinear characteristics of the rainfall sequence.
The website optimization method based on high-dimensional Copula entropy and Krigin is adopted to construct a C-Vine Copula tree structure, and the parameters are estimated using the maximum likelihood estimation method, combining multivariate mutual information and high-dimensional C-Vine Copula density function relationship, and the dynamic rainfall website optimization is performed using standardized MiK-MiT-MaJ index and sliding window method.
The precise simulation of the dependency structure between high-dimensional sites is realized, the total amount of information and estimation errors are optimized, and the dynamic site network optimization can be carried out in the context of climate change, improving site optimization efficiency and accuracy.
Smart Images

Figure CN114595556B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for optimizing a rain gauge network, and particularly to a method for optimizing a rain gauge network based on high-dimensional Copula entropy and Kriging. Background Art
[0002] A hydrometeorological station is a grass-roots hydrological agency established on a river or a basin, mainly used for observing and collecting hydrological and meteorological data related to water bodies such as rivers, lakes and reservoirs. Through the complete collection and control of the measured data in the early stage, it provides sufficient data support for the later work of exploring the basic hydrological laws, and largely meets the basic needs of hydrological forecasting, hydrological information, water resources evaluation work and water science research. Therefore, a reasonable hydrological station network can fully reflect the hydrological spatio-temporal variation characteristics, so that it can collect accurate and detailed hydrological information. It is necessary to explore a more objective theoretical method to support the reasonable planning of the hydrological station network.
[0003] However, traditional information entropy methods or geostatistical methods, on the one hand, one-sidedly consider the rainfall information amount or the areal rainfall estimation error, without jointly considering the optimal information amount and areal rainfall estimation accuracy; on the other hand, based on the premise of the assumption of no trend in the rainfall sequence, the obtained optimized results of the rain gauge network are limited to a specific historical stage and cannot cope with the dynamic characteristics of the optimized results of the rain gauge network caused by the trend of the rainfall sequence under the background of climate change.
[0004] At present, there is also a method for optimizing a rain gauge network using Copula entropy, mainly with the Archimedean Copula function as the core, and there are certain deficiencies in the characterization accuracy of high-dimensional redundant information amount. Especially due to the increase in dimension, the existence of non-linearity makes the dependence structure of the sequence unable to match the symmetry of the Archimedean Copula function, resulting in a large fitting error. Summary of the Invention
[0005] Objects of the Invention: The present invention aims to provide a method for optimizing a rain gauge network based on high-dimensional Copula entropy and Kriging, to solve the difficulty and deficiency of traditional rain gauge network evaluation methods in characterizing the probability density function in high-dimensional cases, and to evaluate the dynamic characteristics of the optimized results of the rain gauge network caused by the time-varying characteristics of the rainfall sequence affected by climate change.
[0006] Technical Solution: The method for optimizing a rain gauge network based on high-dimensional Copula entropy and Kriging according to the present invention includes the following steps:
[0007] (1) Construct a C-Vine Copula tree structure of the hydrological station network;
[0008] (2) Estimate the C-Vine Copula parameters by the maximum likelihood estimation method;
[0009] (3) Obtain the multivariate mutual information through the functional relationship between the multivariate mutual information and the high-dimensional C-Vine Copula density;
[0010] (4) Optimize the dynamic rain gauge network by normalizing the MiK-MiT-MaJ index and the sliding window method.
[0011] Step (1) includes the following steps:
[0012] (11) Calculate the pairwise dependence parameters of all variables The conditional set D of Tree L is an empty set;
[0013] (12) Through the traversal method, select the nodes with the largest sum of dependence parameters for all nodes to form a tree;
[0014] (13) Select the nodes {j, k|D} on Tree L to form edges, and calculate the corresponding C-Vine Copula type and parameters, where L = 1,..., 3(d - 1), d is the number of stations, and j and k are tree node numbers;
[0015] (14) Obtain the pseudo-observations according to step (13) and represents traversing the edges of each tree;
[0016] (15) The tree node j = j + 1, and repeat steps (11) to (14) until the entire tree structure is determined.
[0017] Step (2)
[0018] The maximum likelihood function is:
[0019]
[0020] where, [x1, x2,..., x d represents the variable combination of the network G composed of d stations, d is the number of stations in the initial network; x d represents the t-th observation value of the station numbered i, t = 1, 2,..., n; n is the size of the data samples collected at each station; Θ is the C-Vine Copula parameter set, c it represents the two-dimensional Pair Copula density; F(x j,j+1|1,...,j-1 |x jt x 1t,..., x (j-1)t ) is the conditional joint distribution function, that is, the joint distribution function given the variable set {x1, x2,...x j-1};
[0021] Solve the maximum likelihood function to obtain the parameter Θ
[0022] Θ = arg max ln L(Θ) (2).
[0023] Step (3)
[0024] The multivariate mutual information TC is obtained from equations (3) and (4).
[0025] TC(x1, x2,..., x d ) = -H c (u1, u2,..., u d ) (3)
[0026] H c (u1, u2,..., u d ) = -∫c(u1, u2,..., u d ) log c(u1, u2,..., u d ) dU
[0027] = -E[log c(u1, u2,..., u d )] (4)
[0028] where d represents the number of stations in the station combination, and u i represents the marginal distribution function of the i-th station; G d represents the station combination composed of d stations.
[0029] Step (4) includes the following steps:
[0030] (41) Dynamically analyze the station network optimization result from the perspective of the trend of the rainfall sequence using the sliding window method; divide the original station network sequence G d into m station network subsequences according to different window widths where the station network subsequence and the original station network sequence G d only differ in the length of the rainfall sequence, and the number of stations is still d stations;
[0031] (42) Convert the multi-objective optimization problem into a single-objective optimization using the standardized MiK-MiT-MaJ index, and for the station network subsequence obtain the optimal station combination of c stations selected from d stations by maximizing Stand MKTJ Stand
[0032] Stand MKTJ is
[0033]
[0034]
[0035]
[0036]
[0037] Based on the principle of minimizing the KSE value as the optimization objective function, the KSE is as follows:
[0038]
[0039] where h 0j is the distance between any one site position in the set of G c (c sites) and any other point outside G c , and γ(h 0j ) is the mutation value at the distance h 0j , μ x is the Lagrange multiplier, N is the number of sites with unknown spatial distribution, and w j is the weight value;
[0040] The joint information entropy value JE is:
[0041]
[0042] where p(x1,x2,…,x c ) is the joint probability density function under c sites;
[0043] (43) Select c site combinations from the obtained d sites to generate k different site optimization result datasets, and calculate the occurrence frequency of each site according to the number of times each site appears:
[0044]
[0045] where is the frequency of the site x i entering the optimal site combination, and finally the optimal frequency of each site can be obtained
[0046] Beneficial effects: Compared with the prior art, the present invention has the following remarkable advantages:
[0047] (1) Using the C-Vine Copula function to more accurately simulate the high-dimensional dependence structure between multiple sites to obtain a Copula entropy value that is more accurate than the traditional two types of Copula functions, thereby realizing the two objective functions of the station network optimization, the total information volume (JE) and the total correlation quantity (TC); using the Kriging standard error value commonly used in geostatistics as the key objective function to achieve the dual purposes of the optimal estimation error and the optimal rainfall information of the rainfall station network.
[0048] (2) The dynamic optimization of the rainfall station network is carried out by standardizing the MiK-MiT-MaJ index and using the sliding window method to consider the time-varying characteristics of rainfall. On the one hand, the standardized MiK-MiT-MaJ index simplifies the multi-objective optimization problem into a single-objective optimization, improving the efficiency of site optimization; on the other hand, the sliding window method effectively considers the dynamic characteristics of the station network optimization results caused by the time-varying characteristics of the rainfall sequence under the background of climate change. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] Figure 1 It is a schematic flow chart in the present invention;
[0050] Figure 2 It is a tree structure diagram of the present invention;
[0051] Figure 3 It is the original site layout diagram of the present invention;
[0052] Figure 4 It is the comparison of the total correlation differences calculated by the C-Vine Copula function and the Archimedean Copula in the present invention;
[0053] Figure 5 It is the analysis diagram of the mutual information estimation error of the histogram method and the Copula entropy method under different correlation coefficients in the present invention;
[0054] Figure 6 It is the calculation results of the post-evaluation indexes MNCE and NSC in the embodiment of the present invention;
[0055] Figure 7 It is the station network optimization result based on a sliding window width of 10 in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0056] The technical solution of the present invention will be further described below with reference to the accompanying drawings.
[0057] For the convenience of understanding the present invention, the following description is made:
[0058] Entropy is a measure of the uncertainty of a random variable in statistics. Let x be a discrete random variable with a value space of U and a probability density function of p(x). Its information entropy H(x) is defined as
[0059]
[0060] For the case of two variables, the amount of information transfer I(x i , x j ) can be calculated by mutual information:
[0061] I(x i , x j ) = H(x i ) + H(xj ) - H(x i , x j )
[0062] According to Sklar's theorem, the d - dimensional Copula is as follows:
[0063]
[0064] where P represents the multi - dimensional cumulative distribution function, and is the marginal probability distribution function of the corresponding variable x i . θ is the Coupla parameter. From this, the Copula density function is deduced as follows. The relationship between the two - dimensional Copula entropy and mutual information is as follows:
[0065]
[0066] I(x i , x j ) = - H c (u i , u j )
[0067] From the above formula, it can be obtained that the two - dimensional Copula entropy is the negative value of mutual information, and this conclusion is extended to the high - dimensional case, that is:
[0068] TC(x1, x2,..., x d ) = - H c (u1, u2,..., u d )
[0069] where TC(x1, x2,…, x d ) represents the TC value of the d - dimensional variable.
[0070] Therefore, the available high - dimensional C - Vine Copula entropy value can be used to deduce the information redundancy under multiple stations, which can be regarded as a good combination of the Copula function and information entropy.
[0071] As Figure 1 shown, the station network optimization method based on high - dimensional Copula entropy and Kriging of the present invention includes the following steps:
[0072] Step 1. Determine the C - Vine Copula tree structure of the hydrological station network.
[0073] First, for a hydrological station network composed of d stations, there are many possible C - Vine Copula pair structures. According to the "strongest dependence" criterion, the Kendall's tau value is used to optimize the tree structure.
[0074] For tree Tree L, where L = 1, …, 3(d - 1), the specific steps can be refined as follows:
[0075] Step 1: Calculate the pairwise dependence parameters of all variables Note that for Tree L, the conditional set D is an empty set;
[0076] Step 2: Through the traversal method, for all nodes, select the nodes with the largest sum of dependence parameters to form a tree;
[0077] Step 3: On Tree L, select nodes {j, k|D} (L = 1, …, 3(d - 1)) to form edges, and calculate the corresponding C-Vine Copula type and parameters, where d is the number of stations, and j, k are tree node numbers;
[0078] Step 4: Obtain the pseudo-observations according to step (3) and where i represents traversing the edges of each tree;
[0079] Step 5: Let the tree node j = j + 1, and repeat steps 1 to 4 until the entire tree structure is determined.
[0080] In this embodiment, taking d = 4 as an example, the determination process of the C-Vine Copula tree structure is illustrated.
[0081] Calculate the parameters between every two of the four variables for Tree L According to the principle of the strongest dependence, determine that the core variable is 1, and there is a leading and being-led relationship between variables 2, 3, 4 and 1. Considering the principles of "ensuring that the edges are connected between the nodes with the strongest correlation for the initial nodes" and "ensuring that each node is connected to at least one of these edges", the first tree structure diagram can be determined, that is, construct the structure as Figure 2 shown; for Tree2, calculate respectively According to determine the structure of Tree2; for Tree3, calculate determine the structure as
[0082] In step 2, the maximum likelihood estimation method is used to estimate the C-Vine Copula parameters.
[0083] Assume that [X1, X2,..., X N represents a station network composed of d stations, and d is the number of stations in the initial station network; x itDenote the t-th observation value of the station numbered i, where t = 1, 2,..., n; n is the size of the data sample collected at each station; after determining the structure of the C-Vine Copula tree, estimate the parameters of the N-dimensional C-vine Copula according to the maximum likelihood method. The maximum likelihood function is:
[0084]
[0085] Solving the above formula according to the maximum likelihood method gives the parameter Θ:
[0086] Θ = arg max ln L(Θ)
[0087] where Θ is the C-Vine Copula parameter set, and c j,j+1|1,...,j-1 denotes the two-dimensional Pair Copula density; F(x jt |x 1t ,..., x (j-1)t ) is the conditional joint distribution function, that is, the joint distribution function given the variable set {X1, X2,... X j-1}.
[0088] Step 3 Obtain the high-dimensional mutual information through the functional relationship between the multivariate mutual information and the high-dimensional C-Vine Copula density.
[0089] The multivariate mutual information TC is derived from the following formula:
[0090] TC(x1, x2,..., x d ) = -H c (u1, u2,..., u d ) (3)
[0091] H c (u1, u2,..., u d ) = -∫c(u1, u2,..., u d )logc(u1, u2,..., u d )dU
[0092] = -E[log c(u1, u2,..., u d )] (4)
[0093] where d represents the number of stations in the station combination and d ≤ N, and u i represents the marginal distribution function of the i-th station; G d represents the station combination composed of d stations.
[0094] Step 4. Consider the time-varying characteristics of rainfall and optimize the dynamic rainfall station network by standardizing the MiK-MiT-MaJ index and the sliding window method.
[0095] (1) To consider the impact of climate change on the optimization results of rainfall station networks, the sliding window method is used to dynamically analyze the optimization results of the station network from the perspective of the trend of rainfall sequences. In this embodiment, window widths of 10 years, 5 years, and 2 years are adopted. Different window widths divide the original station network sequence G d into m station network subsequences where the station network subsequence and the original station network sequence G d only differ in the length of the rainfall sequence, and the number of stations remains m.
[0096] (2) The standardized MiK-MiT-MaJ index is used to transform the common multi-objective optimization problem into a single-objective optimization structure. For a specific station network subsequence To select c (c < d) stations from d stations to form a station network, it is necessary to calculate Stand MKTJ as
[0097]
[0098]
[0099]
[0100]
[0101] Based on the principle of the minimum KSE value as the optimization objective function, KSE is:
[0102]
[0103] where h 0j is the distance between any site position in the set of G c (c sites) and any other point outside G c , γ(h 0j ) is the mutation value at distance h 0j , μ x is the Lagrange multiplier, N is the number of spatially distributed unknown sites, and w j is the weight value;
[0104] The joint information entropy value JE is:
[0105]
[0106] where p(x1,x2,…,x c ) is the joint probability density function under c sites;
[0107] (3) Select c sites from the obtained d sites to generate k different optimized result datasets of sites, and calculate the occurrence frequency of each site according to the number of times each site appears:
[0108]
[0109] Among them, is the frequency of site x i entering the optimal site combination, and finally the optimal selection frequency of each site can be obtained
[0110] Example 1: Optimize the dynamic rain gauge network composed of 43 sites in the Huaihe River Basin in this example as an actual application
[0111] Taking the rain gauge network composed of 43 rain gauges in the Huaihe River Basin as an example, with the daily precipitation observation sequence from 1992 to 2018 as the research object, the optimization method of the present invention is used to evaluate and optimize this rain gauge network.
[0112] (1) Basin overview
[0113] The Huaihe River Basin is located in the climate transition zone between the humid and semi-humid regions of the East Asian monsoon. It is the overlapping area of the three transition zones of north-south climate, high and low latitudes, and land and sea. The weather system is complex and changeable. The large-scale circulation and water vapor transport background also have a very significant impact on the climate characteristics of the Huaihe River Basin. It is a "sensitive area" of climate change in China, forming the typical flood and drought characteristics of the Huaihe River Basin, namely "drought without precipitation, flood with precipitation, and flood with heavy precipitation". During the period from 1949 to 2009, 11 major floods and 13 major droughts occurred. Flood and drought disasters have become the main natural disasters in the Huaihe River Basin. Therefore, it is very urgent to select the rain gauge network in this basin to consider the optimization research of the dynamic rain gauge network under the influence of climate change. As Figure 3 can be seen, the rainfall sequences of the rain gauge network composed of 43 original sites generally have a trend characteristic with a significance level of 90%. Therefore, it is necessary to use the sliding window method for dynamic optimization.
[0114] (2) Model operation
[0115] The optimal station combination under the premise of selecting 13 stations from 43 can be obtained through the station network optimization model combined with the C-Vine Copula entropy theory, as shown in Table 1. The first station combination (12, 3) in Table 2 means that among the combinations of selecting 2 stations from 13 stations, the combination of station 12 and station 3 has the smallest information redundancy, the smallest areal rainfall estimation, and the largest total information volume, which can meet the MiKMiTMaJ criterion. Similarly, the optimal station combinations in the cases of selecting 3 stations from 13, 4 stations from 13,..., 12 stations from 13 can be obtained, as shown in Table 1. It can be found that as the number of stations increases, the information redundancy is continuously increasing, while the total information volume increases rapidly under the station combinations with the number of stations from 1 to 4, but when the number of stations increases further, the growth rate tends to be flat. The simulation parameter information of the C-Vine Copula under each optimal station combination is shown in Tables 2 - 12.
[0116] Table 1 Multi-objective optimization results with the C-Vine Copula entropy as the core index under the optimal combination
[0117]
[0118]
[0119] Table 2 2D C-Vine Copula optimization results
[0120]
[0121] In Table 2, t: student t Copula; SG: Survival Gumbel Copula; F: Frank Copula; Gau: Gaussian Copula; SC: Survival Clayton Copula; C: Clayton Copula; R90G: Rotated 90° Gumbel Copula; R90J: Rotated 90° Joe Copula; G: Gumbel Copula; R270J: Rotated 270° Gumbel Copula. Edges 1 and 2 represent the rainfall data sets of stations 12 and 3 respectively.
[0122] Table 3 3D C-Vine Copula optimization results
[0123]
[0124] In Table 3, edges 1, 2, and 3 represent the rainfall data sets of stations 12, 3, and 6 respectively.
[0125] Table 4 4D C-Vine Copula optimization results
[0126]
[0127] In Table 4, edges 1, 2, 3, and 4 represent the rainfall datasets of stations 12, 3, 6, and 4, respectively.
[0128] Table 5 Optimization results of 5-dimensional C-Vine Copula
[0129]
[0130] In Table 5, edges 1, 2, 3, 4, and 5 represent the rainfall datasets of stations 12, 3, 6, 4, and 10, respectively.
[0131] Table 6 6-dimensional C-Vine Copula optimization results
[0132]
[0133]
[0134] In Table 6, edges 1, 2, 3, 4, 5, and 6 represent the rainfall datasets of stations 12, 3, 6, 4, 10, and 1, respectively.
[0135] Table 7 7-dimensional C-Vine Copula optimization results
[0136]
[0137] In Table 7, edges 1, 2, 3, 4, 5, 6, and 7 represent the rainfall datasets for stations 12, 3, 6, 4, 10, 1, and 5, respectively.
[0138] Table 8 8-dimensional C-Vine Copula optimization results
[0139]
[0140]
[0141] In Table 8, edges 1, 2, 3, 4, 5, 6, 7, and 8 represent the rainfall datasets of stations 12, 3, 6, 4, 10, 1, 5, and 7, respectively.
[0142] Table 9 9-dimensional C-Vine Copula optimization results
[0143]
[0144]
[0145] In Table 9, edges 1, 2, 3, 4, 5, 6, 7, 8, and 9 represent the rainfall datasets of stations 12, 3, 6, 4, 10, 1, 5, 7, and 2, respectively.
[0146] Table 10 Optimal Results of 10-Dimensional C-Vine Copula
[0147]
[0148]
[0149] In Table 10, edges 1, 2, 3, 4, 5, 6, 7, 8, 9, and 10 represent the rainfall data sets of stations 12, 3, 6, 4, 10, 1, 5, 7, 2, and 13 respectively.
[0150] Table 11 Optimal Results of 11-Dimensional C-Vine Copula
[0151]
[0152]
[0153]
[0154] In Table 11, edges 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, and 11 represent the rainfall data sets of stations 12, 3, 6, 4, 10, 1, 5, 7, 2, 13, and 8 respectively.
[0155] Table 12 Optimal Results of 12-Dimensional C-Vine Copula
[0156]
[0157]
[0158] In Table 12, edges 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, and 12 represent the rainfall data sets of stations 12, 3, 6, 4, 10, 1, 5, 7, 2, 13, 8, and 9 respectively.
[0159] (3) Simulation and Evaluation of C-Vine Copula Function
[0160] As can be seen from Tables 2 - 12, with the continuous increase in the final number of stations, the tree structure of the high-dimensional C-Vine Copula function becomes increasingly complex. Although the structure of the model is complex, the C-Vine Copula function can mix various different types of Pair Copula functions and can represent a more complex dependence structure compared to the traditional Archimedean and elliptical family Copula function families.
[0161] At the same time, based on the C-Vine Copula function and the Archimedean Copula function, the multi-dimensional Copula entropy is derived according to formulas (4) and (5), and then the TC value is obtained. The specific comparison results are shown inFigure 4 As can be seen from the figure, the TC value derived from the C-Vine Copula function shows that the redundant information increases as the number of stations increases, which is in line with the fact that the redundant information will relatively increase with the increase in the number of stations. However, due to the insufficient simulation of the high-dimensional joint distribution by the Archimedean Copula function, the TC value calculated by it does not show an obvious increasing trend as the number of stations increases, and there is a certain degree of volatility. This also well reflects that as the dimension increases, the Archimedean Copula function cannot effectively simulate the probability density function of the joint distribution compared with the C-Vine Copula function. Therefore, the C-Vine Copula function is adopted in the present invention to calculate the information redundancy of multiple stations, and its estimation effect is significantly better than the commonly used Archimedean function, especially in the case of higher dimensions, the difference is more obvious.
[0162] (4) Mutual information comparison
[0163] In order to compare the accuracy of the mutual information values estimated by Copula entropy, a simulation experiment method was also adopted. Five-dimensional Gaussian random numbers were randomly generated as the data source for analysis. The Figure 5 - Figure 6 comparative analysis results of the mutual information values obtained by the Copula entropy-based method and the mutual information estimated values obtained by the joint histogram are shown. It can be seen from this that the mutual information estimated value based on C-Vine Copula entropy is better than the joint distribution histogram method.
[0164] (5) Dynamic station network optimization results considering the time-varying characteristics of the sequence
[0165] Finally, the station network optimization results based on a sliding window width of 10 are shown in Figure 7 . As can be seen from the figure, since the window width is 10 and the step size is 0.25 years, 68 sub-sequence station networks are generated from the original sequence. According to formula (11), the optimal frequency of each station can be obtained. Figure 6 The optimal frequency of the stations in is proportional to the size of the solid circles.
[0166] The present invention adopts a station network optimization method combining C-Vine Copula entropy and geostatistics, which effectively improves the simulation accuracy of high-dimensional joint distribution while taking into account the factors of the total amount of information and spatial estimation accuracy. While objectively and reasonably obtaining the optimal station network combination, it also has the function of dynamically optimizing the rainfall station network under the background of climate change.
Claims
1. A station network optimization method based on high-dimensional Copula entropy and Kriging, characterized in that: It includes the following steps: (1) Construct the C-Vine Copula tree structure of the hydrological station network; (2) Estimate the C-Vine Copula parameters by the maximum likelihood estimation method; (3) Obtain the multivariate mutual information through the functional relationship between the multivariate mutual information and the high-dimensional C-Vine Copula density; (4) Optimize the dynamic rainfall station network by the standardized MiK-MiT-MaJ index and the sliding window method; Step (4) includes the following steps: (41) The sliding window method is used to dynamically analyze the station network optimization results from the perspective of rainfall sequence trend; the original station network sequence G is converted into d Divide into m station network subsequences The station network subsequence and the original station network sequence G d The only difference is the length of the rainfall series, and the number of stations is still d stations; (42) Transform the multi-objective optimization problem into a single-objective optimization by using the standardized MiK-MiT-MaJ index, and for the station network subsequence By maximizing Stand MKTJ Obtain the optimal station combination of selecting c stations from d stations Stand MKTJ For Based on the principle of the minimum KSE value as the optimization objective function, where KSE is: Among them, h 0j is the position of any one of the c sites in the G c set and the distance between any other point outside G c , and γ(h 0j ) is the mutation value under the distance h 0j , μ x is the Lagrange multiplier, and w j is the weight value; The joint information entropy value JE is: Among them, p(x1, x2, …, x c ) is the joint probability density function under c sites; (43) Select c stations from the obtained d stations to generate k different station optimization result datasets, and calculate the occurrence frequency of each station according to the number of times each station appears: Among them, is the frequency of site x i entering the optimal site combination, and finally the preferred frequency of each site can be obtained 2. The station network optimization method based on high-dimensional Copula entropy and Kriging according to claim 1, characterized in that: Step (1) includes the following steps: (11) Calculate the pairwise dependence parameters of all variables The conditional set D of Tree L is an empty set; (12) By the traversal method, select the node with the largest sum of dependence parameters for all nodes to form a tree; (13) Select nodes {j, k|D} on Tree L to form edges, and calculate the corresponding C-Vine Copula type and parameters, where L = 1, …, 3(d - 1), d is the number of stations, and j and k are tree node numbers; (14)Obtain the pseudo-observations according to step (13) and means traversing the edges of each tree; (15) The tree node j = j + 1, repeat steps (11) to (14) until the entire tree structure is determined.
3. The station network optimization method based on high-dimensional Copula entropy and Kriging according to claim 1, wherein: Step (2) The maximum likelihood function is: Among them, [x1, x2, …, x d represents the variable combination of the station network G d composed of d stations, where d is the number of stations in the initial station network; x it represents the t-th observation value of the station numbered i, t = 1, 2, …, n; n is the size of the data samples collected at each station; Θ is the C-Vine Copula parameter set, and c j,j+1|1,…,j-1 represents the two-dimensional Pair Copula density; F(x jt |x 1t , …, x (j-1)t ) is the conditional joint distribution function, that is, the joint distribution function given the variable set {x1, x2, … x j-1}; Solve the maximum likelihood function to obtain the parameter Θ Θ = arg max ln L(Θ) (2).
4. The station network optimization method based on high-dimensional Copula entropy and Kriging according to claim 1, characterized in that: Step (3) The multivariate mutual information TC is obtained from formulas (3) and (4) TC(x1,x2,…,x d )=-H c (u1,u2,…,u d ) (3) H c (u1,u2,…,u d ) = -∫c(u1,u2,…,u d ) log c(u1,u2,…,u d ) dU =-E[log c(u1,u2,…,u d )] (4) where d represents the number of stations in the station combination, and u i represents the marginal distribution function of the i-th station.
Citation Information
Patent Citations
Optimization model for hydrologic station networks based on combination of kriging method and information entropy theory
CN107315722A
Based on Copula function and information entropy theory, a comprehensive hydrological network evaluation model
CN109460526A