A precipitation rate prediction method based on climate network
Patent Information
- Application Number
- CN202310414599.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-18
- Publication Date
- 2026-09-15
- Estimated Expiration
- 2043-04-18
AI Technical Summary
这些理论通常基于统计学模型、历史气候数据及模型来预测未来的气候模式,但无法考虑复杂的气候现象机制和原理
[0052] This technical solution utilizes complex networks, a powerful data analysis tool, to establish a Pearson time-delay cross-correlation network of global node-to-node temperature anomalies. It identifies a key region for predicting total precipitation during the rainy season in rainy areas such as Hainan Province, and constructs a robust and adaptive prediction model for the time series of network indicators in this region using ridge regression. Therefore, the model's advantage lies in its close integration with physical mechanisms, employing a simple statistical model to achieve a mean percentage error (MAPE) of less than 6.8% and an RMSE of less than 0.0002138 kg/m³ one year in advance. 2The total precipitation rate during the rainy season in tropical rainy regions is predicted at /s. This provides a new approach for long-term climate prediction of regional precipitation rates in tropical rainy regions.
Smart Images

Figure CN116579463B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of precipitation rate prediction technology, specifically to a method for predicting precipitation rate during the rainy season in tropical rainy regions based on climate networks. Background Technology
[0002] Rainy regions generally fall within tropical monsoon climate zones, primarily influenced by the southeast and southwest monsoons. The amount and distribution of rainfall directly impact agricultural production, the ecological environment, and tourism. Excessive or insufficient rainfall affects crop growth and yields, as well as ecosystem functions such as soil moisture and the hydrological cycle. Furthermore, rainfall significantly impacts tourism; excessive rainfall can deter tourists, while insufficient rainfall can lead to water shortages and forest fires. Therefore, accurately predicting rainfall in rainy tropical regions, such as Hainan Province in my country, is crucial for agriculture, the ecological environment, and tourism.
[0003] However, with the intensifying trend of global warming, the climate system has become more complex and unstable, and the uncertainties in climate models have become more significant. Consequently, the accuracy and precision of traditional climate prediction methods have declined, and their predictive power for long-term or medium-term climate change has been limited. Furthermore, current methods for predicting precipitation rates primarily utilize theories such as statistics, machine learning, neural networks, and grey systems. These theories typically rely on statistical models, historical climate data, and models to predict future climate patterns, but they cannot account for the complex mechanisms and principles of climate phenomena. Additionally, global warming means that historical data can no longer fully reflect future climate patterns.
[0004] Complex network theory provides a powerful tool for studying the structure, dynamics, and function of complex systems. Due to the nonlinear interactions and feedback loops within and between different levels of the climate system, it is a typical complex adaptive system. Climate networks offer a novel theoretical framework for the study of complex climate systems and have been successfully used to reveal and predict a variety of important climate mechanisms and phenomena. Summary of the Invention
[0005] Purpose of the invention: To provide a method for predicting precipitation rate in rainy areas during the rainy season based on climate networks.
[0006] Technical solution: A precipitation rate prediction method based on climate networks, comprising the following steps:
[0007] 1) Data acquisition and data preprocessing:
[0008] 1-1) Data Acquisition: Global regional Y1 to Y2 data were acquired using the NCEP-NCAR Reanalysis I dataset.n The annual precipitation rate R, monthly average data, and daily air temperature T at a height of 2 meters;
[0009] 1-2) Select N points at equal intervals covering the entire globe G Each grid point serves as a node in the climate network: the Earth is assumed to be a perfect sphere with a radius of R. E Then, under a fixed meridional resolution HSR(°), equal intervals are selected along the equator. N E 1 grid point, of which Therefore, at any latitude θ, the corresponding perimeter is R. E On the circumference of cosθ, select [N] E [cosθ] grid points such that the distance between any two adjacent points is d0, thus
[0010] 1-3) Calculation of climate anomalies: N is calculated using the following formula. G Among the grid points, the original temperature time series of grid point i on day d in year y. Calculate outlier T i y (d) To remove seasonal trends:
[0011] in
[0012]
[0013] σ(T i y (d) Calculate similarly;
[0014] 2) Selection and processing of prediction targets:
[0015] First, select N, a representative of rainy areas. H Each of the grid points is used to calculate the values from Y1 to Y2. n Every year, during the rainy season, N H Average total precipitation rate per grid point:
[0016]
[0017] As the total precipitation rate of the rainy season in each year, among which This represents the monthly average precipitation rate data for grid point i in month m of year y, thus obtaining the time series set of total precipitation rate data for the rainy season in rainy areas: And use it as the prediction target of the prediction scheme;
[0018] 3) Establishment of a climate network:
[0019] 3-1) Calculate the air temperature anomaly time series T between nodes i and j by the following formula i y (d) Pearson time-lagged cross-correlation function:
[0020]
[0021]
[0022] wherein the square brackets represent the average value over the past N days, that is
[0023]
[0024] where τ∈[0,τ max , y represents the year, d represents the end point of the time series, and N is the length of the time series; next, calculate the following formula
[0025]
[0026] and assign weights to network edges based on this value;
[0027] 3-2) Define the minimum value of the Pearson cross-correlation function in formula (3.4), the time delay τ corresponding to is that is and through θ i,j determine the direction of the edge between nodes i and j by the sign of ; next, define the degree of network node i as the average weight of directed edges pointing from node i to other nodes:
[0028]
[0029] Here, the adjacency matrix of the climate network
[0030]
[0031] 4) Determination of prediction index:
[0032] 4-1) Divide the precipitation rate data time series, totaling n years, into a learning phase which is the first m years of data, where m<n, and a testing phase; in the learning phase, the annual precipitation rate and the network index of the previous year the degree of correlation and statistical significance between the network indicators of different nodes and the precipitation rate are evaluated through the linear correlation between them, so as to preliminarily determine the geographical region for predicting precipitation rate, that is, network nodes; in the testing phase, the network indicators of the previously selected region are used for prediction and the prediction effect is evaluated; wherein y≤Y m ;
[0033] 4-2) In the learning phase, to quantify the correlation between the precipitation rate in a certain phase and the network indicators, select the precipitation rates of the last L (L<m) years of this learning phase, that is, from the (m-L+1)-th year to the m-th year, recorded as Correspondingly, for each node i in the network, select its network indicators from the (m-L)-th year to the (m-1)-th year, that is Then Student's t-test is used to quantify the statistical significance of the linear correlation between the network indicator of any network node i and the precipitation rate, which is recorded as It ranges from 0% to 100%. A smaller value indicates that the linear correlation between the precipitation rate and the network indicator at node i is more significant; thus, point sets with a 99% confidence level under different lengths L+μ, μ=0,1,...,m-L are obtained:
[0034] 4-3) Define the above m-L+1 point sets S L+μ , among μ=0,1,...,m-L, points that appear repeatedly at least half of the time, that is (≥(m-L) / 2) times, are robust points, recorded as i r , the set composed of these points is recorded as S r ; the annual network indicators of these nodes show a relatively stable linear correlation with the precipitation rate in rainy regions in the next year, which are used as potential indicators for predicting precipitation rate;
[0035] 4-4) Test the prediction ability of different nodes i in set S r different nodes i in r : To predict the precipitation rate in rainy regions for the γ-th year (m<γ<n), select the time series of precipitation rate in this region for the previous α years and the time series of network indicators one year in advance of node i r to establish a univariate linear regression model: establish a univariate linear regression model:
[0036] Y=βX+c, (4.1)
[0037] wherein β is the fitting parameter; substitute the network indicator of the (γ-1)-th year into formula (4.1), so as to obtain the predicted precipitation rate for the γ-th year, that is
[0038] 4-5) Calculate the root mean square error RMSE between the predicted value and the observed value, that is
[0039]
[0040] quantitatively evaluate the prediction performance, wherein a smaller RMSE indicates a higher prediction accuracy; finally, change the value of α, and calculate the corresponding RMSE for any α=5,6,7,...,m iα Then, the average RMSE is calculated with respect to the learning length α, i.e.
[0041]
[0042] The node with the smallest average RMES is selected as the optimal node for prediction, and the set of these nodes is denoted as S. pred .
[0043] 5) Optimization and implementation of the prediction model:
[0044] Select the set S of the first α years of the learning phase pred The time series of network indicators of the middle nodes and the corresponding time series of precipitation rates in rainy areas were used to establish a ridge regression model, the parameters of which were estimated as follows:
[0045]
[0046] in λ is the ridge parameter, I α Let be the α-order identity matrix; the RMSE was calculated using equation (4.2) in (4-5). α,λ The ridge regression parameters of the final prediction model are determined by using the parameters corresponding to the minimum RMSE values. and optimal learning length
[0047] A ridge regression model is established using the predictor factors and model parameters determined through the above process.
[0048]
[0049] Here |S pred | represents set S pred The number of elements in the model; using this model, the precipitation rate in year γ-1 is predicted using the network index, and its value is...
[0050]
[0051] Beneficial effects
[0052] This technical solution utilizes complex networks, a powerful data analysis tool, to establish a Pearson time-delay cross-correlation network of global node-to-node temperature anomalies. It identifies a key region for predicting total precipitation during the rainy season in rainy areas such as Hainan Province, and constructs a robust and adaptive prediction model for the time series of network indicators in this region using ridge regression. Therefore, the model's advantage lies in its close integration with physical mechanisms, employing a simple statistical model to achieve a mean percentage error (MAPE) of less than 6.8% and an RMSE of less than 0.0002138 kg / m³ one year in advance. 2The total precipitation rate during the rainy season in tropical rainy regions is predicted at / s. This provides a new approach for long-term climate prediction of regional precipitation rates in tropical rainy regions. Attached Figure Description
[0053] Figure 1 A flowchart illustrating a precipitation rate prediction method based on a climate network, provided for an embodiment of this application;
[0054] Figure 2a This application provides an embodiment of an optimal climate network prediction index, which uses a ridge regression model to predict data in a test set under different parameters and training set lengths. The RMSE is the title of the graph, the horizontal axis is the logarithmic scale, representing the ridge regression parameter λ, and the vertical axis is the length α of the learning phase.
[0055] Figure 2b A method for providing an embodiment of this application Figure 2a The local zoom-in of RMSE heatmap within the black dashed box; where Local zoom of RMSE is the title, the horizontal axis is the linear scale representing the ridge regression parameter λ, and the vertical axis is the length of the learning phase α.
[0056] Figure 3 An average network metric provided for embodiments of this application Linear relationship with total precipitation rate; where the horizontal axis represents the average network index, the vertical axis represents the total precipitation rate, the circles correspond to the learning stage, the triangles correspond to the testing stage, the solid line is the learning stage, that is, the regression curve of the circles in the figure, and the dotted line is the testing stage, that is, the regression curve of the triangles in the figure.
[0057] Figure 4 The graph shows the prediction effect of a regression model in the test set, as provided in the embodiments of this application. The solid line and dots represent the observed precipitation rate, the dashed line and squares represent the predicted precipitation rate, the horizontal axis represents the predicted year, and the vertical axis represents the total precipitation rate. Detailed Implementation
[0058] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0059] The technical solutions provided by the various embodiments of this application are described in detail below with reference to the accompanying drawings. Taking Hainan Province, my country as an example, this application provides a precipitation rate prediction method based on a climate network, and its specific implementation process is illustrated in the schematic diagram below. Figure 1As shown, it includes the following steps:
[0060] i. Data acquisition and data preprocessing:
[0061] 1. Data Acquisition: Using the NCEP-NCAR Reanalysis I dataset,
[0062] (https: / / psl.noaa.gov / data / gridded / data.ncep.reanalysis.html) Obtain monthly average data of global regional precipitation rate R (Precipitation rate) and daily data of air temperature T (Air temperature) at a height of 2 meters from 1948 to 2022;
[0063] 2. Selecting 6242 grid points covering the globe at equal intervals as nodes of the climate network: The dataset mentioned in the first step has a Gaussian spatial distribution, that is, it has 192×94 grid points covering the globe; based on this, we define the average meridional resolution as 3.8°, and thus select 6242 nodes covering the globe at equal intervals.
[0064] 3. Climate anomaly calculation: The original temperature time series of grid point i on day d in year y is calculated using the following formula. Calculate outlier T i y (d) To remove seasonal trends:
[0065]
[0066] in
[0067]
[0068] σ(T i y (d) Calculate similarly;
[0069] ii. Selection and processing of prediction targets:
[0070] First, four grid points representing Hainan Province were selected; the average total precipitation rate for each of the four grid points during the rainy season (May-October) from 1950 to 2022 was calculated:
[0071]
[0072] Using approximations as the total rainfall rate during the rainy season in Hainan Province for each year, among which Let represent the monthly average precipitation rate data of grid point i in month m of year y. Therefore, we can obtain the time series data set of total precipitation rate during the rainy season in Hainan Province. And take it as the prediction target of the prediction scheme proposed in this application.
[0073] iii. Establishment of a climate network:
[0074] 1. Calculate the time series T of temperature anomalies between nodes i and j using the following formula. i y (d) Pearson time-delay cross-correlation function:
[0075]
[0076]
[0077] The square brackets represent the average value over the past 365 days.
[0078]
[0079] Where τ∈[0,200], y represents the year, here y=1949,1950,...,2021; d represents the end point of the time series, i.e. December 31st of each year; next, calculate the following formula.
[0080]
[0081] This value is then used to assign weights to the network edges;
[0082] On the other hand, we define the Pearson cross-correlation function. The minimum value, i.e., in equation (3.4) The corresponding time delay τ is Right now and through θ i,j The sign of the edge determines the direction of the edge between nodes i and j; next, we define the degree of network node i as the average weight of the directed edges from node i to the other nodes:
[0083]
[0084] Here, the climate network adjacency matrix
[0085]
[0086] iv. Determination of the forecast index:
[0087] 1. First, we divided the precipitation rate data time series (1950 to 2022, a total of 73 years) into a learning phase (i.e., the first 52 years of data) and a testing phase; in the learning phase, we used the annual precipitation rate... (where y≤2001) and its network indicators from the previous year We will use the linear correlation between the network indicators and precipitation rate to evaluate the degree of association and statistical significance between different nodes, so as to initially determine the geographical areas (network nodes) for predicting precipitation rate; in the testing phase, we will use the network indicators of the previously selected areas for prediction and evaluate its prediction effect.
[0088] 2. During the learning phase, to quantify the correlation between precipitation rate and network indicators for a specific period, we first selected the precipitation rate for the last five years of that learning phase (i.e., 1997 to 2001), denoted as . Accordingly, for each node i in the network, we select its network metrics from 1996 to 2000, i.e. Then, Student's t-test is used to quantify the statistical significance of the linear correlation between the network index of any network node i and the precipitation rate, and denoted as . Here, values range from 0% to 100%. A larger value indicates a more significant linear correlation between the precipitation rate and the network index at node i. Therefore, we can obtain point sets with 99% confidence for different lengths L+p, p = 0, 1, ..., 47:
[0089] Furthermore, according to the definition in the technical solution, in this set of 48 points S L+p In the range p = 0, 1, ..., 47, a point that appears at least half of the time (i.e., 24 times or more) is a robust point, denoted as i. r The set of these points is denoted as S. r These nodes show a relatively stable linear correlation between their annual network indicators and the precipitation rate in Hainan the following year, and therefore can be used as a potential indicator of Hainan's precipitation rate.
[0090] 3. Next, we examine set S. r Different nodes i r Predictive power: To predict the precipitation rate in Hainan Province in year γ (52 < γ ≤ 73), we selected the time series of Hainan precipitation rates from the previous α years. With node i r Network indicator time series one year in advance Establish a univariate linear regression model:
[0091] Y = βX + c, (4.1) where β is the fitting parameter. The network metrics for year γ-1 are... Substituting into equation (4.1), we obtain the predicted precipitation rate for year γ, i.e. Furthermore, by calculating the mean square error (RMSE) between the predicted and observed values, i.e.
[0092]
[0093] The predictive performance is quantitatively evaluated, with a smaller RMSE indicating a higher predictive accuracy. Finally, the value of α is varied, and the corresponding values are calculated for any α = 5, 6, 7, ..., 52. Then, the average RMSE is calculated for the learning length α, i.e.
[0094]
[0095] The node with the smallest average RMSE is selected as the optimal node for prediction.
[0096] v. Optimization and implementation of the prediction model:
[0097] From a climatological perspective, we tend to study inter-regional interactions; therefore, the point set selected in the above process will also appear in a specific region. However, due to the high similarity of climatic characteristics within this specific region, we need to avoid multicollinearity and overfitting when building the regression model. Therefore, to predict the precipitation rate from year 53 to 73, we first selected the set S of the first α years of the learning phase. pred The time series of network indicators of the middle nodes and the corresponding time series of precipitation rates in Hainan Province were used to establish a ridge regression model, the parameters of which were estimated as follows:
[0098]
[0099] in λ is the ridge parameter, I α Let be the α-order identity matrix. The RMSE was calculated with reference to equation (4.2) in Part 3 of Section IV. α,λ And through the corresponding parameter of the minimum RMSE value (such as Figure 2a (As shown in Figure 2b) the ridge regression parameters of the final prediction model were determined. and optimal learning length
[0100] A ridge regression model is established using the predictor factors and model parameters determined through the above process.
[0101]
[0102] Using this model, we can predict the precipitation rate in year γ using network indicators from year γ-1. The table below illustrates the predictive performance of this model using the past seven years (2016 to 2022) as an example:
[0103]
[0104] Figure 4 The predictions of the model during the testing phase are shown.
Claims
1. A precipitation rate prediction method based on a climate network, characterized in that, Includes the following steps: 1) Data acquisition and data preprocessing: 1-1) Data Acquisition: Global regions were acquired through the NCEP-NCAR Reanalysis I dataset. Year The annual precipitation rate R, monthly average data, and daily air temperature T at a height of 2 meters; 1-2) Selecting equidistant areas covering the entire globe Each grid point serves as a node in the climate network: The Earth is assumed to be a perfect sphere with a radius of... Then, at a fixed meridional resolution Below, equal intervals are selected on the equator. of 1 grid point, of which This represents the target spatial resolution of climate network nodes along the horizontal direction on the Earth's surface, i.e., the arc length spacing between adjacent grid points, thus enabling spatial resolution at any latitude. The corresponding perimeter is On the circumference of the circle, select There are lattice points such that the distance between any two neighboring points is . ,thereby ; 1-3) Calculation of climate anomalies: The following formula is used to calculate... Among the grid points, the original temperature time series of grid point i on day d in year y. Calculate outliers To remove seasonal trends: (1.1) in, (1.2) 2) Selection and processing of prediction targets: First, select representative rainy areas. Each of the grid points is calculated separately. Year Every year, during the rainy season Average total precipitation rate per grid point: (2.1) As the total precipitation rate of the rainy season in each year, among which This represents the monthly average precipitation rate data for grid point i in month m of year y, thus obtaining the time series set of total precipitation rate data for the rainy season in rainy areas: And use it as the prediction target of the prediction scheme; 3) Establishment of a climate network: 3-1) Calculate the nodes using the following formula Time series of inter-temperature anomalies Pearson time-delay cross-correlation function: (3.1) (3.2) Where the square brackets represent the average value over the past N days, i.e. (3.3) in y represents the year, d represents the end point of the time series, and N is the length of the time series; next, calculate the following formula. (3.4) This value is then used to assign weights to the network edges; 3-2) Define the Pearson cross-correlation function The minimum value, i.e., in equation (3.4) Corresponding time delay for ,Right now ; and through The symbol determines the node The direction of the edges; next, the degree of network node i is defined as the average weight of the directed edges pointing from node i to the other nodes: (3.5) Here, the climate network adjacency matrix (3.6) 4) Determination of the forecast index: 4-1) Divide the time series of precipitation rate data, with a total of n years, into a learning phase, which is the first m years of the data where m < n, and a testing phase; in the learning phase, the annual precipitation rate and the network index of its previous year the degree of association and statistical significance between the network indicators of different nodes and the precipitation rate are evaluated through the linear correlation between them, so as to preliminarily determine the geographical region, i.e., network nodes, for predicting the precipitation rate; in the testing phase, the network indicators of the previously selected regions are used for prediction, and the prediction effect is evaluated; wherein ; 4-2) During the learning phase, in order to quantify the correlation between precipitation rate and network indicators at a certain stage, the precipitation rate of the last L years of that learning phase is selected, i.e., the precipitation rate of the Lth year. From year m, , recorded as Correspondingly, for each node i in the network, select its... Year Annual network metrics, i.e. Then, Student's t-test is used to quantify the statistical significance of the linear correlation between the network index of any network node i and the precipitation rate, and denoted as . The value ranges from 0% to 100%. A smaller value indicates a more significant linear correlation between the precipitation rate and the network indicators at node i; thus, different lengths... The following point set has a 99% confidence level: ; 4-3) Define the above Point set In, at least half of them are repeated, that is The point of the second order is a robust point, denoted as . The set of these points is denoted as The network indicators of these nodes show a relatively stable linear correlation with the precipitation rate of the rainy areas in the following year, serving as a potential indicator for predicting precipitation rate. 4-4) Test set Different nodes Predictive power: In order to predict the first Rainfall rate in rainy areas in a given year Before selection Time series of precipitation rate in this region in 2018 , with nodes Network indicator time series one year in advance Establish a univariate linear regression model: ,(4.1) in For the fitting parameters; the first... Annual network metrics Substituting into equation (4.1), we obtain the first... The annual precipitation rate forecast, i.e. ; 4-5) Calculate the mean square error (RMSE) between the predicted and observed values, i.e. (4.2) Quantitative evaluation of prediction performance is conducted, with a lower RMSE indicating a higher prediction rate; finally, changes are made. The size, for any Calculate the corresponding Furthermore, regarding the length of study Calculate the average RMSE, i.e. (4.3) The node with the smallest average RMES is selected as the optimal node for prediction, and the set of these nodes is denoted as . ; 5) Optimization and implementation of the prediction model: Before selecting a learning stage Annual Collection The time series of network indicators of mid-nodes and the corresponding time series of precipitation rates in rainy areas were used to establish a ridge regression model, the parameters of which were estimated as follows: (5.1) in For ridge parameters, for The identity matrix of order 1; calculated with reference to equation (4.2) in (4-5). The ridge regression parameters of the final prediction model are determined by using the parameters corresponding to the minimum RMSE values. and optimal learning length ; A ridge regression model is established using the predictors and model parameters determined through the above process. (5.2) Here Represents a set The number of elements in The intercept parameter of the ridge regression model is used to shift and correct the overall prediction level of the model, ensuring that the model can still characterize the average level or baseline bias of the target precipitation rate sequence beyond the contributions of predictor factors. Through this model, the first... Network metrics forecast for the year The annual precipitation rate, its value is 。