A rainfall prediction method and system based on key influence factor sliding clustering

By establishing a matrix of climate and meteorological influencing factors and employing sliding clustering forecasting technology, the problem of insufficient accuracy in rainfall forecasting in existing technologies has been solved, resulting in more accurate and reliable daily rainfall forecasts and extending the lead time.

CN120891564BActive Publication Date: 2026-02-10BUREAU OF HYDROLOGY CHANGJIANG WATER RESOURCES COMMISSION +2
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202511415279.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-30
Publication Date
2026-02-10
Estimated Expiration
2045-09-30

AI Technical Summary

Technical Problem

Existing rainfall forecasting technologies lack accuracy and are subject to uncertainty in forecasting complex terrain and short-term heavy rainfall events at small and medium scales, as well as in forecasting large-scale medium- and long-term rainfall under the influence of climate change, making it difficult to achieve accurate daily quantitative forecasts.

Method used

By establishing a matrix of climate and meteorological influencing factors at different time scales, and employing key influencing factor identification, correlation learning, and sliding clustering forecasting techniques, daily rolling rainfall forecasts are conducted to improve the accuracy of rainfall forecasts and extend the lead time.

Benefits of technology

It enables accurate forecasting of rainfall events at different time scales, improving the accuracy and reliability of rainfall forecasts and extending the lead time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120891564B_ABST
    Figure CN120891564B_ABST
Patent Text Reader

Abstract

The application provides a rainfall prediction method and system based on key influence factor sliding clustering, constructs a long sequence weather database by collecting natural geography and hydrological meteorological data of a research area; secondly, a rainfall influence factor matrix is constructed, and factors with large order of magnitude deviation are normalized, and key influence factors are screened out; the spatio-temporal correlation between rainfall and key factors is analyzed by using a statistical method, and key factors with significant correlation are determined; clustering analysis is performed on historical rainfall processes, the key factor characteristics of various rainfall processes are counted, and the matching relationship between the rainfall process and the influence factor is established; finally, based on the key factors of the future N day, the corresponding rainfall process category is matched, the tendency value, lower limit and upper limit of the prediction result are determined, and the prediction result is updated in a rolling manner with a day as a step. The method combines data driving and statistical analysis, realizes fine description and rolling prediction of the rainfall process, and has high applicability and prediction accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of rainfall prediction, and in particular to a rainfall prediction method and system based on key influence factor sliding clustering. BACKGROUND

[0002] Rainfall prediction is an important basis for disaster prevention and reduction, agricultural production and water resources management. Accurate rainfall prediction can effectively reduce the loss caused by natural disasters and ensure the sustainable development of social economy. Therefore, the innovative research of rainfall prediction method has important practical significance and application value.

[0003] At present, remarkable progress has been made in rainfall prediction technology at home and abroad. Through numerical model optimization, machine learning algorithm and multi-source data fusion, the accuracy and timeliness of rainfall prediction have been gradually improved in China; breakthroughs have been made in satellite remote sensing data application, high-resolution numerical model and artificial intelligence technology in the international field. However, despite the continuous progress of technology, the accuracy of rainfall prediction is still insufficient, and the complexity and uncertainty of rainfall prediction still need to be further tackled. Especially in the prediction of complex terrain and small-scale short-term heavy rainfall events, and the prediction of large-scale medium and long-term rainfall under the influence of climate change, the limitations of existing methods are particularly prominent. Therefore, the innovative research of rainfall prediction method, and the exploration of more accurate and reliable prediction technology, are the key to improving the ability of disaster prevention and reduction, and are also the necessary direction to promote the development of meteorological science. The present application focuses on the influence factors affecting rainfall intensity at different time scales, and through key influence factor identification, correlation relationship learning and sliding clustering prediction, it carries out multi-scale rainfall prediction, which is expected to further improve the accuracy of daily quantitative rainfall prediction and prolong the prediction period. SUMMARY

[0004] The present application aims to overcome the shortcomings of the prior art and provides a rainfall prediction method and system based on key influence factor sliding clustering, which establishes the correlation between different time scale climate, meteorological and other influence factors and rainfall process, and realizes rolling prediction of daily rainfall through sliding clustering, thereby improving the accuracy of rainfall prediction and prolonging the prediction period.

[0005] To achieve the above-mentioned purpose, the present application adopts the following technical solutions:

[0006] The present application provides a rainfall prediction method based on key influence factor sliding clustering, comprising:

[0007] S1, basic data preparation: selecting a research area, determining the fine zoning according to the natural geographical characteristics and hydro-meteorological characteristics, collecting and sorting long sequence meteorological historical data, including rainfall station rainfall, regional weather system characteristic index and regional teleconnection climate factor;

[0008] S2. Determination of Influencing Factors: Establishing a matrix of rainfall influencing factors for the study area. Furthermore, factors with significant magnitude deviations were normalized. The expression for the rainfall impact factor matrix is ​​as follows:

[0009] Ψ a × b = [ j 11 j 12 ... j 1 i ... j 1 b j 21 j 22 ... j 2 i ... j 2 b ... ... ... ... ... ... j k 1 j k 2 ... j ki ... j kb ... ... ... ... ... ... j a 1 j a 2 ... j ai ... j ab ] ;

[0010] in, Indicates the first The first partition The values ​​of each influence factor; Indicates the first The first partition The values ​​of each influence factor;

[0011] The normalization formula is as follows:

[0012] ;

[0013] in, For each partition The minimum value of each influencing factor; For each partition The maximum value of each influencing factor; The normalized factor data is in the range [0,1].

[0014] After normalization, the rainfall influencing factor matrix is ​​obtained. The expression is:

[0015] Ψ ' a × b = [ j ' 11 j ' 12 ... j ' 1 i ... j ' 1 b j ' 21 j ' 22 ... j ' 2 i ... j ' 2 b ... ... ... ... ... ... j ' k 1 j ' k 2 ... j ' ki ... j ' kb ... ... ... ... ... ... j ' a 1 j ' a 2 ... j ' ai ... j ' ab ] ;

[0016] S3. Correlation Analysis: Statistical methods were used to analyze the strength of the relationship between rainfall and different factors in the influencing factor matrix. Factors with significant temporal and spatial correlations were selected as key influencing factors, and the key influencing factors for each region were recorded in the matrix. The expression is:

[0017] Ω a × m = [ q ' 11 q ' 12 ... q ' 1 i ... q ' 1 m q ' 21 q ' 22 ... q ' 2 i ... q ' 2 m ... ... ... ... ... ... q ' k 1 q ' k 2 ... q ' ki ... q ' km ... ... ... ... ... ... q ' a 1 q ' a 2 ... q ' ai ... q ' am ] ;

[0018] in, For the first The first partition The normalized values ​​of each influence factor;

[0019] S4. Starting from any time, for any partition, denote the daily rainfall for the next N days as a vector. ,in, For the first The daily rainfall over the past N days is used to perform cluster analysis on the daily rainfall over the next N days. The total number of clusters is denoted as . Category number is denoted as ,in We obtain different types of rainfall events in the next N days, and statistically analyze the characteristics of key influencing factors in the next N days. We then establish and label the correlation between various rainfall events and key influencing factors.

[0020] S5. When conducting rainfall forecasting, based on the key influencing factors for the next N days, match the corresponding rainfall process category, and select the first category from the next N days... The average daily rainfall over the past few days will be used as the basis for this forecast. The outcome tendency value for the day, in the first category of the next N days. The minimum and maximum daily rainfall values ​​for each day will be used as the basis for this forecast. The lower and upper limits of the daily results are calculated using the following formula:

[0021] ;

[0022] ;

[0023] ;

[0024] in, This represents the number of samples in the first category. For the first category, the One sample, ; For the first The first sample Daily rainfall; For the first Daily rainfall trend value for the day; For the first The maximum daily rainfall on the [day name] is the [number]th [day name]. The daily rainfall limit for a given day; For the first The minimum daily rainfall on the 1st day is the 2nd day. The lower limit of daily rainfall;

[0025] S6. Starting from the calculation start time, output the tendency value of the daily rainfall forecast results for the next N days as the forecast result of this rainfall process, output the lower limit and upper limit of the daily rainfall forecast results for the next N days as the forecast interval of the rainfall process, and slide forward with a daily step size as time goes by, carry out forecast operations day by day, output the corresponding forecast results, and realize the rolling rainfall forecast.

[0026] Furthermore, in S2, the method for selecting the influencing factor is as follows:

[0027] For rainfall data from rain gauge stations, the Thiessen polygon method is used to calculate the areal rainfall in different zones based on the density of the monitoring station network and the distribution of the stations.

[0028] Based on the characteristic indicators of regional weather systems, factors that affect the spatial structure of meteorological elements, such as temperature, humidity, wind speed, and wind direction, are selected.

[0029] For regional teleconnection climate factors, atmospheric circulation index, sea ice, sea surface temperature, solar radiation, and Indian Ocean Dipole were selected.

[0030] Furthermore, in S3, the statistical method employs a combination of path analysis and improved Lasso regression to determine the direct and indirect influence relationships between different influencing factors and rainfall, analyze the comprehensive impact of multiple factors on rainfall, and screen key factors through regularization to establish a quantitative relationship model between rainfall and influencing factors, specifically:

[0031] S31. Using path analysis, analyze the relationship between different influencing factors and rainfall. The calculation formula is as follows:

[0032] ;

[0033] in,

[0034] R= R a × n = [ y 11 y 12 ... y 1 v ... y 1 n y 21 y 22 ... y 2 v ... y 2 n ... ... ... ... ... ... y k 1 y k 2 ... y kv ... y kn ... ... ... ... ... ... y a 1 y a 2 ... y av ... y an ] ; Ψ ' = Ψ ' a × b = [ j ' 11 j ' 12 ... j ' 1 v ... j ' 1 b j ' 21 j ' 22 ... j ' 2 v ... j ' 2 b ... ... ... ... ... ... j ' k 1 j ' k 2 ... j ' kv ... j ' kb ... ... ... ... ... ... j ' a 1 j ' a 2 ... j ' av ... j ' ab ] ; δ= δ b × 1 = [ δ 1 δ 2 ... δ v ... δ b ] ; Τ = T 1 × n = [ y 11 y 12 ... y 1 w ′ ... y 1 n ] ;

[0035] in, The daily rainfall matrix for the study area is denoted as . , , All are temporary variables. , ; This is the influence factor matrix; This is the error term matrix; Indicates the first The path coefficients of each influencing factor; This represents the rainfall matrix for any given rainfall event in any given region. This indicates the number of rainfall events in any given rainfall event in the first zone. Rainfall for the day;

[0036] No. The first partition The expression for "day" is:

[0037] ;

[0038] wherein, ζ= ζ 1 × b = [ j ' a 1 j ' a 2 ... j ' aj j ' ab ] is the impact factor matrix of the th partition; represents the daily rainfall of the th partition on the th day;

[0039] S32, by improved Lasso regression, using a regularization method to screen key impact factors, taking the prediction of any partition future N-day rainfall process as an example, specifically:

[0040]

[0041] wherein, represents the rainfall of any day of the th rainfall process of any partition; is the th impact factor of the th rainfall process; represents the intercept term of the regression model, i.e. the constant term; represents the regression coefficient of the th impact factor in the regression model, indicating the impact of the impact factor on the rainfall; represents the number of rainfall processes; represents the number of impact factors; represents the regularization parameter; represents the parameter controlling the regularization proportion, γ = [ 0 , 1 ] is the variable to be optimized, which is a set of parameters, including the intercept term of the regression model and the regression coefficient of the th impact factor in the regression model ; is the path coefficient matrix of the impact factor;

[0042] S33, a quantitative relationship model between rainfall and impact factors is established to determine the key impact factors, including the number and the category to which it belongs, wherein, .

[0043] Further, in the S4, the daily rainfall of the next N days is predicted from any time as the starting point, and according to the job requirements, the rainfall of different time scales in the next month is predicted, and accordingly, the unit of clustering analysis and sliding step is also month.

[0044] Further, in the S4, the clustering analysis adopts the method of k-means clustering, spectral clustering or split clustering; ​​​

[0045] When the rainfall influencing factors in the study area are simple, the sample data is small, and the requirements for clustering categories are not high, the k-means clustering method is selected.

[0046] When the factors influencing rainfall in the study area are complex and diverse, and there is a large amount of sample data, spectral clustering is selected.

[0047] When the factors influencing rainfall in the study area are complex and diverse and the results have hierarchical requirements, split clustering is selected.

[0048] Further, it includes: at least one processor; and a memory communicatively connected to at least one of the processors; wherein,

[0049] The memory stores instructions that can be executed by the processor to implement the rainfall forecasting method based on sliding clustering of key influencing factors.

[0050] The beneficial effects of this invention are as follows: First, a long-sequence meteorological database is constructed by collecting natural geographical and hydrological meteorological data of the study area. Second, a rainfall influencing factor matrix is ​​constructed, and factors with large magnitude deviations are normalized to screen out key influencing factors. Third, statistical methods are used to analyze the spatiotemporal correlation between rainfall and key factors to identify significantly correlated key factors. Then, cluster analysis is performed on historical rainfall processes to statistically analyze the key factor characteristics of various rainfall processes and establish a correlation between rainfall processes and influencing factors. Finally, based on the key factors for the next N days, the corresponding rainfall process categories are matched to determine the trend value, lower limit, and upper limit of the forecast results, and the forecast results are updated on a daily basis. This method, through a combination of data-driven and statistical analysis, achieves a detailed characterization and rolling forecast of rainfall processes, and has high applicability and forecast accuracy. Attached Figure Description

[0051] Figure 1 This is a flowchart of a rainfall forecasting method based on sliding clustering of key influencing factors;

[0052] Figure 2 The example shows the cluster analysis results of the rainfall process at the Wuxi station of the Daning River. Detailed Implementation

[0053] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.

[0054] Please see Figure 1 A rainfall forecasting method based on sliding clustering of key influencing factors includes:

[0055] S1. Basic data preparation: Select the study area, determine the detailed zoning based on natural geographical characteristics and hydro-meteorological features, and collect and organize long-sequence meteorological historical data, including rainfall at rain gauge stations, regional weather system characteristic indicators, and regional teleconnection climate factors.

[0056] S2. Determination of Influencing Factors: Establishing a matrix of rainfall influencing factors for the study area. Furthermore, factors with significant magnitude deviations were normalized. The expression for the rainfall impact factor matrix is ​​as follows:

[0057] Ψ a × b = [ j 11 j 12 ... j 1 i j 1 b j 21 j 22 ... j 2 i j 2 b ... ... ... ... ... j k 1 j k 2 ... j ki j kb j a 1 j a 2 ... j ai j ab ] ;

[0058] in, Indicates the first The first partition The values ​​of each influence factor; Indicates the first The first partition The values ​​of each influence factor;

[0059] The normalization formula is as follows:

[0060] ;

[0061] in, For each partition The minimum value of each influencing factor; For each partition The maximum value of each influencing factor; The normalized factor data is in the range [0,1].

[0062] After normalization, the rainfall influencing factor matrix is ​​obtained. The expression is:

[0063] Ψ ' a × b = [ j ' 11 j ' 12 ... j ' 1 i j ' 1 b j ' 21 j ' 22 ... j ' 2 i j ' 2 b ... ... ... ... ... j ' k 1 j ' k 2 ... j ' ki j ' kb j ' a 1 j ' a 2 ... j ' ai j ' ab ] ;

[0064] S3. Correlation Analysis: Statistical methods were used to analyze the strength of the relationship between rainfall and different factors in the influencing factor matrix. Factors with significant temporal and spatial correlations were selected as key influencing factors, and the key influencing factors for each region were recorded in the matrix. The expression is:

[0065] Ω a × m = [ q ' 11 q ' 12 ... q ' 1 i q ' 1 m q ' 21 q ' 22 ... q ' 2 i q ' 2 m ... ... ... ... ... q ' k 1 q ' k 2 ... q ' ki q ' km q ' a 1 q ' a 2 ... q ' ai q ' am ] ;

[0066] in, For the first The first partition The normalized values ​​of each influence factor;

[0067] S4. Starting from any time, for any partition, denote the daily rainfall for the next N days as a vector. ,in, For the first The daily rainfall over the past N days is used to perform cluster analysis on the daily rainfall over the next N days. The total number of clusters is denoted as . Category number is denoted as ,in We obtain different types of rainfall events in the next N days, and statistically analyze the characteristics of key influencing factors in the next N days. We then establish and label the correlation between various rainfall events and key influencing factors.

[0068] S5. When conducting rainfall forecasting, based on the key influencing factors for the next N days, match the corresponding rainfall process category, and select the first category from the next N days... The average daily rainfall over the past few days will be used as the basis for this forecast. The outcome tendency value for the day, in the first category of the next N days. The minimum and maximum daily rainfall values ​​for each day will be used as the basis for this forecast. The lower and upper limits of the daily results are calculated using the following formula:

[0069] ;

[0070] ;

[0071] ;

[0072] in, This represents the number of samples in the first category. For the first category, the One sample, ; For the first The first sample Daily rainfall; For the first Daily rainfall trend value for the day; For the first The maximum daily rainfall on the [day name] is the [number]th [day name]. The daily rainfall limit for a given day; For the first The minimum daily rainfall on the 1st day is the 2nd day. The lower limit of daily rainfall;

[0073] S6. Starting from the calculation start time, output the tendency value of the daily rainfall forecast results for the next N days as the forecast result of this rainfall process, output the lower limit and upper limit of the daily rainfall forecast results for the next N days as the forecast interval of the rainfall process, and slide forward with a daily step size as time goes by, carry out forecast operations day by day, output the corresponding forecast results, and realize the rolling rainfall forecast.

[0074] In S2, the method for selecting the influencing factor is as follows:

[0075] For rainfall data from rain gauge stations, the Thiessen polygon method is used to calculate the areal rainfall in different zones based on the density of the monitoring station network and the distribution of the stations.

[0076] Based on the characteristic indicators of regional weather systems, factors that affect the spatial structure of meteorological elements, such as temperature, humidity, wind speed, and wind direction, are selected.

[0077] For regional teleconnection climate factors, atmospheric circulation index, sea ice, sea surface temperature, solar radiation, and Indian Ocean Dipole were selected.

[0078] In S3, the statistical method employs a combination of path analysis and improved Lasso regression to determine the direct and indirect relationships between different influencing factors and rainfall, analyze the comprehensive impact of multiple factors on rainfall, and screen key factors through regularization to establish a quantitative relationship model between rainfall and influencing factors. Specifically:

[0079] S31. Using path analysis, analyze the relationship between different influencing factors and rainfall. The calculation formula is as follows:

[0080] ;

[0081] in,

[0082] R= R a × n = [ y 11 y 12 ... y 1 i y 1 n y 21 y 22 ... y 2 i y 2 n ... ... ... ... ... y k 1 y k 2 ... y ki y kn y a 1 y a 2 ... y ai y an ] ; Ψ ' = Ψ ' a × b = [ j ' 11 j ' 12 ... j ' 1 i j ' 1 b j ' 21 j ' 22 ... j ' 2 i j ' 2 b ... ... ... ... ... j ' k 1 j ' k 2 ... j ' ki j ' kb j ' a 1 j ' a 2 ... j ' ai j ' ab ] ; δ= δ b × 1 = [ δ 1 δ 2 ... δ v δ b ] ; Τ = T 1 × n = [ y 11 y 12 ... y 1 w ′ y 1 n ] ;

[0083] in, The daily rainfall matrix for the study area is denoted as . , , All are temporary variables. , ; This is the influence factor matrix; This is the error term matrix; Indicates the first The path coefficients of each influencing factor; This indicates the number of rainfall events in any given rainfall event in the first zone. Rainfall for the day;

[0084] No. The first partition The expression for "day" is:

[0085] ;

[0086] In the formula, ζ= ζ 1 × b = [ j ' a 1 j ' a 2 ... j ' aj j ' ab ] For the first Influence factor matrix of each partition; Indicates the first The first partition Daily rainfall;

[0087] S32. Using an improved Lasso regression and a regularization method, key influencing factors are screened. Taking the forecast of rainfall over the next N days for any given region as an example, the specific steps are as follows:

[0088] ;

[0089] in, Indicates the first partition of any partition Rainfall amount on any day during a rainfall event; For the first The first rainfall event One influencing factor; This represents the intercept term of the regression model, i.e., the constant term; In the regression model, the first... The regression coefficients of each influencing factor represent the degree of influence of that factor on rainfall. Indicates the number of rainfall events; Indicates the number of impact factors; Represents the regularization parameter; This represents a parameter that controls the regularization ratio. γ = [ 0 , 1 ] ; The variable to be optimized is a set of parameters, including the intercept term of the regression model. In the regression model, the first Regression coefficients of each influencing factor ; This is the path coefficient matrix of the influencing factors;

[0090] S33. Establish a quantitative relationship model between rainfall and influencing factors, and identify key influencing factors, including quantity. and its category, where, .

[0091] In S4, starting from any time, the daily rainfall for the next N days is forecasted, and based on operational needs, the future rainfall is forecasted... Rainfall at different time scales per month; correspondingly, the units for cluster analysis and sliding step size are also in months.

[0092] In S4, the cluster analysis uses the k-means clustering method, spectral clustering, or splitting clustering method.

[0093] When the rainfall influencing factors in the study area are simple, the sample data is small, and the requirements for clustering categories are not high, the k-means clustering method is selected.

[0094] When the factors influencing rainfall in the study area are complex and diverse, and there is a large amount of sample data, spectral clustering is selected.

[0095] When the factors influencing rainfall in the study area are complex and diverse and the results have hierarchical requirements, split clustering is selected.

[0096] A rainfall forecasting system based on sliding clustering of key influencing factors includes: at least one processor; and a memory communicatively connected to at least one of the processors; wherein,

[0097] The memory stores instructions that can be executed by the processor to implement a rainfall forecasting method based on sliding clustering of key influencing factors.

[0098] Taking a hydrological control station on a tributary in the Three Gorges section of the upper Yangtze River as an example, data from January 2021 to December 2022 were selected to conduct rolling forecast calculations of daily rainfall from 1 to 7 days, in order to verify the feasibility and effectiveness of the method of the present invention.

[0099] Figure 2 The data plotted the cluster analysis results of the daily rainfall for the next 7 days at the current time of the hydrological station and stored the key influencing factors corresponding to each category. Table 1 shows the number of samples in different categories and their respective cluster results. Table 2 shows the clustering results of a certain rainfall forecast operation and other rainfall events of the same category.

[0100] When conducting rainfall forecasting operations, the current time is used as the forecast start time. Regional rainfall influencing factors corresponding to that time are extracted, and the method of this invention is used to identify key influencing factors. Then, based on this series of influencing factors, the k-means clustering method is used for matching. Figure 2The key influencing factors corresponding to various results are identified, and their corresponding rainfall process categories are extracted as the initial solution for this rainfall forecast. All rainfall processes within each category are statistically analyzed, and the average rainfall on the same day is used as the rainfall tendency value for that forecast day. The maximum and minimum rainfall on the same day are used as the maximum and minimum rainfall values ​​for that forecast day, respectively, thus completing one forecast operation. By gradually rolling forward with a one-day step size, daily rainfall forecasts for any number of days can be completed.

[0101] Table 1. Moving clustering results of daily rainfall over 1-7 days

[0102]

[0103] Table 2. Case Study: Rainfall Forecast Results (Unit: mm)

[0104]

[0105] As shown in Table 2, the method of this invention can quickly obtain the daily rainfall forecast for the next 7 days at the current moment, determine its trend value, maximum value and minimum value, and has high forecast accuracy and good process fit with the actual situation.

[0106] Through the above embodiments, daily rainfall forecasts for the next N days can be quickly achieved for the research object. The method of this invention first constructs a long-sequence meteorological database by collecting natural geographical and hydrological meteorological data of the study area; secondly, it constructs a rainfall influencing factor matrix and normalizes factors with large magnitude deviations to screen out key influencing factors; thirdly, it uses statistical methods to analyze the spatiotemporal correlation between rainfall and key factors to identify significantly correlated key factors; then, it performs cluster analysis on historical rainfall processes, statistically analyzes the key factor characteristics of various rainfall processes, and establishes a correlation between rainfall processes and influencing factors; finally, it matches the corresponding rainfall process categories based on the key factors for the next N days, determines the trend value, lower limit, and upper limit of the forecast results, and updates the forecast results on a daily basis. This method, through a combination of data-driven and statistical analysis, achieves a detailed characterization and rolling forecast of rainfall processes, has high applicability and forecast accuracy, and proves the feasibility and effectiveness of this method. Therefore, this method has superior application effects in daily rainfall forecasting.

[0107] As can be seen from the above analysis, the method of the present invention is highly practical and can effectively improve the accuracy of daily rainfall forecasts at different time scales.

[0108] In summary, this invention has the advantages of practicality and strong operability, and can quickly realize daily rainfall forecasts at different time scales with high forecast accuracy, providing a more scientific, efficient and accurate new method for watershed rainfall forecasting.

[0109] The embodiments described above are merely illustrative of implementation methods of the present invention, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these modifications and improvements all fall within the scope of protection of the present invention. Therefore, the scope of protection of this patent should be defined by the appended claims.

Claims

1. A rainfall forecasting method based on sliding clustering of key influencing factors, characterized in that, include: S1. Basic data preparation: Select the study area, determine the detailed zoning based on natural geographical characteristics and hydro-meteorological features, and collect and organize long-sequence meteorological historical data, including rainfall at rain gauge stations, regional weather system characteristic indicators, and regional teleconnection climate factors. S2. Determination of Influencing Factors: Establishing a matrix of rainfall influencing factors for the study area. Furthermore, factors with significant magnitude deviations were normalized. The expression for the rainfall impact factor matrix is ​​as follows: ; in, Indicates the first The first partition The values ​​of each influence factor; Indicates the first The first partition The values ​​of each influence factor; The normalization formula is as follows: ; in, For each partition The minimum value of each influencing factor; For each partition The maximum value of each influencing factor; The normalized factor data is in the range [0,1]. After normalization, the rainfall influencing factor matrix is ​​obtained. The expression is: ; S3. Correlation Analysis: Statistical methods were used to analyze the strength of the relationship between rainfall and different factors in the influencing factor matrix. Factors with significant temporal and spatial correlations were selected as key influencing factors, and the key influencing factors for each region were recorded in the matrix. The expression is: ; in, For the first The first partition The normalized values ​​of each influence factor; S4. Starting from any time, for any partition, denote the daily rainfall for the next N days as a vector. ,in, For the first The daily rainfall over the past N days is used to perform cluster analysis on the daily rainfall over the next N days. The total number of clusters is denoted as . Category number is denoted as ,in We obtain different types of rainfall events in the next N days, and statistically analyze the characteristics of key influencing factors in the next N days. We then establish and label the correlation between various rainfall events and key influencing factors. S5. When conducting rainfall forecasting, based on the key influencing factors for the next N days, match the corresponding rainfall process category, and select the first category from the next N days... The average daily rainfall over the past few days will be used as the basis for this forecast. The outcome tendency value for the day, in the first category of the next N days. The minimum and maximum daily rainfall values ​​for each day will be used as the basis for this forecast. The lower and upper limits of the daily results are calculated using the following formula: ; ; ; in, This represents the number of samples in the first category. For the first category, the One sample, ; For the first The first sample Daily rainfall; For the first Daily rainfall trend value for the day; For the first The maximum daily rainfall on the [day name] is the [number]th [day name]. The daily rainfall limit for a given day; For the first The minimum daily rainfall on the 1st day is the 2nd day. The lower limit of daily rainfall; S6. Starting from the calculation start time, output the tendency value of the daily rainfall forecast results for the next N days as the forecast result of this rainfall process, output the lower limit and upper limit of the daily rainfall forecast results for the next N days as the forecast interval of the rainfall process, and slide forward with a daily step size as time goes by, carry out forecast operations day by day, output the corresponding forecast results, and realize the rolling rainfall forecast. In S2, the method for selecting the influencing factor is as follows: For rainfall data from rain gauge stations, the Thiessen polygon method is used to calculate the areal rainfall in different zones based on the density of the monitoring station network and the distribution of the stations. Based on the characteristic indicators of regional weather systems, factors that affect the spatial structure of meteorological elements, such as temperature, humidity, wind speed, and wind direction, are selected. For regional teleconnection climate factors, atmospheric circulation index, sea ice, sea surface temperature, solar radiation, and Indian Ocean Dipole were selected. In S3, the statistical method employs a combination of path analysis and improved Lasso regression to determine the direct and indirect relationships between different influencing factors and rainfall, analyze the comprehensive impact of multiple factors on rainfall, and screen key factors through regularization to establish a quantitative relationship model between rainfall and influencing factors. Specifically: S31. Using path analysis, analyze the relationship between different influencing factors and rainfall. The calculation formula is as follows: ; in, ; ; ; ; in, The daily rainfall matrix for the study area is denoted as . , , All are temporary variables. , ; This is the influence factor matrix; This is the error term matrix; Indicates the first The path coefficients of each influencing factor; This represents the rainfall matrix for any given rainfall event in any given region. This indicates the number of rainfall events in any given rainfall event in the first zone. Rainfall for the day; No. The first partition The expression for "day" is: ; In the formula, For the first Influence factor matrix of each partition; Indicates the first The first partition Daily rainfall; S32. Using an improved Lasso regression and a regularization method, key influencing factors are screened. Taking the forecast of rainfall over the next N days for any given region as an example, the specific steps are as follows: ; in, Indicates the first partition of any partition Rainfall amount on any day during a rainfall event; For the first The first rainfall event One influencing factor; This represents the intercept term of the regression model, i.e., the constant term; In the regression model, the first... The regression coefficients of each influencing factor represent the degree of influence of that factor on rainfall. Indicates the number of rainfall events; Indicates the number of impact factors; Represents the regularization parameter; This represents a parameter that controls the regularization ratio. ; The variable to be optimized is a set of parameters, including the intercept term of the regression model. In the regression model, the first Regression coefficients of each influencing factor ; This is the path coefficient matrix of the influencing factors; S33. Establish a quantitative relationship model between rainfall and influencing factors, and identify key influencing factors, including quantity. and its category, where, ; In S4, starting from any time, the daily rainfall for the next N days is forecasted, and based on operational needs, the future rainfall is forecasted... Rainfall at different time scales per month; correspondingly, the units for cluster analysis and sliding step size are also in months.

2. The rainfall forecasting method based on sliding clustering of key influencing factors according to claim 1, characterized in that: In S4, the cluster analysis uses the k-means clustering method, spectral clustering, or splitting clustering method. When the rainfall influencing factors in the study area are simple, the sample data is small, and the requirements for clustering categories are not high, the k-means clustering method is selected. When the factors influencing rainfall in the study area are complex and diverse, and there is a large amount of sample data, spectral clustering is selected. When the factors influencing rainfall in the study area are complex and diverse and the results have hierarchical requirements, split clustering is selected.

3. A rainfall forecasting system based on sliding clustering of key influencing factors, characterized in that, include: At least one processor; and a memory communicatively connected to at least one of the processors; wherein, The memory stores instructions that can be executed by the processor to implement the rainfall forecasting method based on sliding clustering of key influencing factors as described in any one of claims 1 to 2.

Citation Information

Patent Citations

  • Statistics downscaling method based on SVM algorithm

    CN103838979A

  • Traditional rainfall forecast space-time quantification method conforming to rainfall space-time distribution characteristics

    CN118409372A

  • Carbon emission prediction method and system based on electric power big data

    CN119721336A