A multi-mode dynamic integrated precipitation prediction method and system suitable for small and medium rivers in a region
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ANHUI METEOROLOGICAL STATION
- Filing Date
- 2026-05-18
- Publication Date
- 2026-08-07
AI Technical Summary
[0003]区域中小河流流域地形复杂、下垫面异质性强,降水具有局地性强、变率大、突发性强的特点,极易引发山洪、滑坡等次生灾害,对降水预报的精准度、时效性提出了极高要求,而当前单一数值预报模式难以满足业务需求
[0043] This invention provides a multi-model dynamic integrated precipitation forecasting method and system applicable to small and medium-sized rivers in a region. Based on sub-basin division, sliding training period, dynamic ranking of TS scores, neighborhood selection, and heavy rain/light rain threshold elimination technology, it systematically and hierarchically divides small and medium-sized watersheds and establishes a mapping relationship between grid points in numerical model forecasts and the divided sub-basins. This allows for the generation or extraction of precipitation forecasts for each sub-basin, enhancing the ability to capture extreme precipitation events and the timeliness of flood warnings for small and medium-sized rivers.
Smart Images

Figure CN122525692A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of numerical model forecasting technology, specifically relating to a multi-model dynamic integrated precipitation forecasting method and system applicable to small and medium-sized rivers in the region. Background Technology
[0002] The frequent occurrence of localized, sudden, and short-duration heavy rainfall events caused by global climate change in some regions has become a major trigger for floods in small and medium-sized rivers, flash floods, and urban waterlogging, introducing uncertainty into flood control forecasting and emergency response. Currently, the starting point of the flood forecasting chain is precipitation forecasting, which mainly relies on numerical weather prediction models.
[0003] The region's small and medium-sized river basins have complex topography and highly heterogeneous underlying surfaces. Precipitation is characterized by strong locality, high variability, and suddenness, making it highly susceptible to secondary disasters such as flash floods and landslides. This places extremely high demands on the accuracy and timeliness of precipitation forecasts, which current single numerical weather prediction models cannot meet operational needs. On the one hand, the fixed physical parameterization schemes and grid resolutions of single models cannot accurately capture the spatiotemporal evolution characteristics of severe convective precipitation, leading to missed or false alarms, and they also struggle to accurately characterize the impact of local small-scale circulation on precipitation within the basin. On the other hand, different numerical weather prediction models have varying forecasting advantages, and a single model cannot meet the forecasting needs of various precipitation scenarios. Traditional static ensemble methods do not consider the dynamic differences in model performance, resulting in limited ensemble effectiveness. Currently, research on multi-model dynamic ensembles for small and medium-sized river basins remains lacking, and targeted ensemble methods are scarce. Summary of the Invention
[0004] To address the problems existing in the prior art, this invention provides a multi-mode dynamic integrated precipitation forecasting method and system applicable to small and medium-sized rivers in a region.
[0005] The specific technical solution of the present invention is as follows:
[0006] A multi-model dynamic integrated precipitation forecasting method applicable to small and medium-sized rivers in a region includes the following steps:
[0007] The target area is determined, and the river network information of the target area is extracted based on the DEM digital elevation model. The target area is divided into several sub-basins according to the small and medium-sized rivers determined by the river network information. The sub-basins are divided according to the set grid spacing.
[0008] Historical precipitation forecast data and corresponding historical real-time data of various numerical weather prediction models for the target area during the flood season are obtained, interpolated to the set grid spacing, and the data corresponding to the latitude and longitude of each grid point in each sub-basin are taken as the matching model data and the matching real-time data.
[0009] Based on the matched real-time data, the TS score of each matched model data is tested and sorted according to the TS score, so as to obtain the dynamic performance sequence of each forecast time of each model in different sub-basins and different sliding training periods at different magnitudes.
[0010] Based on the dynamic performance sequence, a hierarchical optimization method related to precipitation level is used to determine the grid point assignment strategy in each sub-basin.
[0011] Based on the grid point assignment strategy in each sub-basin, a multi-model dynamic integrated precipitation forecast model is constructed. The model takes the current precipitation forecast data of various numerical forecast models as input and outputs precipitation forecasts for small and medium-sized rivers.
[0012] Furthermore, a hierarchical optimization method related to precipitation levels is adopted to determine the grid point assignment strategy in the coordinate sequence of each sub-basin, specifically as follows:
[0013] Read data from multiple numerical weather prediction models at a specific grid point and sequentially determine the following precipitation levels.
[0014] The rainstorm level determination is to check whether there are two or more patterns with rainfall levels of rainstorm level in the top 3 pattern data. If so, the grid point is determined to be rainstorm level; otherwise, it is determined to be heavy rain level.
[0015] The heavy rainfall level is determined by checking whether there are two or more models with heavy rainfall levels in the top three model data. If so, the grid point is determined to be heavy rainfall level; otherwise, it is determined to be moderate rainfall level.
[0016] For moderate rainfall determination, check whether there are two or more models with moderate rainfall in the top three model data. If so, the grid point is determined to be moderate rainfall; otherwise, proceed to light rainfall determination.
[0017] To determine the level of light rain, we need to check whether there are two or more models with the level of light rain in the top three model data. If so, the grid point is determined to be at the level of light rain; otherwise, it is determined to be at the level of no rain.
[0018] The neighborhood selection method is used to determine the determination and assignment strategy for heavy rain or torrential rain levels; the point-to-point selection method is used to determine the determination and assignment strategy for moderate rain or light rain levels.
[0019] Furthermore, during the process of determining the intensity of heavy rainfall, heavy rainfall atmospheric clearance processing is performed; during the process of determining the intensity of light rainfall, light rainfall atmospheric clearance processing is performed.
[0020] Furthermore, the determination and assignment strategy for the rainfall intensity level (heavy rain or torrential rain) is determined using a neighborhood selection method, specifically as follows:
[0021] When determining the level of heavy rainfall, read the dynamic performance ranking of heavy rainfall levels and the corresponding model forecast values. Use the set grid spacing as the unit grid spacing and determine whether the number of models with heavy rainfall levels within the unit grid spacing (0.05° grid spacing) around the grid point in the top 3 models with heavy rainfall level TS scores is ≥2. If it is determined to be heavy rainfall level, select the maximum forecast value of the model that is determined to be heavy rainfall level and ranks first as the value of the grid point.
[0022] When determining the level of heavy rainfall, the dynamic performance ranking of heavy rainfall levels and the corresponding model forecast values are read. The set grid spacing is used as the unit grid spacing. It is determined whether the number of models with heavy rainfall levels within the unit grid spacing (0.05° grid spacing) around the grid point in the top 3 models with heavy rainfall TS scores is ≥2. If so, the top 3 models are assigned model weights according to the proportion of TS score, and the grid point is assigned a value according to the model weights and model forecast values.
[0023] Furthermore, the determination and assignment strategy for moderate or light rain levels adopts a point-to-point optimal selection method, specifically as follows:
[0024] When determining whether rainfall is moderate or light, the dynamic performance ranking and corresponding model forecast values for that rainfall level are read.
[0025] First, determine whether the number of patterns with moderate or light rain at this grid point is ≥2 among the top 3 patterns in the corresponding magnitude TS score. If so, assign pattern weights to the top 3 patterns according to the proportion of the corresponding magnitude TS score, and determine the assignment strategy for this grid point based on the pattern weights and the pattern forecast values.
[0026] Furthermore, the strategy for determining the grid point's value based on the mode weights and mode prediction values is calculated using the following formula:
[0027] Pattern weights are calculated using the following formula:
[0028]
[0029] In the formula, For the first The weight of each pattern, For the first TS scores for each pattern For the first TS scores for each pattern Take 1, 2, and 3 respectively. Choose from 1 to N, and set N to 3;
[0030] The grid point is assigned the following values:
[0031]
[0032] In the formula, The ensemble predictions for the grid points. For the first The predicted values for each model.
[0033] Furthermore, when data is missing in the dynamic performance sequence, the missing value is replaced by data from the base sequence. The missing data refers to the number of samples at a certain grid point being 0, or the TS score of that grid point being 0.
[0034] Furthermore, the construction method of the basic sequence is as follows: obtain multi-year historical precipitation forecast data and corresponding historical real data of various numerical weather prediction models during the flood season, interpolate the data to the set grid interval, and obtain grid data corresponding to the latitude and longitude of each sub-basin as the matched data. Based on the matched real data, perform TS score verification on the matched model data and sort them according to the TS score. The model forecast performance ranking of each forecast time for each forecast lead of each model, different sub-basins, and different sliding training periods is obtained as the basic sequence.
[0035] A multi-model dynamic integrated precipitation forecasting system applicable to small and medium-sized rivers in a region includes the following modules:
[0036] The sub-basin division module determines the target area, extracts river network information of the target area based on the DEM digital elevation model, and divides the target area into several sub-basins according to the small and medium-sized rivers determined by the river network information. The sub-basins are divided according to the set grid spacing.
[0037] The data matching module acquires historical precipitation forecast data and corresponding historical real-time data of various numerical forecast models for the target area during the flood season, interpolates them to the set grid spacing, and takes the data corresponding to the latitude and longitude of each grid point in each sub-basin as the matching model data and the matching real-time data.
[0038] The dynamic performance sequence construction module performs TS score verification on the matched real-time data and sorts the data according to the TS score to obtain the dynamic performance sequences of each forecast time for different sub-basins and different sliding training periods of each model at different magnitudes.
[0039] The assignment module, based on the dynamic performance sequence, uses a hierarchical optimization method related to precipitation level to determine the assignment strategy of grid points in each sub-basin;
[0040] The integrated forecasting module constructs a multi-model dynamic integrated precipitation forecasting model based on the grid point assignment strategy in each sub-basin. It takes the current precipitation forecast data from various numerical forecasting models as input and outputs precipitation forecasts for small and medium-sized rivers.
[0041] An electronic device includes a memory and a processor, the memory storing a computer program, and the processor being configured to invoke and run the computer program stored in the memory to perform the methods described above.
[0042] Compared with the prior art, the present invention has the following advantages:
[0043] This invention provides a multi-model dynamic integrated precipitation forecasting method and system applicable to small and medium-sized rivers in a region. Based on sub-basin division, sliding training period, dynamic ranking of TS scores, neighborhood selection, and heavy rain / light rain threshold elimination technology, it systematically and hierarchically divides small and medium-sized watersheds and establishes a mapping relationship between grid points in numerical model forecasts and the divided sub-basins. This allows for the generation or extraction of precipitation forecasts for each sub-basin, enhancing the ability to capture extreme precipitation events and the timeliness of flood warnings for small and medium-sized rivers.
[0044] This invention employs a tiered optimization forecasting strategy related to precipitation intensity, using different strategies for different precipitation levels. For heavy rain and above, a neighborhood selection method is used, comprehensively considering the forecasts of multiple high-performing models in the region to better capture heavy precipitation. For example, if more than two of the top three models in the neighborhood forecast heavy rain, the grid point is determined to be a heavy rain area, and the forecast value of the best-performing model among those forecasting heavy rain is assigned. This method maximizes the preservation of heavy precipitation signals from each model, avoiding the loss of important early warning information due to missed forecasts by a single model. For moderate and light rain, a point-to-point TS scoring method is used, which typically has relatively uniform spatial distribution, weak intensity, and wide range, requiring less precise accuracy at the boundaries of precipitation areas compared to heavy precipitation. The neighborhood selection method is not suitable for small precipitation levels as it increases the false alarm rate.
[0045] This invention uses TS scoring to construct a dynamic performance sequence and then determines the grid assignment strategy. Compared with other scoring methods such as ETS, TS scoring is simpler to calculate and more sensitive to changes in forecast bias. In the dynamic ensemble algorithm, using TS scoring as the target can reduce missed and false alarms for key events. It is an innovative method that directly embeds business scoring into the optimization engine.
[0046] This invention addresses the challenges of segmenting small and medium-sized river basins, their small spatial scale, and the prominent local characteristics of precipitation. It innovatively decentralizes the calculation of Threshold Test (TS) scores to each sub-basin unit. By conducting multi-model TS performance evaluations on a sub-basin and precipitation-level basis, a three-dimensional evaluation system of model-sub-basin-precipitation-level is constructed, enabling refined zoning and hierarchical dynamic integrated modeling. This overcomes the inherent limitations of traditional research that assigns global and uniform weights to single models, and expands the application of TS scores in the weight allocation of different models.
[0047] This invention integrates forecast information from multiple independent models to form a superior multi-model integrated forecast product, significantly improving forecast skill. In most cases, the integrated result outperforms any single model involved in the integration in key indicators such as average error and skill score, and this advantage typically covers multiple forecast lead times. Simultaneously, it effectively reduces uncertainty and systematic errors. Different numerical models differ in physical parameterization schemes, initial conditions, and resolutions, leading to varying systematic errors and uncertainties in their forecasts. Multi-model integration can comprehensively utilize the advantages of each model, smoothing out the extreme errors of individual models, thereby reducing the systematic bias of the overall forecast and improving its stability and reliability.
[0048] This invention studies a multi-model dynamic integrated precipitation forecasting method for small and medium-sized rivers in the region. By dynamically integrating the advantages of multiple numerical forecasting models and optimizing the integration strategy, the accuracy and timeliness of precipitation forecasts are improved. This has important practical significance and application value for improving the disaster prevention and mitigation technology system for small and medium-sized rivers and reducing disaster risks. Attached Figure Description
[0049] Figure 1 This is a flowchart of the present invention;
[0050] Figure 2 A schematic diagram of the grid point range for the neighborhood selection method;
[0051] Figure 3 A comparison chart of MAR values for different precipitation levels and forecast times with a 24-hour forecast lead time;
[0052] Figure 4 A comparison chart of FAR values for different precipitation levels and forecast times with a 24-hour forecast lead time;
[0053] Figure 5 A comparison chart of BIAS values for different precipitation levels and forecast times with a 24-hour forecast lead time;
[0054] Figure 6 A comparison chart of TS scores for 12-hour forecast lead time for heavy rainfall;
[0055] Figure 7 A comparison chart of TS scores for 6-hour forecast lead time for heavy rainfall;
[0056] Figure 8 A comparison chart of TS scores for 3-hour forecast lead time for heavy rainfall;
[0057] Figure 9 The distribution of false alarm rates for the ECMWF 0.1mm probability forecast across different probability intervals from 24h to 024h is shown.
[0058] Figure 10The sample size distribution map for each probability interval of the ECMWF 0.1mm probability forecast from 24h to 024h is shown.
[0059] Figure 11 Comparison of precipitation from various forecast products and actual precipitation from CLDAS for sub-basin a from 08:00 on June 23 to 08:00 on June 24. Detailed Implementation
[0060] The present invention will be further described below with reference to specific examples.
[0061] I. Data Source:
[0062] 1. Real-time data
[0063] The real-time data used in this invention comes from the China Meteorological Administration's Land Surface Data Assimilation System (CLDAS V2.0). Since its operational launch in 2015, this system has achieved the fusion of multi-source observations from the ground, radar, and satellite. The resulting gridded real-time product has a spatial resolution of 0.0625° × 0.0625° (approximately 6.25 km) and a temporal resolution of 1 hour, covering elements such as air temperature, air pressure, relative humidity, wind speed, precipitation, and shortwave radiation. This study selects CLDAS precipitation data and accumulates it to obtain real-time data for various forecast lead times, including 3 hours, 6 hours, 12 hours, and 24 hours.
[0064] 2. Model precipitation forecast data
[0065] This invention selects five commonly used operational numerical model precipitation forecast products as the base numerical models for integration: CMA-MESO (China Meteorological Administration Regional Scale Numerical Prediction System), CMA-GFS (China Meteorological Administration Global Medium-Range Weather Forecast System), CMA-SH9 (China Meteorological Administration Shanghai Numerical Prediction Model System), ECMWF (European Centre for Medium-Range Weather Forecasts), and NCEP (National Center for Environmental Prediction, USA Global Forecast System). Details are as follows (Table 1):
[0066] CMA-MESO: In 2001, the China Meteorological Administration initiated the development of GRAPES (Global / Regional Assimilation and Prediction System). It first achieved operational use in 2006 with a horizontal resolution of 30km. With improvements in observational data and computing resources, the model gradually incorporated high spatiotemporal resolution data from radar and satellites, progressively increasing the resolution to 10km between 2010 and 2015, and completing the brand renaming from GRAPES-MESO to CMA-MESO. Subsequently, CMA-MESO underwent a 3km operational upgrade between 2018 and 2020, achieving 1-hour cycle updates for the first time. In July 2024, CMA-MESO's horizontal grid spacing reached 1km.
[0067] CMA-GFS: CMA-GFS is a global numerical weather prediction and assimilation system independently developed by my country. Since its first operational deployment in 2009, the system has undergone several upgrades, including V3.3, V4.0, and V4.2. The grid spacing has been increased from the initial 0.25° (≈25km) to 0.125° (≈12.5km). Version V4.2 was officially launched at the end of 2024, with 89 vertical layers and a top layer of approximately 60km.
[0068] CMA-SH9: The Shanghai Numerical Weather Prediction Model System of the China Meteorological Administration (CMA-SH9) is an improved and upgraded version of the first-generation East China Regional Numerical Weather Prediction Model System. It began operational trial operation in June 2014 and is a high-resolution model for local areas such as East China and North China, with a horizontal resolution of 9km×9km.
[0069] ECMWF: The ECMWF high-resolution product uses a horizontal grid spacing of 0.125°×0.125° (approximately 12.5km). This model is based on a non-hydrodynamic core, and the physical processes include multi-spectral microphysics, improved convection parameterization, and a full-spectrum radiation scheme. It has a 51-layer vertical structure, with the top layer of the model at approximately 0.1hPa.
[0070] NCEP-GFS: NCEP-GFS is one of the earliest and most widely used medium-range forecasting models in the world. Its operational products include two resolutions: 0.5°×0.5° (approximately 50km) and 0.25°×0.25° (approximately 25km). Table 1 shows the models and their corresponding spatial resolutions, with a forecast lead time of up to 384 hours.
[0071]
[0072] Furthermore, the ECMWF's 50mm heavy rainfall probability forecast product was used in the heavy rainfall depletion method. ECMWF provides heavy rainfall probability forecast products through its ensemble forecast system (EPS), which includes precipitation probability (PP) at the grid scale, providing percentage probabilities for different thresholds (e.g., ≥1mm, 5mm, 10mm, 20mm, 50mm). This study uses the probability forecast product above 50mm with a spatial resolution of 0.5°.
[0073] II. Threshold-based emptying
[0074] The process of clearing the skies after heavy rain is as follows:
[0075] To improve the accuracy of heavy rainfall forecasts, a "ghosting" process is performed on the rainfall magnitude, which involves removing forecast samples where heavy rainfall is highly unlikely. Using historical ECMWF heavy rainfall probability forecasts and actual observation data, a ghosting threshold is used as a screening boundary. If a forecast does not reach this threshold, it is automatically excluded and no longer considered a valid heavy rainfall forecast. Numerical weather forecast products are then corrected, including the following steps:
[0076] Data pairing and sample preparation; probability interval division and statistics; calculation of false alarm rate; and determination of false alarm threshold.
[0077] The process of light rain dissipating from the air is as follows:
[0078] To improve the accuracy of light rain forecasts, a "ghost removal" process is performed on light rain intensity, which involves eliminating forecast samples where light rain is highly unlikely. The ghost removal threshold for light rain forecasts is determined using ECMWF precipitation forecast products, following the same steps as for heavy rain ghost removal.
[0079] Example 1:
[0080] like Figure 1 As shown, the present invention provides a multi-model dynamic integrated precipitation forecasting method applicable to small and medium-sized rivers in a region, comprising the following steps:
[0081] The target area is determined, and the river network information of the target area is extracted based on the DEM digital elevation model. The target area is divided into several sub-basins according to the small and medium-sized rivers determined by the river network information. The sub-basins are divided according to the set grid spacing.
[0082] Historical precipitation forecast data and corresponding historical real data of various numerical weather prediction models during the flood season are obtained for the target area, interpolated to the set grid spacing, and the data corresponding to the latitude and longitude of each grid point in each sub-basin are taken as the matching model data and the matching real data.
[0083] Based on the matched real-time data, the TS score of each matched model data is tested and sorted according to the TS score, so as to obtain the dynamic performance sequence of each forecast time of each model in different sub-basins and different sliding training periods at different magnitudes.
[0084] Based on dynamic performance sequences, a hierarchical optimization method related to precipitation levels is used to determine the grid point assignment strategy in each sub-basin.
[0085] A multi-model dynamic integrated precipitation forecasting model is constructed based on the grid point assignment strategy in each sub-basin. The model takes the current precipitation forecast data from various numerical forecasting models as input and outputs precipitation forecasts for small and medium-sized rivers.
[0086] This invention selects the TS score to evaluate the performance of various models. The design combines practicality and strategy optimization. In precipitation forecast evaluation, simply using "accuracy" (PC = (number of hits + number of correct rejections) / total number of samples) will lose its discriminative power due to too many "rainless" days. Using other complex scores is not easy to directly convert into model weights. At the same time, the core of the TS score is to focus on the forecasting ability of "whether an event has occurred" (e.g., ≥50mm precipitation) at a certain magnitude. This selection innovatively aligns the evaluation indicators with the core research objective (forecasting disastrous precipitation), avoiding the dispersion of evaluation resources. This allows subsequent dynamic weight optimization to fully improve the forecasting skills for key thresholds, thereby improving the accuracy of the integrated forecast.
[0087] Example 2:
[0088] This example is further designed to use a hierarchical optimization method related to precipitation levels to determine the grid point assignment strategy in the coordinate sequence of each sub-basin, specifically:
[0089] Read data from multiple numerical weather prediction models at a specific grid point and sequentially determine the following precipitation levels.
[0090] The rainstorm level determination is to check whether there are two or more patterns with rainfall levels of rainstorm level in the top 3 pattern data. If so, the grid point is determined to be rainstorm level; otherwise, it is determined to be heavy rain level.
[0091] The heavy rainfall level is determined by checking whether there are two or more models with heavy rainfall levels in the top three model data. If so, the grid point is determined to be heavy rainfall level; otherwise, it is determined to be moderate rainfall level.
[0092] For moderate rainfall determination, check whether there are two or more models with moderate rainfall in the top three model data. If so, the grid point is determined to be moderate rainfall; otherwise, proceed to light rainfall determination.
[0093] To determine the level of light rain, we need to check whether there are two or more models with the level of light rain in the top three model data. If so, the grid point is determined to be at the level of light rain; otherwise, it is determined to be at the level of no rain.
[0094] The neighborhood selection method is used to determine the determination and assignment strategy for heavy rain or torrential rain levels; the point-to-point selection method is used to determine the determination and assignment strategy for moderate rain or light rain levels.
[0095] Heavy rain intensity determination is processed by clearing up the atmosphere during the heavy rain intensity determination process; light rain intensity determination is processed by clearing up the atmosphere during the light rain intensity determination process.
[0096] Example 3:
[0097] This example is further designed so that the determination and assignment strategy for the rainfall intensity level (heavy rain or torrential rain) adopts the neighborhood selection method, specifically as follows:
[0098] When determining the level of heavy rainfall, read the dynamic performance ranking of heavy rainfall levels and the corresponding model forecast values. Use the set grid spacing as the unit grid spacing and determine whether the number of models with heavy rainfall levels within the unit grid spacing (0.05° grid spacing) around the grid point in the top 3 models with heavy rainfall level TS scores is ≥2. If it is determined to be heavy rainfall level, select the maximum forecast value of the model that is determined to be heavy rainfall level and ranks first as the value of the grid point.
[0099] When determining the level of heavy rainfall, the dynamic performance ranking of heavy rainfall levels and the corresponding model forecast values are read. The set grid spacing is used as the unit grid spacing. It is determined whether there are ≥2 models with heavy rainfall levels within the unit grid spacing (0.05° grid spacing) around the grid point among the top 3 models with heavy rainfall TS scores. If so, the top 3 models are assigned model weights according to the proportion of TS score, and the grid point is assigned a value according to the model weights and model forecast values.
[0100] Neighborhood selection method for grid point range selection:
[0101] like Figure 2 As shown, the neighborhood selection method searches within a 0.05° (approximately 5km) grid interval around a grid point (forecast grid point). The diagram typically includes the forecast grid point, four grid points at the midpoints of each side of the rectangle, and four grid points at each corner of the rectangle, for a total of nine grid points. Assuming that the top three models in the sub-basin for a given precipitation level are Model a1, Model a2, and Model a3, the search is performed on these nine grid points. If more than two of the top three models forecast heavy rainfall, then the grid point is considered to be at the heavy rainfall level. For each model, if one of the nine grid points reaches the heavy rainfall level, then that model is considered to be at the heavy rainfall level.
[0102] Example 4:
[0103] This example is further designed to use a point-to-point optimal selection method to determine and assign values for moderate or light rainfall levels, specifically as follows:
[0104] When determining whether rainfall is moderate or light, the dynamic performance ranking and corresponding model forecast values for that rainfall level are read.
[0105] First, determine whether the number of patterns with moderate or light rain at this grid point is ≥2 among the top 3 patterns in the corresponding magnitude TS score. If so, assign pattern weights to the top 3 patterns according to the proportion of the corresponding magnitude TS score, and determine the assignment strategy for this grid point based on the pattern weights and the pattern forecast values.
[0106] The assignment strategy for this grid point is determined based on the model weights and model prediction values, calculated using the following formula:
[0107] Pattern weights are calculated using the following formula:
[0108]
[0109] In the formula, For the first The weight of each pattern, For the first TS scores for each pattern For the first TS scores for each pattern Take 1, 2, and 3 respectively. Choose from 1 to N, and set N to 3;
[0110] The grid point is assigned the following values:
[0111]
[0112] In the formula, The ensemble predictions for the grid points. For the first The predicted values for each model.
[0113] For example, when integrating the five forecast models of this invention, taking the calculation of the heavy rainfall level weight as an example, during a certain sliding training period, the TS score of each model in the heavy rainfall level of a certain grid point in a certain sub-basin is calculated. Assuming that the top three models are ranked as a, b, and c, and the forecast values of the grid point for models a, b, and c are all heavy rainfall levels, namely Pa, Pb, and Pc respectively, the weight of model a is calculated as follows, and the weights of models b and c can be obtained in the same way.
[0114]
[0115] The integrated forecast value for that grid point:
[0116] .
[0117] Example 5:
[0118] This example is further designed so that when data is missing in the dynamic performance sequence, data from the base sequence is used to replace the missing values. Missing data means that the number of samples at a certain grid point is 0, or the TS score of that grid point is 0.
[0119] The basic sequence is constructed as follows: multi-year historical precipitation forecast data and corresponding historical real data of various numerical weather prediction models during the flood season are obtained, the data are interpolated to a set grid interval, and grid data corresponding to the latitude and longitude of each sub-basin are obtained as matched data. The matched model data are tested by TS score based on the matched real data and sorted according to the TS score. The model forecast performance ranking of each forecast time for each forecast lead of each model, different sub-basin, and different sliding training period is obtained as the basic sequence.
[0120] Example 6:
[0121] A multi-model dynamic integrated precipitation forecasting system applicable to small and medium-sized rivers in a region includes the following modules:
[0122] The sub-basin division module determines the target area, extracts river network information of the target area based on the DEM digital elevation model, and divides the target area into several sub-basins according to the small and medium-sized rivers determined by the river network information. The sub-basins are divided according to the set grid spacing.
[0123] The data matching module acquires historical precipitation forecast data and corresponding historical real-time data from various numerical weather prediction models for the target area during the flood season, interpolates them to a set grid spacing, and takes the data corresponding to the latitude and longitude of each grid point in each sub-basin as the matching model data and the matching real-time data.
[0124] The dynamic performance sequence construction module performs TS score verification on the matched real-time data and sorts the data according to the TS score to obtain the dynamic performance sequences of each forecast time for different sub-basins and different sliding training periods of each model at different magnitudes.
[0125] The assignment module, based on dynamic performance sequences, uses a hierarchical optimization method related to precipitation levels to determine the assignment strategy for grid points in each sub-basin;
[0126] The integrated forecasting module constructs a multi-model dynamic integrated precipitation forecasting model based on the grid point assignment strategy in each sub-basin. It takes the current precipitation forecast data from various numerical forecasting models as input and outputs precipitation forecasts for small and medium-sized rivers.
[0127] An electronic device includes a memory and a processor. The memory stores a computer program, and the processor is used to call and run the computer program stored in the memory to perform the methods described above.
[0128] Application examples:
[0129] I. Division of Sub-basins
[0130] For example, taking a certain region as the target area, the target area is further subdivided into 49 rivers and 92 sub-basins. This target area spans three major basins: the first basin covers 67,200 km², the second 66,700 km², and the third 6,200 km², accounting for 48.0%, 47.6%, and 4.4% of the target area, respectively. Among the rivers in the target area, there are 901 small and medium-sized rivers with a basin area of over 50 km², including 214 rivers with a basin area of 200–3000 km², 25 major tributaries with a basin area of over 3000 km², and 128 lakes with a basin area of 1 km² or more. Given the complex natural water system and severe flood control situation in the target area, a systematic and hierarchical division of the small and medium-sized river basins is implemented. 214 rivers are too many for actual forecasting services, while 25 rivers cannot meet the needs of flood control. Therefore, taking into account factors such as topography (watershed), natural distribution of water systems, and administrative divisions, the target area was refined into 49 rivers based on the characteristics of the topography and water system and the needs of flood control services. The more important small and medium-sized rivers were further subdivided, and the target area was finally refined into 92 sub-basins of 49 rivers. The sub-basins were divided according to the set grid spacing, which was 0.05° (about 5km).
[0131] II. Dynamic Performance Sequence
[0132] This invention selects historical precipitation forecast data and historical real-time data from five numerical weather prediction models during the 2024 flood season (June-September). Before testing, the data is preprocessed and interpolated to a set grid spacing (the grid spacing used for sub-basin division). For different sub-basins, different sliding training periods (7d, 14d, 21d, 28d, 35d, where d is short for days), different forecast lead times (3h, 6h, 12h, 24h), different forecast times (forecast times are the same as the base sequence), and different precipitation levels (light rain, moderate rain, heavy rain, torrential rain), daily Threat Score (TS) scores are calculated and dynamic TS tests are performed to obtain real-time dynamic performance sequences according to different sub-basins, different sliding training periods, different forecast lead times, different forecast times, and different precipitation levels.
[0133] III. Basic Sequence
[0134] This invention selects historical data from five numerical weather prediction models during the flood season (June-September) of 2020-2023 as the base series for performance evaluation. Before testing, the data is preprocessed and interpolated to a set grid spacing (the grid spacing used for sub-basin division). Forecast lead times cover four levels: 3h, 6h, 12h, and 24h. Specifically, the 3h forecast lead time includes 16 forecast times: 003, 006, 009, 012, 015, 018, 021, 024, 027, 030, 033, 036, 039, 042, 045, and 048; the 6h forecast lead time includes 006, 012, 018, 024, 030, 036, 042, 048, 054, 0... The forecast includes 12 forecast times: 060, 066, 072, etc.; 12-hour forecasts include 10 forecast times: 012, 024, 036, 048, 060, 072, 084, 096, 108, 120, etc.; and 24-hour forecasts include 9 forecast times: 024, 036, 048, 060, 072, 084, 096, 108, 120, etc. Historical CLDAS grid data from the 2020-2023 flood season (June-September) were used for validation. TS scoring was applied at different precipitation levels (light rain, moderate rain, heavy rain, and torrential rain) to obtain basic sequences according to different sub-basins, different sliding training periods, different forecast times, different forecast times, and different magnitudes, providing supplementary data for subsequent dynamic performance sequences when they are missing.
[0135] Application Example 1:
[0136] This invention provides verification and analysis of integrated forecast products with multiple forecast lead times, multiple forecast durations, and different training periods.
[0137] The integrated forecast products with five dynamic sliding training periods (7d, 14d, 21d, 28d, and 35d, where 'd' is an abbreviation for days) of the method of this invention were tested and analyzed against five single numerical weather prediction model products (ECMWF, CMA-SH9, CMA-MESO, CMA-GFS, and NCEP). The main testing methods were TS, ETS, false alarm rate, missed alarm rate, and bias. The five testing methods for 24-hour forecast leads were analyzed in detail, while the analysis for 12-hour, 6-hour, and 3-hour forecast leads mainly focused on TS scores.
[0138] I. 24-hour forecast lead time verification analysis
[0139] 1. TS score:
[0140] As shown in Table 1, the integrated product demonstrates an advantage in light rain forecasts at 24-hour intervals: 0.6189 for 21-day and 28-day forecasts, 0.6184 for 35-day forecasts, 0.6125 for 7-day forecasts, and 0.619 for 14-day forecasts, all higher than all single models (the highest being CMA-SH9 at 0.5846, CMA-GFS at 0.5679, ECMWF at 0.5024, and NCEP at 0.4886). As the forecast duration increases, the integrated product score slowly decreases, but remains around 0.547 at 84-hour intervals (28-day and 35-day forecasts), significantly higher than ECMWF (0.5187) and CMA-GFS (0.4958). At 120-hour intervals, the integrated product score is mostly in the 0.494-0.502 range, still outperforming most single models. This demonstrates that the integrated approach has a sustained advantage in forecasting weak precipitation coverage, especially with a slower attenuation as the lead time increases.
[0141] Table 1. TS Scores of Integrated Forecast Products and Individual Model Products for Light Rainfall Levels During Training Periods
[0142] (Bold text indicates the highest level; the same applies to the following tables.)
[0143]
[0144] For moderate rainfall (Table 2), among the single models over 24 hours, ECMWF (0.5285) and NCEP (0.513) were relatively high, followed by CMA-GFS (0.518), while CMA-SH9 (0.4627) and CMA-MESO (0.4513) were relatively low. Among the integrated products, 7 days had an ECMWF of 0.5209, which was slightly lower than ECMWF and ranked second. The ECMWF values for 14 days (0.5172), 21 days (0.514), 28 days (0.5114), and 35 days (0.512) were all within 1.5% of the optimal single model, and their distribution was concentrated with small fluctuations. By 48 hours, integrated products outperformed most single-mode models, with 21-day, 28-day, and 35-day results all above 0.497. ECMWF decreased to 0.4758, CMA-GFS was 0.4518, and NCEP was 0.4791, indicating that integrated products performed better during this period. For longer durations (84–120 hours), integrated products remained in the 0.38–0.41 range. Although the overall value decreased, it was still generally higher than that of single-mode models, demonstrating the advantage of integrated models in extending the duration of moderate-intensity precipitation.
[0145] Table 2 shows the TS scores of integrated forecast products and individual model products for different rainfall levels during each training period.
[0146]
[0147] For heavy rainfall (Table 3), NCEP (0.4486) was the highest over 24 hours, followed by ECMWF (0.4427), while CMA-GFS (0.4055), CMA-MESO (0.3894), and CMA-SH9 (0.3822) were lower. The integrated product's 7-day (0.4385) values were close to the ECMWF level, while the 14-day values were slightly lower (0.4307). The 21-day (0.4284), 28-day (0.4224), and 35-day (0.4211) values were about 3–6% lower than the optimal single model, and they were close to each other, indicating high stability. As the forecast duration increased, the scores of individual models generally declined and fluctuated significantly. For example, the ECMWF score dropped from 0.4427 to 0.2715 at 84 hours, and the CMA-GFS score dropped to 0.2823. The integrated product, however, maintained a score of 0.333–0.337 at 84 hours, outperforming both CMA-GFS and ECMWF. At 108–120 hours, the integrated product score mostly remained in the 0.24–0.34 range, higher than that of individual models, indicating that the integrated approach can effectively suppress the rapid decay of individual models in long-term forecasts of heavy precipitation.
[0148] Table 3. TS Scores of Integrated Forecast Products and Individual Model Products for Heavy Rainfall Levels During Each Training Period
[0149]
[0150] Rainfall intensity (Table 4): Rainfall intensity is the level most likely to trigger meteorological risks in small and medium-sized rivers. For the 24-hour period, all five integrated forecast products outperformed all individual models, indicating that fusion over multiple training periods can improve the probability of capturing extreme precipitation events. Extending the timeframe to 36–48 hours, the integrated product scores further improved (0.3243 for 14 days, 0.3236 for 21 days, 0.32 for 28 days, and 0.3123 for 35 days), significantly outperforming ECMWF (0.2799→0.2583) and CMA-GFS (0.2376→0.2582), and even better than NCEP (0.2468→0.1929). During the long-term period, although the integrated product gradually decreased, it still remained at about 0.17-0.18 at 84h, which is higher than CMA-GFS (0.1532) and NCEP (0.1242), indicating that the availability of basic integrated forecasts for rainstorms is significantly higher than that of single models.
[0151] Table 4. TS Scores of Integrated Forecast Products and Individual Model Products for Rainfall Intensity During Each Training Period
[0152]
[0153] In summary, the integrated product demonstrated more balanced skill and stronger robustness across all four precipitation levels: it comprehensively outperformed individual models in light rain forecasts, ensuring high coverage; in moderate and heavy rain forecasts, it showed minimal and more stable performance compared to the best individual model, avoiding the significant shortcomings of individual models in certain timeframes; and in heavy rain forecasts, it outperformed individual models, highlighting the excellent applicability of the integrated approach to heavy rain forecasting. In contrast, among individual models, ECMWF and NCEP showed some advantages in moderate and heavy rain, but exhibited greater fluctuations; CMA-GFS was superior in the early stages of light and moderate rain but decayed rapidly over longer periods; the integrated method, through a sliding training period, retained the high skill of heavy precipitation members while compensating for the deficiencies of weak precipitation members, and decayed more slowly during the extended lead time, making it an effective means to improve the overall forecast reliability.
[0154] 2. ETS Score: Light Rainfall (Table 5), similar to TS, the integrated product's advantage is most prominent and stable. Throughout the entire forecast period from 24h to 120h, the ETS scores of all integrated products ranked in the top five almost all timeframes, and repeatedly achieved the highest score (in red). At 24h, the 14-day product ranked first with a score of 0.619, and the scores of the other integrated products were also very close (0.6189, 0.6189, 0.6184, 0.6125), significantly higher than the best-performing single model CMA-SH9 (0.5846). As the forecast lead time increased, the leading advantage of the integrated products remained obvious. At 84h, the integrated product score was around 0.5475, while the better-performing single models ECMWF and CMA-GFS were 0.5187 and 0.4958, respectively. This indicates that the ensemble method significantly improves the forecast accuracy for light precipitation (with or without rain) by fusing information from multiple training periods, and this advantage is maintained throughout the entire 5-day forecast period.
[0155] Table 5. ETS Scores of Integrated Forecast Products and Individual Model Products for Light Rainfall Levels During Each Training Period
[0156]
[0157] For moderate rainfall (Table 6), the integrated product's advantages become apparent in the later stages. At 24 hours, the best product was the single-model ECMWF (0.5285), closely followed by the integrated product 7-days (0.5209), with other integrated product scores very close. However, starting from the 48-hour forecast, the integrated product began to outperform the others. At 48 hours, the integrated product (21-days: 0.4975) surpassed all single models (the best being NCEP: 0.4791). By 120 hours, the integrated product score stabilized at around 0.38, while the best-performing single-model ECMWF had dropped to 0.3447. This indicates that for moderate rainfall, the integrated product is comparable to the best model in short-term forecasts, but becomes a better choice in medium-term forecasts due to its superior stability.
[0158] The intensity of heavy rainfall is similar to that of moderate rainfall, showing a pattern of similar intensity in the short term but leading in the middle and later stages.
[0159] Table 6. ETS Scores of Integrated Forecast Products and Individual Model Products for Moderate to Heavy Rainfall Levels During Each Training Period
[0160]
[0161] Regarding the intensity of heavy rainfall (Table 7), TS and ETS are similar, with the ensemble showing a significant advantage over single models. The top three scores for each forecast time are all in the ensemble product. Even in the later forecast stages (96-108h), the ensemble product (e.g., 35 days at 108h: 0.1800) maintains relatively high skill, while single models such as NCEP have dropped below 0.1. This indicates that ensemble forecasting methods have better ability to capture extreme precipitation events and more persistent forecast skill.
[0162] Table 7. ETS Scores of Integrated Forecast Products and Individual Model Products for Each Training Period at Different Rainfall Intensities
[0163]
[0164] 3. Missing report rate (e.g.) Figure 3 , Figure 4 (as shown)
[0165] In terms of light rain intensity, the single-model NCEP and ECMWF performed best in terms of missed detection rate (MAR), with the lowest values (0.0298 for NCEP and 0.048 for ECMWF in 24h). CMA-SH9 and CMA-MESO had significantly higher MARs. All integrated products had MARs between these two categories, significantly lower than CMA-SH9 and CMA-MESO, but higher than NCEP and ECMWF. In terms of false alarm rate (FAR), the opposite was true: NCEP and ECMWF had the highest FARs (0.504 and 0.4846 respectively in 24h), meaning they generated a large number of invalid light rain forecasts; CMA-SH9 had a relatively low FAR (0.3501 in 24h). Integrated products had significantly lower FARs than NCEP and ECMWF, comparable to or even better than CMA-SH9 (approximately 0.311-0.313 for 24h integrated products). In summary, while integrated products are slightly inferior to NCEP and ECMWF in capturing all light rain events (with extremely low false negatives), they achieve higher forecast efficiency and quality by significantly reducing unnecessary false negatives and avoiding false negatives.
[0166] For moderate rainfall events, integrated products demonstrated a significant advantage in reducing the false alarm rate. Across all forecast durations, the false alarm rate (MAR) of integrated products was significantly lower than that of any single model. The 24-hour false alarm rate of integrated products ranged from 0.1594 to 0.174, while single models generally exceeded 0.24 (ECMWF 0.242, CMA-SH9 0.2734). This indicates that integrated products more effectively reflected actual moderate rainfall events. Regarding false alarm rate (FAR), integrated products and single models showed varying degrees of accuracy. Between 24 and 48 hours, ECMWF and CMA-GFS had the lowest false alarm rates, with integrated products slightly higher; however, after 72 hours, the false alarm rate of integrated products was comparable to, or even lower than, ECMWF and NCEP (e.g., at 120 hours, integrated products had approximately 0.517-0.5235, lower than ECMWF's 0.5266 and CMA-GFS's 0.5501). This indicates that for moderate rain, the integrated product achieved a significant reduction in the missed detection rate with a slight increase in the acceptable false alarm rate, thus significantly improving the ability to detect precipitation events.
[0167] During periods of heavy rainfall, integrated products demonstrate a significant advantage in reducing missed reports. At all times, the missed report rate (MAR) of integrated products is far lower than any single model. Over 24 hours, the integrated product has the lowest MAR at 0.151 (35 days), while the best single model, CMA-SH9, has a MAR of 0.2641, and ECMWF and NCEP are as high as 0.3517 and 0.4271, respectively. This advantage persists throughout the forecast period. Regarding false alarm rate (FAR), in the short term (24-36 hours), NCEP has the lowest false alarm rate (0.3261-0.3467), while integrated products have the highest (0.5062-0.5448). However, as the lead time increases, after 84 hours, the false alarm rate of integrated products tends to be close to or even lower than that of ECMWF and CMA-GFS (e.g., in the 120-hour period, the false alarm rate of integrated products is about 0.619-0.6338, which is lower than ECMWF's 0.66 and CMA-GFS's 0.6687).
[0168] For heavy rainfall events with severe impacts, the integrated product demonstrated the greatest advantage in terms of missed detection rate (MAR). In almost all timeframes, the integrated product exhibited the lowest MAR. For example, in the 24-hour timeframe, the integrated product had the lowest MAR at 0.2343 (35 days), while the best single-model CMA-SH9 had an MAR of 0.3408 and an ECMWF as high as 0.5793. This means that the integrated product missed far fewer extreme precipitation events than single-model forecasts. In terms of false alarm rate (FAR), the integrated product is on the same order of magnitude as the better-performing single models (such as ECMWF and NCEP). For example, at 24h, NCEP has the lowest false alarm rate (0.4211), while the integrated product has a higher rate (0.6609-0.6843). However, at 72h, the NCEP false alarm rate drops sharply to 0.4008, which is lower than the integrated product (approximately 0.6949-0.7131). By 108h, the false alarm rates of each product are close (ECMWF 0.7417, integrated product approximately 0.7259-0.7268).
[0169] ④Bias score (e.g.) Figure 5 (As shown):
[0170] For light rain, the integrated product demonstrated the best bias control capability. Across all forecast periods (24-120 hours), the bias values of all integrated products were highly concentrated, remaining stable between 1.16 and 1.30 (e.g., approximately 1.23-1.25 for 24 hours and approximately 1.20-1.25 for 120 hours). This range is closest to the ideal value of 1, indicating that the integrated product provided the most appropriate forecast for light rain. In contrast, the biases of individual models were more extreme and dispersed: NCEP and ECMWF showed significant over-reporting, with their biases consistently high at 1.60-2.02 throughout the forecast period; CMA-GFS also showed some over-reporting (1.42-1.50); CMA-SH9 (approximately 1.31-1.46), while better than NCEP and ECMWF, was still higher than the integrated product; the CMA-MESO value was also relatively high.
[0171] For moderate rainfall, the integrated product's bias showed a stable, moderately positive bias. Across all timeframes, the bias value of the integrated product was also highly concentrated, ranging from approximately 1.33 to 1.57, with a slight increasing trend as the training period lengthened (from 7 days to 35 days). This indicates that the integrated product tends to slightly overreport the frequency or extent of moderate rainfall. The performance of individual models varied significantly: the biases of ECMWF and NCEP ranged from 1.11 to 1.25, slightly above or close to 1; the bias of CMA-GFS fluctuated more (1.03-1.30); and the CMA-SH9 data were also relatively high (approximately 1.25-1.35).
[0172] During heavy rainfall, the bias values of integrated products were concentrated and significantly greater than 1, ranging from 1.29 to 1.87 (higher in the early stages, decreasing to 1.06-1.17 in later stages, such as 120h), showing a clear positive bias. Single-mode: CMA-SH9 showed strong overreporting (up to 1.66); while NCEP and CMA-GFS frequently showed underreporting (bias mostly less than 1, as low as 0.59); ECMWF's bias was relatively moderate but fluctuating, ranging from 1.07 to 1.26.
[0173] The most pronounced contrasts were observed in the biases for heavy rainfall intensity. The integrated product showed a bias significantly greater than 1 across all timeframes (except 120h), ranging from 0.78 to 2.43. Particularly in the 24-72h forecasts, the bias was generally above 1.37, indicating a strong tendency to over-report. Individual models showed polarized performance: CMA-SH9 exhibited extreme over-reporting (bias as high as 2.24); while NCEP and CMA-GFS showed severe under-reporting (bias frequently below 0.75, even as low as 0.25); ECMWF's bias was between 0.92 and 1.20, relatively closest to 1 but with slightly insufficient forecast coverage.
[0174] II. 12-hour forecast lead time (TS) verification analysis
[0175] A comparative analysis was conducted on the TS scores of various products (ECMWF, CMA-SH9, CMA-MESO, CMA-GFS, NCEP, 7-days, 14-days, 21-days, 28-days, and 35-days) for different forecast times of 12-hour forecast lead for heavy rainfall intensity. Figure 6 As shown in the figure, the models generally score higher in shorter forecast periods (such as around 12 hours), and then decrease as the forecast period lengthens. The integrated forecast products show relatively consistent TS scores across different time periods, and in most time periods, they are higher than or close to the best-performing single model, demonstrating their robustness and advantages in heavy rainfall forecasting.
[0176] In the 24-hour forecast, ensemble products generally achieved scores of around 0.30 or higher, even surpassing ECMWF and CMA-SH9. The ensemble method, through multi-training-period fusion, suppressed the extreme high or low scores caused by sample fluctuations in single models, resulting in a smoother and more predictable scoring curve. As the forecast period extended from 12 hours to 36 hours, the TS score of the ensemble products decreased more slowly, remaining generally better than single models. This indicates that the ensemble products can still maintain effective identification of heavy rain events even after extending the forecast lead time, reducing the problem of rapid skill decay due to extended lead times. Although the ensemble products ranked high in TS scores in most forecasts, in the shortest forecast period of 12 hours, CMA-MESO ranked first, and the five ensemble products ranked second. This is because ensembles, in the process of multi-member fusion, compromise the extreme advantages of individual members, meaning that in short-term scenarios where the best single member performs well, it may not necessarily surpass that member.
[0177] III. 6-hour forecast lead time verification analysis
[0178] A comparative analysis was conducted on the TS scores of various products (ECMWF, CMA-SH9, CMA-MESO, CMA-GFS, NCEP, 7-days, 14-days, 21-days, 28-days, and 35-days) for different forecast times of 6-hour forecast lead for heavy rainfall intensity. Figure 7As shown in the figure, only CMA-MESO (0.2798) outperforms the integrated product at the 6-hour forecast lead time. At the 12-hour forecast lead time, CMA-SH9 (0.1157) and CMA-GFS (0.1268) are relatively high, with the integrated product ranking in the middle, from third to seventh. ECMWF (0.0518), NCEP (0.0659), and CMA-MESO (0.0845) are all not high enough. At the 18-hour forecast lead time, ECMWF and NCEP rank first and second, with the integrated forecast following closely behind, all outperforming other single models. The integrated forecast maintains a relative advantage over single models from 24 to 36 hours. Therefore, at the 6-hour forecast lead time, each model has a relatively high score at shorter forecast times (6hTS), and the score gradually decreases as the forecast lead time increases. This is consistent with the general rule that short-term forecasts are more skillful due to more accurate initial fields. Compared with the single model, the integrated product maintained a relatively stable high score in most cases, and showed a slight upward trend as the training period lengthened (from 7 days to 35 days), demonstrating the robustness of the integrated method in short-term heavy rain forecasting and avoiding the extreme low values and large fluctuations of the single model.
[0179] IV. 3-hour forecast lead time verification analysis
[0180] A comparative analysis was conducted on the TS scores of various products (ECMWF, CMA-SH9, CMA-MESO, CMA-GFS, NCEP, 7-days, 14-days, 21-days, 28-days, and 35-days) for different forecast times of 3-hour forecast lead for heavy rainfall intensity. Figure 8 As shown in the figure, the average TS of the integrated forecast products (7 days to 35 days) ranges from 0.110 to 0.114, with the highest TS for 7 days (0.11367), followed by 14 days (0.11207), 21 days (0.11143), 28 days (0.11105), and 35 days (0.11045). The TS shows a slight decreasing trend with increasing training period, but the overall level is close to and significantly higher than most individual models. Among the individual models, CMA-SH9 has the highest average TS (0.10839), followed by CMA-MESO (0.0946), NCEP (0.08658), and ECMWF (0.07438), while CMA-GFS has the lowest (0.05827). It is evident that the integrated products lead all individual models on average, and their internal performance is stable with relatively small differences.
[0181] The performance of the forecast times is as follows: The 3-12h integrated product can achieve a TS of 0.214 to 0.229 in the 3h time (0.2295 in the 7 days), which is significantly better than the single mode. CMA-SH9 and CMA-MESO also have good performance in the 3h time (0.2134 and 0.2529, respectively), but CMA-GFS is extremely low (0.0275). As the forecast time increased to 6-12 hours, the TS values of individual models generally decreased, while the integrated product remained in the range of 0.056-0.155, which was better than most individual models. From 15-36 hours, the TS value of the integrated product at the 18-hour time rose to 0.138-0.149. Among the individual models, CMA-SH9 and CMA-MESO fluctuated more significantly, while ECMWF and NCEP were relatively high at certain times during this period (e.g., NCEP reached 0.1845 at 18 hours). From 24-36 hours, the TS value of the integrated product remained stable in the range of 0.13-0.16, and reached its peak at 27 hours (approximately 0.1708 and 0.1701 at 28 days and 35 days, respectively), demonstrating the better ability of the integrated product to capture heavy rainfall events from 24 to 36 hours.
[0182] Application Example 2:
[0183] Analysis of the effect of rainstorm air dissipation
[0184] Taking the 24-hour forecast lead time (024 forecast time) of various models during the 2024 flood season as an example, based on historical statistics of ECMWF probability forecast products above 50mm, the threshold for rainstorm clearing was determined to be 2%. For integrated forecast products with different sliding training periods, the TS score was used to calculate the rainstorm level score, and the effect of the TS score before and after rainstorm clearing was analyzed. The results are shown in Table 8.
[0185] Table 8 Comparison of Rainfall Dissipation Effects During Sliding Training Periods at Different Rainfall Levels
[0186]
[0187] Analysis shows that the effect of clearing false alarms in heavy rain exhibits two main characteristics: First, clearing has a positive effect on improving the accuracy of heavy rain forecasts. The TS score after clearing is slightly higher than that before clearing in each sliding training period, with the difference remaining stable between 0.0002 and 0.0004. This indicates that clearing effectively suppresses false alarms in non-heavy rain areas while protecting the forecast sensitivity of heavy rain areas, thus achieving the core requirements of flood control forecasting: "reducing false alarms and ensuring timeliness." Second, the length of the sliding training period significantly affects the clearing effect. As the training period extends from 7 days to 35 days, the TS score after clearing shows a continuous decreasing trend, from 0.2984 to 0.2879. The score remains the same from 14 to 21 days (both are 0.2924), indicating that the improvement in clearing effect weakens significantly after the training period exceeds 14 days, while the clearing effect is optimal with a 7-day sliding training period. This indicates that a shorter training period is better suited to the time-sensitive characteristics of weather patterns during the flood season, enabling the atmospheric cancellation algorithm to more accurately capture changes in precipitation systems and improve the targeting of rainstorm forecasts. In summary, a 2% threshold for ECMWF rainstorm probability forecasts can effectively optimize forecast quality and achieve effective atmospheric cancellation of rainstorms.
[0188] Application Example 3:
[0189] Analysis of the air dissipation effect of light rain
[0190] This study used two types of data (ECMWF precipitation probability forecast products above 0.1 mm and ECMWF precipitation forecast products) and employed historical false alarm rate analysis and PSS analysis methods to obtain the light rain clearing thresholds. Taking a 24-hour forecast lead time of 0-24 hours as an example, the acquisition of the two light rain clearing thresholds was analyzed.
[0191] The method for obtaining the blanking threshold based on ECMWF probabilistic forecast products is as follows: Spatial interpolation is performed on the original ECMWF probabilistic forecast data to make it consistent with the actual grid spacing (the same as the grid spacing of the sub-basin division) to improve the matching accuracy. Then, the forecast probability is divided into multiple intervals (e.g., 0–2%, left open and right closed). The comparison between the forecast and the actual situation in each interval is statistically analyzed, and the blanking rate of each interval is calculated. The blanking rate is controlled below 95%, that is, about 5% of the forecasts are retained as high-confidence signals. The blanking threshold for light rain is obtained by analyzing various distributions such as the blanking rate distribution, sample size distribution, and blanking rate distribution.
[0192] Results Analysis: This analysis focuses on light rain forecasts from 24 hours to 24 hours, with a total sample size of 32,023,460.
[0193] The distribution of empty reports shows ( Figure 9The false alarm rate shows a clear decreasing trend across different probability intervals. In low probability intervals (e.g., 0-2%), the false alarm rate is extremely high, approaching 100%, meaning that almost all "rainy" forecasts in this interval are false. As the forecast probability increases, the false alarm rate gradually decreases. As the probability interval rises, the false alarm rate continues to decrease, reaching 24% in the 90-92% interval. The figure marks a reference line with a target false alarm rate of 95.0%, and based on this 95% target false alarm rate, the false alarm elimination threshold is set at 6%.
[0194] Further analysis of the sample size distribution shows that ( Figure 10 The samples are mainly concentrated in two intervals with extremely low and extremely high forecast probabilities, while the sample size in the intermediate probability interval is relatively small. With 6% as the threshold, the sample size in the low probability interval (≤6%) is limited (the sample size in the 0-2% interval is about 0.15e7), while the sample size in the high probability interval (such as 98-100%) is extremely large (close to 1.2e7), which is the main concentration area of the samples. This distribution shows that filtering the low probability (≤6%) interval will not significantly reduce the total sample size, and can accurately eliminate high false alarm rate samples.
[0195] Application Example 4:
[0196] Sub-basin a analysis
[0197] From late June to early July 2024, sub-basin a experienced a continuous precipitation process, with a basin-wide rainstorm on the 23rd. The maximum precipitation was recorded at Yiqi Station at 189.3 mm, and the average precipitation in the basin was 117.8 mm.
[0198] Analysis of individual forecast model products (CMA-MESO, CMA-GFS, CMA-SH9, NCEP, ECMWF) and dynamic training period 7-day integrated forecast products (since the forecast areas of the five dynamic training period integrated precipitation products are similar, only the 7-day forecast map is used for comparison here) compares the 24-hour forecast lead time and 0-24-hour forecast time precipitation areas with the actual precipitation observed by CLDAS. Figure 11 As shown, the actual rainfall was only moderate in some southern areas, while other areas experienced heavy rain or more, with most areas experiencing torrential rain. The CMA-SH9 forecast primarily predicted moderate rain, with heavy rain in the central basin and localized torrential rain, a significant difference from the actual rainfall. Both ECMWF and CMA-MESO forecast basin-wide heavy rain, but ECMWF primarily predicted heavy rain, with only torrential rain in some western areas, a much smaller area than the actual rainfall. CMA-MESO's forecast for torrential rain also differed significantly from the actual rainfall. The NCEP forecast map showed heavy rain in the central and eastern parts of the basin and moderate rain in the west, both smaller than the actual rainfall. Compared to individual models, the integrated forecast product's rainfall distribution was more similar to the actual rainfall, outperforming individual models.
[0199] Further analysis of the TS scores of various models and integrated forecast products for heavy rain and torrential rain levels is shown in Table 9. The five integrated models all scored the same (0.9815). For heavy rain levels, except for ECMWF, the scores of the other models were lower than the integrated products, with CMA-SH9 having the lowest score at only 0.0469. For torrential rain levels, CMA-SH9, CMA-GFS, and NCEP all scored highly, consistent with the localized forecast maps. ECMWF scored only 0.0323, and the highest score among the single models was CMA-MESO, at only 0.1606. Compared to the single-model integrated products, the scores for torrential rain levels were more ideal, all at 0.3404. This indicates that integrated forecasts have a certain predictive value for heavy rain, especially torrential rain.
[0200] Table 9 - 24-hour Forecast Lead Time (0:24:00) for June 23: Scores of Forecast Products for Rainstorm and Heavy Rainstorm Intensities
[0201]
[0202] This method constructs an intelligent decision-making framework of "real-time evaluation and dynamic adjustment." Based on the TS score and using the most recent performance (sliding window) as the evaluation criterion, this framework allows the integrated forecasting system to flexibly allocate resources daily according to the latest "state" of each model, thereby continuously and stably outputting the best integrated forecast product. Analysis of integrated forecast products and single models for different sliding training periods, forecast lead times, and forecast durations during the 2024 flood season (June-September) reveals that the 24-hour forecast lead integrated product exhibits more balanced forecast skill and stronger robustness across the four precipitation levels (light rain, moderate rain, heavy rain, and torrential rain). The integration method, through its sliding training period design, retains the high skill of heavy precipitation forecast members while compensating for the deficiencies of weak precipitation forecast members, resulting in a smoother decay during lead extension, making it an effective means of improving overall forecast reliability. For the 12-hour forecast lead, each model generally scores higher in shorter forecast durations (e.g., 12 hours), and the scores decrease as the forecast duration increases. The integrated forecast products showed strong stability in TS scores at various time points. The 3-hour and 6-hour forecast leads further verified the robustness of the integrated method in short-term heavy rainfall forecasts, effectively avoiding the extreme low values and large fluctuations that occur with single models. By selecting heavy rainfall cases from the above-mentioned sub-basins for analysis, it can be seen that the integrated forecast is significantly better than the single model in forecasting heavy rainfall and torrential rainfall.
Claims
1. A multi-model dynamic integrated precipitation forecasting method applicable to small and medium-sized rivers in a region, characterized in that, Includes the following steps: The target area is determined, and the river network information of the target area is extracted based on the DEM digital elevation model. The target area is divided into several sub-basins according to the small and medium-sized rivers determined by the river network information. The sub-basins are divided according to the set grid spacing. Historical precipitation forecast data and corresponding historical real-time data of various numerical weather prediction models for the target area during the flood season are obtained, interpolated to the set grid spacing, and the data corresponding to the latitude and longitude of each grid point in each sub-basin are taken as the matching model data and the matching real-time data. Based on the matched real-time data, the TS score of each matched model data is tested and sorted according to the TS score, so as to obtain the dynamic performance sequence of each forecast time of each model in different sub-basins and different sliding training periods at different magnitudes. Based on the dynamic performance sequence, a hierarchical optimization method related to precipitation level is used to determine the grid point assignment strategy in each sub-basin. Based on the grid point assignment strategy in each sub-basin, a multi-model dynamic integrated precipitation forecast model is constructed. The model takes the current precipitation forecast data of various numerical forecast models as input and outputs precipitation forecasts for small and medium-sized rivers.
2. The multi-model dynamic integrated precipitation forecasting method applicable to small and medium-sized rivers in a region according to claim 1, characterized in that, A hierarchical optimization method related to precipitation levels is used to determine the grid point assignment strategy in the coordinate sequence of each sub-basin, specifically as follows: Read data from multiple numerical weather prediction models at a specific grid point and sequentially determine the following precipitation levels. The rainstorm level determination is to check whether there are two or more patterns with rainfall levels of rainstorm level in the top 3 pattern data. If so, the grid point is determined to be rainstorm level; otherwise, it is determined to be heavy rain level. The heavy rainfall level is determined by checking whether there are two or more models with heavy rainfall levels in the top three model data. If so, the grid point is determined to be heavy rainfall level; otherwise, it is determined to be moderate rainfall level. For moderate rainfall determination, check whether there are two or more models with moderate rainfall in the top three model data. If so, the grid point is determined to be moderate rainfall; otherwise, proceed to light rainfall determination. To determine the level of light rain, we need to check whether there are two or more models with the level of light rain in the top three model data. If so, the grid point is determined to be at the level of light rain; otherwise, it is determined to be at the level of no rain. The neighborhood selection method is used to determine the determination and assignment strategy for heavy rain or torrential rain levels; the point-to-point selection method is used to determine the determination and assignment strategy for moderate rain or light rain levels.
3. The multi-model dynamic integrated precipitation forecasting method applicable to small and medium-sized rivers in a region according to claim 2, characterized in that, Heavy rain intensity determination is processed by clearing up the atmosphere during the heavy rain intensity determination process; light rain intensity determination is processed by clearing up the atmosphere during the light rain intensity determination process.
4. The multi-model dynamic integrated precipitation forecasting method applicable to small and medium-sized rivers in a region according to claim 3, characterized in that, The determination and assignment strategy for the rainfall intensity level (heavy rain or torrential rain) adopts the neighborhood selection method, specifically as follows: When determining the level of heavy rainfall, read the dynamic performance ranking of heavy rainfall levels and the corresponding model forecast values. Use the set grid interval as the unit grid interval and determine whether the number of models with heavy rainfall levels within the unit grid interval around the grid point is ≥2 among the top 3 models with heavy rainfall level TS scores. If it is determined to be heavy rainfall level, select the maximum forecast value of the model that is determined to be heavy rainfall level and ranks first as the value of the grid point. When determining the level of heavy rainfall, the dynamic performance ranking of heavy rainfall levels and the corresponding model forecast values are read. The set grid spacing is used as the unit grid spacing. It is determined whether the number of models with heavy rainfall levels within the unit grid spacing around the grid point in the top 3 models with heavy rainfall TS scores is ≥2. If so, the top 3 models are assigned model weights according to the proportion of TS score, and the grid point is assigned a value according to the model weights and model forecast values.
5. The multi-model dynamic integrated precipitation forecasting method applicable to small and medium-sized rivers in a region according to claim 4, characterized in that, The determination and assignment strategy for moderate or light rain levels adopts a point-to-point optimal selection method, specifically as follows: When determining whether rainfall is moderate or light, the dynamic performance ranking and corresponding model forecast values for that rainfall level are read. First, determine whether the number of patterns with moderate or light rain at this grid point is ≥2 among the top 3 patterns in the corresponding magnitude TS score. If so, assign pattern weights to the top 3 patterns according to the proportion of the corresponding magnitude TS score, and determine the assignment strategy for this grid point based on the pattern weights and the pattern forecast values.
6. The multi-model dynamic integrated precipitation forecasting method applicable to small and medium-sized rivers in a region according to claim 5, characterized in that, The strategy for determining the grid point's value based on model weights and model prediction values is calculated using the following formula: Pattern weights are calculated using the following formula: ; In the formula, For the first The weight of each pattern, For the first TS scores for each pattern For the first TS scores for each pattern Take 1, 2, and 3 respectively. Choose from 1 to N, and set N to 3; The grid point is assigned the following values: ; In the formula, For the integrated prediction values of the grid points, For the first The predicted values for each model.
7. The multi-model dynamic integrated precipitation forecasting method applicable to small and medium-sized rivers in a region according to claim 1, characterized in that, When data is missing in the dynamic performance sequence, the missing value is replaced by data from the base sequence. The missing data refers to the number of samples at a certain grid point being 0, or the TS score of that grid point being 0.
8. The multi-model dynamic integrated precipitation forecasting method applicable to small and medium-sized rivers in a region according to claim 1, characterized in that, The construction method of the basic sequence is as follows: obtain multi-year historical precipitation forecast data and corresponding historical real data of various numerical weather prediction models during the flood season, interpolate the data to the set grid interval, and obtain grid data corresponding to the latitude and longitude of each sub-basin as the matched data. Perform TS score verification on the matched model data according to the matched real data and sort them according to the TS score. The model forecast performance ranking of each forecast time for each forecast lead of each model, different sub-basins, and different sliding training periods is obtained as the basic sequence.
9. A multi-model dynamic integrated precipitation forecasting system applicable to small and medium-sized rivers in a region, characterized in that, Includes the following modules: The sub-basin division module determines the target area, extracts river network information of the target area based on the DEM digital elevation model, and divides the target area into several sub-basins according to the small and medium-sized rivers determined by the river network information. The sub-basins are divided according to the set grid spacing. The data matching module acquires historical precipitation forecast data and corresponding historical real-time data of various numerical forecast models for the target area during the flood season, interpolates them to the set grid spacing, and takes the data corresponding to the latitude and longitude of each grid point in each sub-basin as the matching model data and the matching real-time data. The dynamic performance sequence construction module performs TS score verification on the matched real-time data and sorts the data according to the TS score to obtain the dynamic performance sequences of each forecast time for different sub-basins and different sliding training periods of each model at different magnitudes. The assignment module, based on the dynamic performance sequence, uses a hierarchical optimization method related to precipitation level to determine the assignment strategy of grid points in each sub-basin; The integrated forecasting module constructs a multi-model dynamic integrated precipitation forecasting model based on the grid point assignment strategy in each sub-basin. It takes the current precipitation forecast data from various numerical forecasting models as input and outputs precipitation forecasts for small and medium-sized rivers.
10. An electronic device, characterized in that, The electronic device includes a memory and a processor, the memory storing a computer program, and the processor being configured to invoke and run the computer program stored in the memory to perform the method as described in any one of claims 1 to 9.