A method for predicting precipitation extreme based on multi-source historical data
By fusing multi-source historical data and performing real-time environmental field similarity analysis, the problem of systematic underestimation in traditional extreme precipitation forecasting has been solved, achieving high-quality extreme precipitation forecasting and providing accurate risk assessment and decision support.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-04-07
- Publication Date
- 2026-07-03
AI Technical Summary
Traditional numerical model forecasting methods suffer from systematic underestimation or underestimation in extreme precipitation forecasting. They lack physical mechanism support and fail to effectively integrate multi-source historical data, resulting in limited application value and interpretability of forecast products in disaster prevention and mitigation decision-making.
By establishing a multi-source historical data fusion database, we can identify historical extreme precipitation events in the target area, extract environmental field characteristics, generate a sample library, and calculate a similarity index based on the real-time environmental field to construct a probability density function for future precipitation intensity and quantify the risk probability under different return periods.
It significantly improves the physical foundation and historical traceability of the forecasting system, accurately quantifies the tail probability of low-probability extreme events, and generates structured, business-oriented forecasting products, providing scientific and practical quantitative risk basis for disaster prevention and mitigation decision-making.
Smart Images

Figure CN122332953A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of meteorological forecasting technology, specifically to a method for forecasting extreme precipitation values based on multi-source historical data. Background Technology
[0002] In existing extreme precipitation forecasting technologies, traditional numerical model forecasting methods are limited by model systematic errors and uncertainties in physical parameterization schemes. They are generally insufficient in forecasting low-probability, high-impact nonlinear events such as extreme precipitation, often resulting in systematic underestimation or underreporting. They often lack sufficient physical mechanism support, do not make enough use of actual observation sequences of similar historical extreme events, and find it difficult to establish a dynamic and quantitative correlation between the environmental field and the intensity of extreme precipitation. They also fail to effectively integrate multi-source historical data and quantify the risk probability under different return periods, resulting in limited direct application value and interpretability of forecast products in operational disaster prevention and mitigation decision-making. Summary of the Invention
[0003] To address the aforementioned technical problems, a method for predicting extreme precipitation values based on multi-source historical data is provided. This technical solution solves the problems mentioned above.
[0004] To achieve the above objectives, the technical solution adopted by the present invention is as follows: A method for predicting extreme precipitation values based on multi-source historical data includes: S1. Establish a multi-source historical data fusion database of reanalysis data and gridded precipitation products, identify historical extreme precipitation events in the target area, extract the environmental field characteristics of historical extreme precipitation, and generate a sample library of historical extreme events. S2. Obtain real-time multidimensional environmental field data of the target area, establish a structured set of environmental field feature variables, calculate the contribution weight of each feature variable to extreme precipitation, associate and pair historical extreme event sample databases, calculate the environmental field similarity index of historical-real-time extreme precipitation events in the target area, and establish a similarity set of extreme precipitation events. S3. Based on the similarity set of extreme precipitation events, extract the actual observed precipitation intensity sequence of each historical case within the target forecast period, construct the probability density function of future precipitation intensity in the target area, quantify the probability of precipitation extremes occurring at different return periods, and generate a precipitation extreme probability forecasting scheme.
[0005] Preferably, step S1 specifically includes: Collect high-resolution satellite radar fusion precipitation products of the target area over the past ten years, obtain MERRA-2 reanalysis data of the target area over the past forty years, extract three-dimensional field variables of dynamics, thermodynamics, and water vapor associated with precipitation, align timestamps, perform data preprocessing, and establish a spatiotemporally aligned reanalysis data-gridized precipitation product multi-source historical data fusion database. Based on a multi-source historical data fusion database of reanalysis data and gridded precipitation products, high spatiotemporal resolution precipitation data from the past thirty years were extracted, including CMORPH and GPCP gridded precipitation products. ERA5 and NCEP / NCAR reanalysis data from the same period were also obtained. Specific humidity, vertical velocity, wind field, geopotential height, convective available potential energy, and water vapor flux divergence were calculated to generate a set of historical environmental field variables for the target area. Bilinear interpolation was performed on gridded data from different sources and resolutions to unify the spatial grid and temporal frequency.
[0006] Preferably, step S1 further includes: Based on the reanalysis data-gridized precipitation product multi-source historical data fusion database, the annual precipitation sequences of all grid points in the target area are statistically analyzed, and the 95th percentile value of the sequence is calculated as the extreme precipitation threshold of the grid point. The historical precipitation data of the grid points are scanned daily. If the daily precipitation of a certain grid point is greater than or equal to the extreme precipitation threshold of the grid point, the grid point is marked as an extreme precipitation grid point on that day. The total area of all extreme precipitation grid points in the target area is calculated. The number of days with adjacent extreme precipitation grid points in time is defined as a single candidate event. For each candidate event, the spatial distribution of its daily extreme precipitation grid points is examined. Using the 8-neighborhood connectivity method, multiple spatially independent rain clusters are aggregated from the daily extreme precipitation grid points. If two candidate events are less than or equal to 2 days apart in time and the dominant rain clusters have significant spatial overlap, then these two candidate events are merged into one extreme precipitation event. A minimum duration t is set, and historical extreme precipitation events in the target area are identified to obtain their unique ID, start and end dates, spatial range, continuous area with the largest cumulative rainfall, maximum daily precipitation of all grid points during the event period, and average total rainfall of the target area. The significant folding includes: the distance between the centroid of the dominant rain cluster and the threshold h.
[0007] Preferably, step S1 further includes: For identified historical extreme precipitation events in the target area, the 24 hours before the event is selected as the precursor period. Corresponding environmental field variables are extracted, the spatial average value of each variable in the target area is calculated, and a multi-dimensional feature vector of historical extreme precipitation environmental field is generated. Combined with the corresponding actual precipitation intensity and occurrence time, a sample library of historical extreme events is generated.
[0008] Preferably, step S2 specifically includes: Based on the ECMWF weather forecast center, multidimensional environmental field grid data of the target area at the current time and the forecast field for the next 6-24 hours are obtained. The real-time environmental field variables are subjected to quality control and interpolation. The three-dimensional field variables of dynamics, thermodynamics and water vapor associated with precipitation are extracted. According to the precursor period, the spatial average value of each variable in the target area in the next 24 hours is calculated, and a set of structured characteristic variables of the environmental field is generated.
[0009] Preferably, step S2 further includes: Using environmental field characteristic variables from a historical extreme event sample database as input, the input data is sampled with replacement to generate multiple environmental field characteristic variable subsets. A decision tree is trained for each environmental field characteristic variable subset. When splitting nodes, some features are randomly selected to find the optimal split point, and an extreme precipitation prediction model is established. The hyperparameters of the number of decision trees and the maximum depth are optimized through grid search and cross-validation. The extreme precipitation observation label corresponding to each historical case is used as the output, with 1 indicating that extreme precipitation occurred and 0 indicating that extreme precipitation did not occur. The Gini importance score of each environmental field characteristic variable is calculated and normalized to obtain the contribution weight of each characteristic variable to extreme precipitation.
[0010] Preferably, step S2 further includes: By associating and pairing historical extreme event sample databases, the weighted cosine similarity between historical environmental field feature variables and real-time environmental field feature variables of the target area is calculated to obtain the environmental field similarity index of historical-real-time extreme precipitation events in the target area. The similarity indices are sorted in descending order, and the top N historical extreme precipitation events in the target area with the highest similarity indices are selected to establish a similarity set of extreme precipitation events under the current real-time environmental field conditions.
[0011] Preferably, step S3 specifically includes: Based on the extreme precipitation event similarity set and combined with the historical extreme event sample database, a unique ID is extracted for each historical case. Based on the event ID, complete observed precipitation data of the historical extreme event in the target area is extracted. For each similar historical extreme event, the daily precipitation sequence of all grid points in the target forecast area during the actual occurrence period of the event is extracted. The average precipitation intensity sequence of each event in the target area is calculated as the actual observed precipitation intensity sequence corresponding to the similar event.
[0012] Preferably, step S3 further includes: The peak values of the actual observed precipitation intensity sequence corresponding to the similar event are extracted as the sample point set. The environmental field similarity index of the historical-real-time extreme precipitation event in the target area corresponding to each sample point is transformed into positive weights as sample weights through an exponential function. A Gaussian kernel function is used as the smoothing kernel, and the optimal bandwidth is determined by the Silverman rule. The weighted kernel density estimation probability density function of the regional average precipitation intensity in the future forecast period of the target area is constructed to obtain the distribution of precipitation of different intensities in the future based on similar historical extreme events under the current environmental field conditions.
[0013] Preferably, step S3 further includes: Based on the weighted kernel density estimation probability density function of the regional average precipitation intensity during the future forecast period of the target area, the probability that the precipitation intensity of the target area exceeds a specific threshold R is calculated. Based on long-term historical climate data, a precipitation intensity-return period relationship table is established. The design precipitation extremes for different return periods F in the target area are found. Using the design precipitation extremes as thresholds, the probability that precipitation intensity will exceed the design precipitation extremes in the future target period under the current forecast situation and similar historical extreme events is calculated. A structured precipitation extreme probability forecast product is generated, including the creation of probability forecast tables, listing the exceedance probability of different precipitation levels, providing the maximum precipitation intensity estimate corresponding to different return periods, and generating extreme precipitation probability forecast schemes.
[0014] Compared with the prior art, the beneficial effects of the present invention are as follows: This invention proposes a precipitation extreme value forecasting scheme based on multi-source historical data. By fusing reanalysis data and gridded precipitation products to construct a high-quality historical sample library, and by dynamically calculating the similarity index based on the real-time environmental field, the physical basis and historical traceability of the forecasting system are significantly improved. By utilizing actual observation sequences of similar historical extreme events, a probability density function of future precipitation intensity is constructed through weighted kernel density estimation, overcoming the defect of traditional numerical models that systematically underestimate extreme precipitation systems, and more accurately quantifying the distribution tail probability of low-probability extreme events. By directly linking probability forecasts with return periods and design precipitation extreme values, structured and operationally oriented forecast products are generated, providing a scientific and practical quantitative risk basis for disaster prevention and mitigation decision-making. Attached Figure Description
[0015] Figure 1 This is a flowchart of a precipitation extreme value forecasting method based on multi-source historical data. Detailed Implementation
[0016] The following description is intended to disclose the invention and enable those skilled in the art to implement it. The preferred embodiments described below are merely examples, and other obvious variations will occur to those skilled in the art.
[0017] Reference Figure 1 As shown, a method for predicting extreme precipitation values based on multi-source historical data includes: S1. Establish a multi-source historical data fusion database of reanalysis data and gridded precipitation products, identify historical extreme precipitation events in the target area, extract the environmental field characteristics of historical extreme precipitation, and generate a sample library of historical extreme events. Step S1 specifically includes: Collect high-resolution satellite radar fusion precipitation products of the target area over the past ten years, obtain MERRA-2 reanalysis data of the target area over the past forty years, extract three-dimensional field variables of dynamics, thermodynamics, and water vapor associated with precipitation, align timestamps, perform data preprocessing, and establish a spatiotemporally aligned reanalysis data-gridized precipitation product multi-source historical data fusion database. Based on a multi-source historical data fusion database of reanalysis data and gridded precipitation products, high spatiotemporal resolution precipitation data from the past thirty years were extracted, including CMORPH and GPCP gridded precipitation products. ERA5 and NCEP / NCAR reanalysis data from the same period were also obtained. Specific humidity, vertical velocity, wind field, geopotential height, convective available potential energy, and water vapor flux divergence were calculated to generate a set of historical environmental field variables for the target area. Bilinear interpolation was performed on gridded data from different sources and resolutions to unify the spatial grid and temporal frequency.
[0018] Step S1 also includes: Based on the reanalysis data-gridized precipitation product multi-source historical data fusion database, the annual precipitation sequences of all grid points in the target area are statistically analyzed, and the 95th percentile value of the sequence is calculated as the extreme precipitation threshold of the grid point. The historical precipitation data of the grid points are scanned daily. If the daily precipitation of a certain grid point is greater than or equal to the extreme precipitation threshold of the grid point, the grid point is marked as an extreme precipitation grid point on that day. The total area of all extreme precipitation grid points in the target area is calculated. The number of days with adjacent extreme precipitation grid points in time is defined as a single candidate event. For each candidate event, the spatial distribution of its daily extreme precipitation grid points is examined. Using the 8-neighborhood connectivity method, multiple spatially independent rain clusters are aggregated from the daily extreme precipitation grid points. If two candidate events are less than or equal to 2 days apart in time and the dominant rain clusters have significant spatial overlap, then these two candidate events are merged into one extreme precipitation event. A minimum duration t is set, and historical extreme precipitation events in the target area are identified to obtain their unique ID, start and end dates, spatial range, continuous area with the largest cumulative rainfall, maximum daily precipitation of all grid points during the event period, and average total rainfall of the target area. The significant folding includes: the distance between the centroid of the dominant rain cluster and the threshold h.
[0019] Step S1 also includes: For identified historical extreme precipitation events in the target area, the 24 hours before the event is selected as the precursor period. Corresponding environmental field variables are extracted, the spatial average value of each variable in the target area is calculated, and a multi-dimensional feature vector of historical extreme precipitation environmental field is generated. Combined with the corresponding actual precipitation intensity and occurrence time, a sample library of historical extreme events is generated.
[0020] When using it, please refer to the steps outlined above: In existing technologies, the fusion of reanalysis data and multi-source historical data of gridded precipitation products often faces challenges such as heterogeneous data sources, mismatched spatiotemporal resolution, and highly subjective methods for identifying extreme precipitation events. This results in deficiencies in the construction of historical extreme precipitation event sample databases, including vague event definitions, incomplete extraction of environmental field features, and insufficient spatiotemporal consistency, limiting the accuracy of extreme precipitation mechanism analysis and forecast model training. This step establishes a spatiotemporally aligned multi-source fusion database and employs an objective identification method based on statistical thresholds and spatiotemporal connectivity criteria. This enables the systematic and standardized extraction of historical extreme precipitation events and their precursor environmental field features, significantly improving the spatiotemporal consistency and physical representativeness of event samples, and providing high-quality, multi-dimensional basic data support.
[0021] S2. Obtain real-time multidimensional environmental field data of the target area, establish a structured set of environmental field feature variables, calculate the contribution weight of each feature variable to extreme precipitation, associate and pair historical extreme event sample databases, calculate the environmental field similarity index of historical-real-time extreme precipitation events in the target area, and establish a similarity set of extreme precipitation events. Step S2 specifically includes: Based on the ECMWF weather forecast center, multidimensional environmental field grid data of the target area at the current time and the forecast field for the next 6-24 hours are obtained. The real-time environmental field variables are subjected to quality control and interpolation. The three-dimensional field variables of dynamics, thermodynamics and water vapor associated with precipitation are extracted. According to the precursor period, the spatial average value of each variable in the target area in the next 24 hours is calculated, and a set of structured characteristic variables of the environmental field is generated.
[0022] Step S2 also includes: Using environmental field characteristic variables from a historical extreme event sample database as input, the input data is sampled with replacement to generate multiple environmental field characteristic variable subsets. A decision tree is trained for each environmental field characteristic variable subset. When splitting nodes, some features are randomly selected to find the optimal split point, and an extreme precipitation prediction model is established. The hyperparameters of the number of decision trees and the maximum depth are optimized through grid search and cross-validation. The extreme precipitation observation label corresponding to each historical case is used as the output, with 1 indicating that extreme precipitation occurred and 0 indicating that extreme precipitation did not occur. The Gini importance score of each environmental field characteristic variable is calculated and normalized to obtain the contribution weight of each characteristic variable to extreme precipitation.
[0023] Step S2 also includes: By associating and pairing historical extreme event sample databases, the weighted cosine similarity between historical environmental field feature variables and real-time environmental field feature variables of the target area is calculated to obtain the environmental field similarity index of historical-real-time extreme precipitation events in the target area. The similarity indices are sorted in descending order, and the top N historical extreme precipitation events in the target area with the highest similarity indices are selected to establish a similarity set of extreme precipitation events under the current real-time environmental field conditions.
[0024] When using it, please refer to the steps outlined above: Extreme precipitation prediction often relies on single numerical model outputs or a few empirical factors, resulting in insufficient extraction and integration of multidimensional and structured features of the environmental field. It lacks a coordinated quantitative representation of dynamic, thermal, and water vapor conditions, and the weights of feature variables often depend on subjective experience or simple statistics, failing to objectively quantify the dynamic contribution of each environmental element to extreme precipitation based on machine learning. This leads to weak physical consistency and specificity in matching historical similar events. This step acquires multidimensional environmental field grid data in real time and constructs a structured feature variable set. It then uses a random forest-based machine learning method to objectively quantify the contribution weights of each feature variable to extreme precipitation. Finally, it combines weighted cosine similarity to accurately match a set of physically meaningful similar events from a historical sample database. This significantly improves the comprehensiveness, objectivity, and physical interpretability of the extreme precipitation environmental field representation, providing more targeted and reliable historical similarity evidence for subsequent forecasts.
[0025] S3. Based on the similarity set of extreme precipitation events, extract the actual observed precipitation intensity sequence of each historical case within the target forecast period, construct the probability density function of future precipitation intensity in the target area, quantify the probability of precipitation extremes occurring at different return periods, and generate a precipitation extreme probability forecast scheme. Step S3 specifically includes: Based on the extreme precipitation event similarity set and combined with the historical extreme event sample database, a unique ID is extracted for each historical case. Based on the event ID, complete observed precipitation data of the historical extreme event in the target area is extracted. For each similar historical extreme event, the daily precipitation sequence of all grid points in the target forecast area during the actual occurrence period of the event is extracted. The average precipitation intensity sequence of each event in the target area is calculated as the actual observed precipitation intensity sequence corresponding to the similar event.
[0026] Step S3 also includes: The peak values of the actual observed precipitation intensity sequence corresponding to the similar event are extracted as the sample point set. The environmental field similarity index of the historical-real-time extreme precipitation event in the target area corresponding to each sample point is transformed into positive weights as sample weights through an exponential function. A Gaussian kernel function is used as the smoothing kernel, and the optimal bandwidth is determined by the Silverman rule. The weighted kernel density estimation probability density function of the regional average precipitation intensity in the future forecast period of the target area is constructed to obtain the distribution of precipitation of different intensities in the future based on similar historical extreme events under the current environmental field conditions.
[0027] Step S3 also includes: Based on the weighted kernel density estimation probability density function of the regional average precipitation intensity during the future forecast period of the target area, the probability that the precipitation intensity of the target area exceeds a specific threshold R is calculated. Based on long-term historical climate data, a precipitation intensity-return period relationship table is established. The design precipitation extremes for different return periods F in the target area are found. Using the design precipitation extremes as thresholds, the probability that precipitation intensity will exceed the design precipitation extremes in the future target period under the current forecast situation and similar historical extreme events is calculated. A structured precipitation extreme probability forecast product is generated, including the creation of probability forecast tables, listing the exceedance probability of different precipitation levels, providing the maximum precipitation intensity estimate corresponding to different return periods, and generating extreme precipitation probability forecast schemes.
[0028] When using it, please refer to the steps outlined above: In the field of extreme precipitation probability forecasting, existing technologies mostly rely on single numerical models or simple statistical models. Their shortcomings include: difficulty in accurately characterizing extreme precipitation as a low-probability, high-impact nonlinear event, resulting in insufficient forecasting capability for low-probability extreme events; lack of systematic integration and weighted fusion of actual observation sequences of similar historical extreme events, making forecast results susceptible to model systematic biases and insufficient objectivity and stability of probability density estimation; and traditional methods typically only provide point estimates or simple probabilities, failing to effectively establish a dynamic probabilistic relationship between precipitation intensity and return period, thus limiting their direct application in disaster prevention and mitigation decision-making. Value: This step constructs a weighted kernel density estimation probability density function based on similarity sets. Its beneficial effect lies in effectively integrating environmental field similarity weights by utilizing actual observation sequences of historical extreme precipitation events, thereby improving the reliability and physical consistency of probability density estimation. By combining probability forecasts with return periods and design precipitation extremes, structured and operationally oriented probability forecast products are directly generated. This not only quantifies the probability of occurrence of precipitation of different intensities but also clarifies the significance of the return period of extreme precipitation events, significantly improving the practicality and decision support capabilities of forecast results, and providing more accurate and intuitive risk probability information for disaster prevention and mitigation.
[0029] Based on the above, the specific implementation method is as follows: Using the middle and lower reaches of the Yangtze River as the target area, high-resolution CMORPH satellite radar fusion precipitation products (0.1°×0.1°, hourly) from 2014 to 2023 and MERRA-2 reanalysis data (0.5°×0.5°, 6-hourly) from 1984 to 2023 were collected. The precipitation products and reanalysis data were time-stamp aligned and missing values were imputed. Bilinear interpolation was used to unify the reanalysis data to a 0.1° grid, and a spatiotemporally consistent multi-source fusion database was established. Based on daily precipitation data from the past 30 years (1994-2023), the 95th percentile of the precipitation sequence for each grid point was calculated as the extreme precipitation threshold, and extreme precipitation grid points were identified daily. Using temporal adjacency (interval ≤ 2 days) and spatial connectivity (8-neighbor clusters, centroid distance of dominant rain cluster ≤ 200 km) rules, historical extreme precipitation events in the region were objectively identified. A persistent extreme precipitation event from July 4-8, 2020 was identified and assigned the unique ID "YRD20200704", recording its start and end times, spatial extent, and the center of maximum cumulative rainfall. For this event, environmental field variables such as specific humidity, vertical velocity, and convective available potential energy were extracted from the ERA5 reanalysis data of the 24 hours before the event (00:00-24:00 on July 3). The regional spatial average was calculated to form a multidimensional feature vector (average specific humidity 14.2 g / kg, average vertical velocity -0.8 Pa / s), which was then correlated with the actual precipitation intensity of the event (regional average rainfall 45 mm / day) and stored in the historical extreme event sample database.
[0030] At 08:00 on July 10, 2024, the forecast field for the target area for the next 24 hours was obtained from ECMWF (the forecast time started at 00:00 on July 10). Grid data such as specific humidity, vertical velocity, wind field, and geopotential height were extracted and interpolated to a 0.1° grid after quality control. The spatial average value of each variable during the precursor period of the region (08:00 on July 10 to 08:00 on July 11) was calculated to construct a set of structured characteristic variables for the real-time environmental field. Using 500 historical extreme events and their corresponding environmental field features from a historical sample database, a random forest extreme precipitation prediction model (number of decision trees = 200, maximum depth = 10) was trained. The contribution weights of each feature variable were calculated using Gini importance (specific humidity weight 0.35, vertical velocity weight 0.28, water vapor flux divergence weight 0.18). Based on these weights, the weighted cosine similarity between the real-time environmental field features and the environmental field features of each historical event was calculated. After sorting the similarity in descending order, the top 20 most similar historical events (the one with the highest similarity is historical event "YRD20200704", similarity index 0.92) were selected to form the extreme precipitation event similarity set under the current real-time environmental field.
[0031] Based on 20 similar historical events, the daily average precipitation intensity sequence of each event during its actual occurrence period in the target area was extracted from the historical precipitation observation database. The peak value of each sequence was taken as the sample point (the peak intensity of the event “YRD20200704” was 45 mm / day). The environmental field similarity index corresponding to each sample point was converted into weights through an exponential function. The probability density function of the regional average precipitation intensity in the next 24 hours was constructed by weighted kernel density estimation (Gaussian kernel, bandwidth adaptively determined by Silverman rule). This function calculates the probability of precipitation intensity exceeding different thresholds: the probability of ≥50 mm / day is 18%, and the probability of ≥100 mm / day is 5%. Combining this with a precipitation intensity-return period table established from long-term historical climate data (24-hour precipitation extremes: 120 mm / day for a 50-year return period and 150 mm / day for a 100-year return period), the probability of current precipitation intensity exceeding the design return period is calculated: the probability of exceeding the 50-year return period is 3%, and the probability of exceeding the 100-year return period is 1%. Finally, a structured probability forecast table is generated, listing the exceedance probabilities of different precipitation levels (30, 50, 80, 100 mm / day) and the extreme value estimates corresponding to different return periods, forming a precipitation extreme value probability forecasting scheme that can directly serve flood control departments.
[0032] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the claimed invention. The scope of protection claimed by the appended claims and their equivalents is defined.
Claims
1. A method for precipitation extreme forecast based on multi-source historical data, characterized in that, include: S1. Establish a multi-source historical data fusion database of reanalysis data and gridded precipitation products, identify historical extreme precipitation events in the target area, extract the environmental field characteristics of historical extreme precipitation, and generate a sample library of historical extreme events. S2. Obtain real-time multidimensional environmental field data of the target area, establish a structured set of environmental field feature variables, calculate the contribution weight of each feature variable to extreme precipitation, associate and pair historical extreme event sample databases, calculate the environmental field similarity index of historical-real-time extreme precipitation events in the target area, and establish a similarity set of extreme precipitation events. S3. Based on the similarity set of extreme precipitation events, extract the actual observed precipitation intensity sequence of each historical case within the target forecast period, construct the probability density function of future precipitation intensity in the target area, quantify the probability of precipitation extremes occurring at different return periods, and generate a precipitation extreme probability forecasting scheme.
2. The method of claim 1, wherein the method is characterized by, Step S1 specifically includes: Collect high-resolution satellite radar fusion precipitation products of the target area over the past ten years, obtain MERRA-2 reanalysis data of the target area over the past forty years, extract three-dimensional field variables of dynamics, thermodynamics, and water vapor associated with precipitation, align timestamps, perform data preprocessing, and establish a spatiotemporally aligned reanalysis data-gridized precipitation product multi-source historical data fusion database. Based on a multi-source historical data fusion database of reanalysis data and gridded precipitation products, high spatiotemporal resolution precipitation data from the past thirty years were extracted, including CMORPH and GPCP gridded precipitation products. ERA5 and NCEP / NCAR reanalysis data from the same period were also obtained. Specific humidity, vertical velocity, wind field, geopotential height, convective available potential energy, and water vapor flux divergence were calculated to generate a set of historical environmental field variables for the target area. Bilinear interpolation was performed on gridded data from different sources and resolutions to unify the spatial grid and temporal frequency.
3. The method of claim 2, wherein the method is characterized by, Step S1 also includes: Based on the reanalysis data-gridized precipitation product multi-source historical data fusion database, the annual precipitation sequences of all grid points in the target area are statistically analyzed, and the 95th percentile value of the sequence is calculated as the extreme precipitation threshold of the grid point. The historical precipitation data of the grid points are scanned daily. If the daily precipitation of a certain grid point is greater than or equal to the extreme precipitation threshold of the grid point, the grid point is marked as an extreme precipitation grid point on that day. The total area of all extreme precipitation grid points in the target area is calculated. The number of days with adjacent extreme precipitation grid points in time is defined as a single candidate event. For each candidate event, the spatial distribution of its daily extreme precipitation grid points is examined. Using the 8-neighborhood connectivity method, multiple spatially independent rain clusters are aggregated from the daily extreme precipitation grid points. If two candidate events are less than or equal to 2 days apart in time and the dominant rain clusters have significant spatial overlap, then these two candidate events are merged into one extreme precipitation event. A minimum duration t is set, and historical extreme precipitation events in the target area are identified to obtain their unique ID, start and end dates, spatial range, continuous area with the largest cumulative rainfall, maximum daily precipitation of all grid points during the event period, and average total rainfall of the target area. The significant folding includes: the distance between the centroid of the dominant rain cluster and the threshold h.
4. The method for predicting extreme precipitation values based on multi-source historical data according to claim 3, characterized in that, Step S1 also includes: For identified historical extreme precipitation events in the target area, the 24 hours before the event is selected as the precursor period. Corresponding environmental field variables are extracted, the spatial average value of each variable in the target area is calculated, and a multi-dimensional feature vector of historical extreme precipitation environmental field is generated. Combined with the corresponding actual precipitation intensity and occurrence time, a sample library of historical extreme events is generated.
5. The method for predicting extreme precipitation values based on multi-source historical data according to claim 1, characterized in that, Step S2 specifically includes: Based on the ECMWF weather forecast center, multidimensional environmental field grid data of the target area at the current time and the forecast field for the next 6-24 hours are obtained. The real-time environmental field variables are subjected to quality control and interpolation. The three-dimensional field variables of dynamics, thermodynamics and water vapor associated with precipitation are extracted. According to the precursor period, the spatial average value of each variable in the target area in the next 24 hours is calculated, and a set of structured characteristic variables of the environmental field is generated.
6. The method for predicting extreme precipitation values based on multi-source historical data according to claim 5, characterized in that, Step S2 also includes: Using environmental field characteristic variables from a historical extreme event sample database as input, the input data is sampled with replacement to generate multiple environmental field characteristic variable subsets. A decision tree is trained for each environmental field characteristic variable subset. When splitting nodes, some features are randomly selected to find the optimal split point, and an extreme precipitation prediction model is established. The hyperparameters of the number of decision trees and the maximum depth are optimized through grid search and cross-validation. The extreme precipitation observation label corresponding to each historical case is used as the output, with 1 indicating that extreme precipitation occurred and 0 indicating that extreme precipitation did not occur. The Gini importance score of each environmental field characteristic variable is calculated and normalized to obtain the contribution weight of each characteristic variable to extreme precipitation.
7. The method for predicting extreme precipitation values based on multi-source historical data according to claim 6, characterized in that, Step S2 also includes: By associating and pairing historical extreme event sample databases, the weighted cosine similarity between historical environmental field feature variables and real-time environmental field feature variables of the target area is calculated to obtain the environmental field similarity index of historical-real-time extreme precipitation events in the target area. The similarity indices are sorted in descending order, and the top N historical extreme precipitation events in the target area with the highest similarity indices are selected to establish a similarity set of extreme precipitation events under the current real-time environmental field conditions.
8. The method for predicting extreme precipitation values based on multi-source historical data according to claim 7, characterized in that, Step S3 specifically includes: Based on the extreme precipitation event similarity set and combined with the historical extreme event sample database, a unique ID is extracted for each historical case. Based on the event ID, complete observed precipitation data of the historical extreme event in the target area is extracted. For each similar historical extreme event, the daily precipitation sequence of all grid points in the target forecast area during the actual occurrence period of the event is extracted. The average precipitation intensity sequence of each event in the target area is calculated as the actual observed precipitation intensity sequence corresponding to the similar event.
9. A method for predicting extreme precipitation values based on multi-source historical data according to claim 8, characterized in that, Step S3 also includes: The peak values of the actual observed precipitation intensity sequence corresponding to the similar event are extracted as the sample point set. The environmental field similarity index of the historical-real-time extreme precipitation event in the target area corresponding to each sample point is transformed into positive weights as sample weights through an exponential function. A Gaussian kernel function is used as the smoothing kernel, and the optimal bandwidth is determined by the Silverman rule. The weighted kernel density estimation probability density function of the regional average precipitation intensity in the future forecast period of the target area is constructed to obtain the distribution of precipitation of different intensities in the future based on similar historical extreme events under the current environmental field conditions.
10. A method for predicting extreme precipitation values based on multi-source historical data according to claim 9, characterized in that, Step S3 also includes: Based on the weighted kernel density estimation probability density function of the regional average precipitation intensity during the future forecast period of the target area, the probability that the precipitation intensity of the target area exceeds a specific threshold R is calculated. Based on long-term historical climate data, a precipitation intensity-return period relationship table is established. The design precipitation extremes for different return periods F in the target area are found. Using the design precipitation extremes as thresholds, the probability that precipitation intensity will exceed the design precipitation extremes in the future target period under the current forecast situation and similar historical extreme events is calculated. A structured precipitation extreme probability forecast product is generated, including the creation of probability forecast tables, listing the exceedance probability of different precipitation levels, providing the maximum precipitation intensity estimate corresponding to different return periods, and generating extreme precipitation probability forecast schemes.