Machine learning based radar precipitation estimation adaptive calibration method and system
By employing a machine learning-based adaptive calibration method, utilizing dual-polarization radar network data and ground rain gauge information, event-scale modeling and adaptive error correction for rainfall intensity are constructed. This solves the problem of overestimating weak precipitation and underestimating heavy precipitation in radar precipitation estimation, achieving high-precision and interpretable precipitation estimation.
Patent Information
- Application Number
- CN202610785819.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-02
- Publication Date
- 2026-08-25
AI Technical Summary
Existing radar precipitation estimation methods suffer from poor interpretability, weak generalization ability, incomplete error correction, and insufficient data utilization. In particular, there are structural errors such as overestimation of weak precipitation and underestimation of strong precipitation under different rainfall intensity levels.
An adaptive calibration method based on machine learning was adopted. By acquiring CAPPI three-dimensional reflectivity products from a dual-polarization weather radar network and ground rain gauge data, an event-scale modeling dataset was constructed. Three-dimensional radar reflectivity structural features were extracted, multiple machine learning baseline models were trained, and a physical constraint symbolic regression method was used to mine explicit precipitation estimation formulas. Adaptive error correction was performed for rainfall intensity.
It improves the consistency and stability of precipitation estimation results at different rainfall intensity levels, meets the requirements of meteorological operations for physical consistency, enhances the estimation accuracy of weak and heavy precipitation, and solves the limitation of traditional methods that cannot adapt to different precipitation intensity characteristics.
Smart Images

Figure CN122632200A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of radar meteorological detection technology, specifically to an adaptive calibration method and system for radar precipitation estimation based on machine learning. Background Technology
[0002] Quantitative precipitation estimation (QPE) using radar is one of the core technologies for weather forecasting, hydrological monitoring, and disaster early warning. Currently, the mainstream radar precipitation estimation methods are mainly divided into two categories: one is the traditional method based on empirical ZR relationships, and the other is the black-box model method based on machine learning.
[0003] Traditional ZR relationship methods use globally unified empirical formulas, which cannot adapt to precipitation characteristics of different precipitation types, regions, and seasons. They generally suffer from structural errors such as overestimation of weak precipitation and underestimation of strong precipitation, and are difficult to utilize the structural information of three-dimensional radar reflectivity.
[0004] While existing machine learning methods have achieved high accuracy on local datasets, they suffer from the following key drawbacks:
[0005] 1) Poor interpretability: Black box models cannot provide clear physical relationships, making it difficult to meet the requirements of meteorological operations for physical consistency, and also hindering error tracing and model optimization;
[0006] 2) Weak generalization ability: Most methods use random sample partitioning for training and validation, which leads to a significant decrease in model performance when applied across events and regions;
[0007] 3) Incomplete error correction: Existing error correction methods are mostly global corrections, which cannot specifically address systematic deviations at different rainfall intensity levels;
[0008] 4) Insufficient data utilization: Most methods only use reflectivity data from a single height layer, failing to fully explore the vertical structural features of three-dimensional radar reflectivity.
[0009] With the advancement of meteorological modernization in my country, the nationwide dual-polarization weather radar network has been basically completed and put into operational use. National weather radar mosaics and CAPPI three-dimensional reflectivity products have been gradually applied to operational monitoring. CAPPI three-dimensional reflectivity products can provide radar reflectivity structure information at multiple altitudes, providing a data foundation for the vertical structure analysis of precipitation clouds and quantitative precipitation estimation. However, existing radar precipitation estimation methods still largely rely on single altitude layers or empirical relationships, failing to fully utilize the reflectivity structure characteristics at multiple altitudes, and still exhibiting structural errors such as overestimation of weak precipitation and underestimation of strong precipitation at different rainfall intensities. Summary of the Invention
[0010] To address the shortcomings of existing technologies, this invention discloses an adaptive calibration method and system for radar precipitation estimation based on machine learning, in order to solve the problems mentioned in the background.
[0011] To achieve the above objectives, the present invention provides the following technical solution: an adaptive calibration method for radar precipitation estimation based on machine learning, comprising:
[0012] S1. Obtain the CAPPI three-dimensional reflectivity product of the dual-polarization weather radar network and the hourly cumulative precipitation data of the corresponding area from the ground rain gauge station within the preset time period;
[0013] S2. Analyze and perform quality control on the radar CAPPI three-dimensional reflectivity data to obtain effective radar reflectivity data;
[0014] S3. Establish the spatial matching relationship between ground rain gauges and radar grids, and extract multi-height radar reflectivity data for each rain gauge location;
[0015] S4. Based on effective radar reflectivity data and hourly precipitation data from ground rain gauges, identify typical precipitation events and construct an event-scale modeling dataset.
[0016] S5. Extract three-dimensional radar reflectivity structural features from the event-scale modeling dataset. The three-dimensional radar reflectivity structural features include low-level multi-height neighborhood average echo features and mid-level average echo features. Using the three-dimensional radar reflectivity structural features as input and the hourly cumulative precipitation at ground level as output, train multiple machine learning baseline models. Select the optimal baseline model through leave-one-out cross-validation and generate error distribution features for each rainfall intensity level based on the optimal baseline model.
[0017] S6. Based on the event-scale modeling dataset, the physical constraint symbolic regression method is used to mine interpretable explicit precipitation estimation formulas that satisfy monotonically increasing physical constraints. The higher the radar reflectivity, the greater the precipitation.
[0018] S7. Based on the joint error distribution characteristics of the optimal baseline model and the initial explicit precipitation estimation formula, the explicit precipitation estimation formula is subjected to adaptive error correction based on rainfall intensity to obtain the final quantitative precipitation estimation model.
[0019] S8. Input the radar CAPPI three-dimensional reflectivity data to be estimated into the final quantitative precipitation estimation model to generate the quantitative precipitation estimation results for the corresponding area.
[0020] Preferably, the specific steps for parsing and quality control of the radar CAPPI three-dimensional reflectivity data in step S2 include:
[0021] S21: Parse the 256-byte fixed-length header of the radar CAPPIbin file and extract the header time, layer number, and height information;
[0022] S22: Remove radar files whose header time is outside the preset time period, whose header year is abnormal, or whose reading errors occur;
[0023] S23: Decompress the bz2 compressed data block, convert the original int16 little-endian data to the logarithmic reflectance factor dBZ, using the following formula:
[0024] ,
[0025] Where r is the original integer value after decompression, and s is the preset scaling factor, which ranges from 8 to 12;
[0026] S24: Convert the logarithmic reflectance factor dBZ to the linear reflectance factor Z using the following formula:
[0027] .
[0028] S25: Aggregate and statistically analyze the linear reflectivity factor Z within the preset time window to obtain radar reflectivity characteristic data on an hourly scale. The aggregation method includes the average value, the maximum value, or the preset quantile.
[0029] S26: Based on the preset reflectivity threshold and effective scanning conditions, perform quality control on weak echoes and system noise, eliminate invalid echoes below the preset noise threshold, and retain weak precipitation echo signals that meet the conditions of temporal and spatial continuity.
[0030] Preferably, the specific steps for establishing the spatial matching relationship between ground rain gauges and radar grids in step S3 include:
[0031] S31: Obtain the latitude, longitude, and altitude information of each ground rain gauge station;
[0032] S32: Based on the Lambert conformal projection parameters used by the nationwide network radar, the latitude and longitude coordinates are converted into the row and column numbers of the radar grid, and the Euclidean distance from each rain gauge station to the center of the corresponding radar grid is calculated.
[0033] S33: Select ground rain gauge stations that are within a preset distance range from the center of the radar grid as valid rain gauge stations;
[0034] S34: Based on the row and column numbers of the effective rain gauges, extract the log reflectance factor data of the corresponding height layers of 500m, 1000m, 1500m, 2000m, 3000m, and 4000m.
[0035] Preferably, the specific steps for constructing the event-scale modeling dataset in step S4 include:
[0036] S41: Based on hourly precipitation data from ground rain gauges, identify periods of continuous precipitation. When there is effective precipitation (hourly precipitation ≥ 0.1 mm) for 3 consecutive hours or more and at least one station has an hourly precipitation greater than 5 mm, it is classified as a typical precipitation event.
[0037] S42: For each typical precipitation event, extract hourly precipitation data from all valid rain gauge stations within the event and multi-height radar reflectivity data for the corresponding time period;
[0038] S43: Calculate the cumulative precipitation at each rain gauge station within the event and filter out samples with cumulative precipitation greater than a preset threshold;
[0039] S44: Perform stratified balanced sampling on the selected samples according to precipitation events and rainfall intensity levels to construct an event-scale modeling dataset. The balanced sampling is used to avoid model bias caused by the dominance of weak precipitation samples or uneven distribution of samples of different rainfall intensity levels.
[0040] Preferably, the specific steps in step S5 for training the machine learning baseline model and generating error distribution features include:
[0041] S51: Extract 3D radar reflectivity structural features from the event-scale modeling dataset, where:
[0042] Low-level multi-height neighborhood average echo characteristics The arithmetic mean of the logarithmic reflectance factor of the 3×3 neighborhood within the 500m, 1000m, and 1500m height layers;
[0043] Mid-level average echo characteristics The arithmetic mean of the logarithmic reflectance factor for the 2000m, 3000m, and 4000m altitude layers, i.e. ,in , , The average log reflectance factors for the 2000m, 3000m, and 4000m altitude layers are respectively.
[0044] Vertical structural features The difference in log reflectivity factor between the 4000m and 2000m height layers;
[0045] S52: Using the three-dimensional radar reflectivity structural features as input and the hourly cumulative precipitation on the ground as output, train various types of machine learning models, including but not limited to random forest models, linear regression models, and log-linear models.
[0046] S53: The leave-one-out cross-validation method is used to evaluate the performance of each machine learning model. That is, one precipitation event is kept as the test set each time, and the remaining precipitation events are used as the training set. After repeating this process multiple times, the average performance index is calculated. The evaluation indexes include root mean square error, mean absolute error, correlation coefficient and Nash efficiency coefficient.
[0047] S54: Select the machine learning model with the best overall performance as the optimal baseline model based on the evaluation results;
[0048] S55: Input the event-scale modeling dataset into the optimal baseline model, generate the predicted value for each sample, and calculate the average deviation, relative deviation, and error distribution curve for each rainfall intensity level to obtain the error distribution characteristics for each rainfall intensity level.
[0049] Preferably, the specific steps in step S6 of mining explicit precipitation estimation formulas using the physical constraint symbolic regression method include:
[0050] S61: Perform a logarithmic transformation on the target variable, cumulative precipitation at ground stations R. The transformation formula is as follows:
[0051] ,
[0052] in, This is a preset minimum value, ranging from 0.001 to 0.01, used to avoid meaningless logarithmic operations;
[0053] S62: Using the three-dimensional radar reflectivity structural features as input variables and the transformed target variable y as output, a physical constraint symbolic regression method is used for searching, that is, the genetic programming symbolic regression method of the PySR library is used to generate multiple candidate explicit formulas. The physical constraints include formulas that monotonically increase with the input features.
[0054] S63: Perform performance evaluation and complexity analysis on multiple candidate explicit formulas, select the formula that achieves the best balance between accuracy, complexity and physical interpretability as the initial explicit precipitation estimation formula, calculate the comprehensive score with an accuracy weight of 0.7 and a complexity weight of 0.3, and select the formula with the highest comprehensive score as the initial explicit precipitation estimation formula.
[0055] S64: The initial explicit precipitation estimation formula can be a linear explicit formula, in the form of:
[0056] ,
[0057] Where 'a' is the first preset coefficient, with a value range of 0.03-0.04; and 'b' is the second preset coefficient, with a value range of 0.6-0.7.
[0058] S65: Convert the initial explicit precipitation estimation formula into a precipitation estimation formula:
[0059] .
[0060] Preferably, the specific steps of the adaptive error correction for rainfall intensity in step S7 include:
[0061] S71: Based on the error distribution characteristics generated by the optimal baseline model, the hourly cumulative precipitation is divided into four precipitation levels: light precipitation (5-10 mm), moderate to light precipitation (10-20 mm), moderate to heavy precipitation (20-50 mm), and heavy precipitation (above 50 mm), denoted as follows: (Weak precipitation) (Moderate to light precipitation) (Moderate to heavy precipitation) (Heavy rainfall);
[0062] S72: Calculate the initial explicit precipitation estimation formula for each precipitation level. Average deviation The calculation formula is:
[0063] ,
[0064] in, Precipitation level The number of samples; Let be the predicted precipitation for the i-th sample; Let be the observed precipitation for the i-th sample;
[0065] S73: For each precipitation level A piecewise linear deviation correction model is constructed. The correction coefficients are obtained by least-squares fitting of the joint error of the optimal baseline model and the initial explicit formula at this level. The joint error is the arithmetic mean of the prediction error of the optimal baseline model and the prediction error of the initial explicit formula. The corrected target variable... for:
[0066] ,
[0067] in, Precipitation level The slope correction factor has a value range of 0.8-1.2; Precipitation level The intercept correction factor has a value range of -0.5 to 0.5.
[0068] S74: For light precipitation levels and Increase the intercept correction factor The negative value range is used to offset the overestimation of weak precipitation; for heavy precipitation levels Increase the slope correction factor The positive range is intended to mitigate the underestimation of heavy rainfall;
[0069] S75: The piecewise linear deviation correction model is applied to the initial explicit precipitation estimation formula to obtain the corrected explicit precipitation estimation formula, which serves as the final quantitative precipitation estimation model.
[0070] ,
[0071] in, Precipitation level The revised precipitation estimate is as follows.
[0072] Preferably, the specific steps for generating the quantitative precipitation estimation result in step S8 include:
[0073] S81: Obtain the radar CAPPI three-dimensional reflectivity data to be estimated, and perform analysis and quality control according to the method in step S2;
[0074] S82: Extract the three-dimensional radar reflectivity structural features of the region to be estimated;
[0075] S83: Based on the three-dimensional radar reflectivity structure characteristics, a preliminary judgment of the precipitation level is made, and a correction coefficient corresponding to the precipitation level is selected;
[0076] S84: Input the three-dimensional radar reflectivity structure characteristics and corresponding correction coefficients into the final quantitative precipitation estimation model to obtain the estimated hourly cumulative precipitation value of the area to be estimated.
[0077] S85: The quantitative precipitation estimation results are validated by rainfall intensity using an independent validation set. The validation indicators include the root mean square error, mean absolute error, bias and correlation coefficient at each rainfall intensity level.
[0078] This invention also provides a machine learning-based adaptive calibration system for radar precipitation estimation, comprising:
[0079] The data acquisition module is used to acquire the CAPPI three-dimensional reflectivity product of the dual-polarization weather radar network and the hourly precipitation data of the corresponding ground rain gauge station within a preset time period.
[0080] The radar analysis and quality control module is used to analyze and control the quality of the CAPPI three-dimensional reflectivity data of the dual-polarization radar network to obtain effective radar reflectivity data.
[0081] The site matching module is used to establish the spatial matching relationship between ground rain gauges and radar grids, and extract multi-height radar reflectivity data for each rain gauge location.
[0082] The event sample construction module is used to identify typical precipitation events and construct an event-scale modeling dataset based on the effective radar reflectivity data and hourly precipitation data from ground rain gauges.
[0083] The machine learning error diagnosis module is used to extract three-dimensional radar reflectivity structural features from the event-scale modeling dataset, train multiple machine learning baseline models, select the optimal baseline model through leave-one-out cross-validation, and generate error distribution features for each rainfall intensity level.
[0084] The symbolic regression module is used to model the dataset based on the event scale and use the physical constraint symbolic regression method to mine interpretable explicit precipitation estimation formulas that satisfy monotonically increasing physical constraints.
[0085] The rainfall intensity adaptive correction module is used to perform rainfall intensity adaptive error correction on the explicit precipitation estimation formula according to the error distribution characteristics of each rainfall intensity level, so as to obtain the final quantitative precipitation estimation model.
[0086] It also includes a result generation module, which is used to input the radar CAPPI three-dimensional reflectivity data to be estimated into the final quantitative precipitation estimation model to generate quantitative precipitation estimation results for the corresponding region based on rainfall intensity.
[0087] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0088] 1. This invention uses a machine learning baseline model to accurately diagnose the error distribution characteristics of different rainfall intensity levels, and constructs a segmented linear deviation correction model for different rainfall intensity levels. This model specifically offsets the overestimation of weak precipitation and alleviates the underestimation of strong precipitation, overcoming the limitation that traditional global unified formulas cannot adapt to different precipitation intensity characteristics, and improving the consistency and stability of precipitation estimation results for different rainfall intensity levels.
[0089] 2. This invention innovatively integrates the high-precision advantages of machine learning with the interpretability advantages of symbolic regression. It uses the physical constraint symbolic regression method to mine explicit estimation formulas that conform to the physical laws of precipitation, avoiding the defect of pure machine learning black box models that cannot provide clear physical relationships. It not only meets the requirements of meteorological operations for physical consistency, but also maintains estimation performance close to that of black box models.
[0090] 3. This invention adopts an evaluation framework of event-scale modeling and leave-one-out cross-validation, avoiding data leakage and overfitting problems caused by random sample partitioning, and enhancing the performance stability of the model in different regions and different types of precipitation events. This invention makes full use of the multi-height layer structure information of the CAPPI three-dimensional reflectivity products of the weather radar network, and improves the utilization of radar reflectivity information by extracting multi-height layer neighborhood features and vertical structure features, further improving the estimation accuracy of weak precipitation and heavy precipitation. Attached Figure Description
[0091] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used together with the embodiments of the invention to explain the invention and do not constitute a limitation thereof.
[0092] In the attached diagram:
[0093] Figure 1 This is an overall flowchart of the adaptive calibration method for radar precipitation estimation based on machine learning in this invention;
[0094] Figure 2 This is a cross-validation of the Top 15 precipitation events with one event left out: a comparison chart of the average deviation of rainfall intensity levels;
[0095] Figure 3 This is a block diagram of the adaptive calibration system for radar precipitation estimation based on machine learning, as described in this invention.
[0096] Figure 4 This is a comparison chart of the average deviation of rainfall intensity using different methods. Detailed Implementation
[0097] The preferred embodiments of the present invention will be described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are for illustration and explanation only and are not intended to limit the present invention.
[0098] Example 1: As Figure 1 As shown, this embodiment provides an adaptive calibration method for radar precipitation estimation based on machine learning. It is implemented using the CAPPI three-dimensional reflectivity product from the national network of dual-polarization weather radars from March 1, 2025 to November 30, 2025, and hourly precipitation data from 62,720 automatic rain gauge stations nationwide. The specific steps are as follows:
[0099] S1: Data Acquisition
[0100] The system acquires CAPPI three-dimensional reflectance products from a nationwide network of dual-polarization weather radars and hourly cumulative precipitation data from corresponding regional ground rain gauges within a preset time period via a dedicated meteorological data line. The radar data is in binary bin format with a time resolution of 5 minutes. Each scan time contains three-dimensional reflectance data from multiple altitude layers. In this embodiment, six altitude layers (500m, 1000m, 1500m, 2000m, 3000m, and 4000m) are selected for feature construction, covering a 4200×6200 grid across the entire country, with a spatial resolution of km×1km. The ground rain gauge data includes station number, longitude, latitude, altitude, and hourly cumulative precipitation, with a time resolution of 1 hour.
[0101] S2: Radar data analysis and quality control, the specific steps of which include:
[0102] S21: Traverse all radar CAPPI bin files one by one, parse the fixed-length header of each file (first 256 bytes), and extract key metadata information such as time, layer number, height, and scan angle from the header.
[0103] S22: Perform time quality control. In one specific embodiment, files with abnormal header year and files with reading errors are removed. Files with abnormal header year include files with the year 2026, and files with reading errors are files that failed to be parsed or decompressed.
[0104] S23: For radar files that have passed time quality control, the bz2 algorithm is used to decompress their suffix data blocks. The decompressed int16 little-endian format raw data is then converted into a logarithmic reflectivity factor (dBZ). The conversion formula is as follows:
[0105] ,
[0106] Where r is the original integer value after decompression, and the preset scaling factor s is 10; S24: In radar meteorology, the standard conversion relationship between the logarithmic reflectivity factor (dBZ) and the linear reflectivity factor (Z) is defined as: dBZ = 10 log 10 (Z), through algebraic transformation, the logarithmic reflectivity factor dBZ is converted to the linear reflectivity factor Z. The conversion formula is:
[0107] ,
[0108] Since dBZ is a logarithmic unit, directly performing an arithmetic average will lead to distortion of the physical meaning. Therefore, it must be converted into a linear reflectivity factor before time aggregation.
[0109] S25: Using natural hours as the time window, the linear reflectivity factor Z of all effective scanning times within each hour is arithmetically averaged to obtain multi-height radar reflectivity characteristic data at the hour scale.
[0110] S26: Based on the high sensitivity of dual-polarization radar, adaptive enhancement processing is performed on weak reflectivity signals below 0dBZ to retain effective weak precipitation signals while eliminating system noise below -10dBZ.
[0111] S3: Establish the spatial matching relationship between ground rain gauges and radar grids, and extract multi-height radar reflectivity data for each rain gauge location;
[0112] S31: Read the station_info.csv file to obtain the longitude, latitude, and altitude information of 62,720 ground rain gauge stations;
[0113] S32: Based on the Lambert conformal projection parameters adopted by the national network radar (central meridian 105°E, standard parallels 25°N and 47°N, origin latitude 30°N, false east 0m, false north 0m), these parameters are a common standard in the meteorological industry, ensuring the spatial consistency of radar data in different regions. The latitude and longitude coordinates of each rain gauge station are converted into the row and column numbers of the radar grid, and the Euclidean distance from the rain gauge station to the center of the corresponding radar grid is calculated.
[0114] S33: Ground rain gauge stations less than 0.5km from the center of the radar grid were selected as effective rain gauge stations, resulting in a total of 62,687 effective rain gauge stations with an average distance of 0.4km from the grid center;
[0115] S34: Based on the row and column numbers of the effective rain gauges, extract the hourly average log reflectance factor data for the six altitude layers at corresponding locations: 500m, 1000m, 1500m, 2000m, 3000m, and 4000m.
[0116] S4: Based on effective radar reflectivity data and hourly precipitation data from ground rain gauges, identify typical precipitation events and construct an event-scale modeling dataset. Specific steps include:
[0117] S41: Based on hourly precipitation data from ground rain gauges, identify periods of continuous precipitation. When there is effective precipitation for several consecutive hours and at least one station has an hourly precipitation greater than 5 mm, it is classified as a typical precipitation event.
[0118] S42: Select 15 representative heavy precipitation events, numbered 3, 5, 11, 21, 23, 25, 36, 40, 43, 57, 100, 101, 105, 124, and 128, and extract hourly precipitation data from all valid rain gauge stations within each event and multi-height radar reflectivity data for the corresponding time period;
[0119] S43: Calculate the cumulative precipitation for each rain gauge station within the corresponding event. After quality control, a total of 784,252 station-event samples were obtained. Further screening was conducted to select samples with cumulative precipitation greater than 5 mm as modeling samples, resulting in approximately 237,074 modeling samples.
[0120] S44: Perform stratified balanced sampling on the selected samples according to precipitation events and rainfall intensity levels, and extract the same number of samples for each event and each rainfall intensity level to construct a balanced event-scale modeling dataset for symbolic regression pre-experiment; in a specific embodiment, the balanced sample size is 48,000 to avoid model bias caused by an excessively high proportion of heavy precipitation samples.
[0121] S5. Machine learning error diagnosis, the specific steps of which include:
[0122] S51: Extract 3D radar reflectivity structural features from the event-scale modeling dataset, including:
[0123] Low-level multi-height neighborhood average echo characteristics Calculate the arithmetic mean of the logarithmic reflectance factor of the 3×3 neighborhood at heights of 500m, 1000m, and 1500m for each station location, and then average the results across the three height levels; for each ground rain gauge corresponding to the radar grid point Its low-level multi-height neighborhood average echo characteristics The calculation steps are as follows:
[0124] (1) Extract data from the three altitude layers of 500m, 1000m, and 1500m respectively, using... The log reflectance factor values for 9 grid points within a 3×3 neighborhood centered on the center;
[0125] (2) Calculate the arithmetic mean of the nine values for each height level, and obtain , , ;
[0126] (3) Calculate the arithmetic mean again for the average values of the three altitude levels, i.e.:
[0127] ,
[0128] This feature can effectively reflect the horizontal distribution characteristics of the lower layer of precipitation clouds, avoiding the influence of random fluctuations in reflectance of single grid points.
[0129] Mid-level average echo characteristics Calculate the arithmetic mean of the logarithmic reflectance factors at altitudes of 2000m, 3000m, and 4000m for each site, i.e. ,in, , , The average log reflectance factors are for the 2000m, 3000m, and 4000m height layers, respectively.
[0130] Vertical structural features Calculate the difference in log reflectivity factor between the 4000m and 2000m altitude layers to reflect the vertical development intensity of precipitation clouds.
[0131] S52: Using the three-dimensional radar reflectivity structure features as input and the hourly cumulative precipitation on the ground as output, train four different types of machine learning models respectively: RF1_lowlevel_mean, RF2_multiheight_plus, M4_multiheight_plus_ols, and M1_lowlevel_mean_loglinear.
[0132] S53: Use leave-one-out cross-validation to evaluate the performance of each machine learning model. That is, each time one event is reserved as the test set and the remaining 14 events are used as the training set. After repeating 15 times, the average performance index is calculated. The evaluation index includes root mean square error (RMSE), mean absolute error (MAE), correlation coefficient (CC), and Nash efficiency coefficient (NSE).
[0133] S54: Based on the comprehensive evaluation results, the M4_multiheight_plus_ols linear model is selected as the optimal baseline model; it performs best in terms of MAE, absolute deviation and correlation coefficient.
[0134] S55: Input the event-scale modeling dataset into the optimal baseline model to generate predicted values for each sample, and statistically analyze the average deviation, relative deviation, and error distribution curves according to rainfall intensity level to obtain the error distribution characteristics for each rainfall intensity level, such as... Figure 2 As shown, this is a comparison of the average deviation of different models in terms of rainfall intensity under cross-validation with one leave event for the Top 15 precipitation events. The vertical axis represents the average deviation (unit: mm), and the horizontal axis represents the rainfall intensity level (unit: mm). The results show that all baseline models have a consistent error trend: the 5-10 mm rainfall intensity range is overestimated by 12-15 mm, the 10-20 mm rainfall intensity range is overestimated by 9-11 mm, the 20-50 mm rainfall intensity range has a smaller deviation (-2.5 to -0.5 mm), and the rainfall intensity range above 50 mm is underestimated by 38-45 mm.
[0135] S6. Based on the event-scale modeling dataset, the physical constraint symbolic regression method is used to mine an interpretable explicit precipitation estimation formula that satisfies the monotonically increasing physical constraints. The specific steps include:
[0136] S61: Perform a logarithmic transformation on the target variable, hourly cumulative precipitation R. The transformation formula is as follows:
[0137] ,
[0138] Among them, the preset minimum value The value is 0.01, used to avoid Logarithmic operations are meaningless at this time;
[0139] S62: Based on the average echo characteristics of the middle layer Using the transformed target variable y as the input variable, this embodiment employs a genetic programming symbolic regression method based on the PySR library for the search. The search process is configured as follows:
[0140] Population size: 100;
[0141] The basic set of operators: addition, subtraction, multiplication, division, and constants;
[0142] Maximum formula complexity: 15 (complexity is defined as the number of operators + the number of variables);
[0143] Loss function: Mean Squared Error (MSE) + Non-monotonic penalty term;
[0144] Physical constraints: Add a non-monotonic penalty term to the loss function to ensure that the formula follows... Monotonically increasing; for input features any two values If the formula output Then a penalty value is added to the loss function. Ensure the formula follows Monotonically increasing;
[0145] Number of iterations: 100 generations; a total of 7 candidate explicit formulas were generated.
[0146] S63: Performance evaluation and complexity analysis were performed on the seven candidate explicit formulas, and the overall score was calculated with a precision weight of 0.7 and a complexity weight of 0.3. The results show that formula C5 has the highest overall score, achieving the optimal balance between precision, complexity, and physical interpretability.
[0147] S64: The initial explicit precipitation estimation formula is a linear formula, in the form of:
[0148] ;
[0149] This embodiment uses a linear formula as the optimal implementation method. Those skilled in the art can choose other explicit formulas that satisfy the monotonically increasing constraint according to actual needs, and all of them fall within the protection scope of this invention.
[0150] S65: Convert the initial explicit precipitation estimation formula into a precipitation estimation formula:
[0151] .
[0152] S7: Adaptive error correction for rainfall intensity, the specific steps include:
[0153] S71: Based on the error distribution characteristics generated by the optimal baseline model, the hourly cumulative precipitation at ground level is divided into four precipitation levels: (5-10mm, light precipitation) (10-20mm, moderate to light precipitation) (20-50mm, moderate to heavy rainfall) (Heavy rainfall exceeding 50mm);
[0154] S72: Calculate the initial explicit precipitation estimation formula for each precipitation level. Average deviation The calculation formula is:
[0155] ,
[0156] in, Precipitation level The number of samples; Let be the predicted precipitation for the i-th sample; Let be the observed precipitation for the i-th sample;
[0157] Calculated mm, mm, mm, mm;
[0158] S73: For each precipitation level A piecewise linear bias correction model is constructed. The correction coefficients are obtained by least-squares fitting of the joint error (arithmetic mean of the prediction errors of the optimal baseline model and the initial explicit formula at this level). The corrected target variable... for:
[0159] ,
[0160] in, Precipitation level The slope correction factor has a value range of 0.8-1.2; Precipitation level The intercept correction factor has a value range of -0.5 to 0.5.
[0161] Rainfall Intensity Correction Coefficient , The fitting objective function is:
[0162] ,
[0163] in, The predicted value is the initial explicit formula. The predicted value is the one from the optimal baseline model. For the observed values, This is the logarithmic output value of the initial explicit formula. The objective function simultaneously fits the error characteristics of both the explicit formula and the baseline model, ensuring that the corrected model combines the advantages of both.
[0164] S74: Adjust the correction coefficients according to the error characteristics of each rainfall intensity level:
[0165] For weak precipitation level and Increase the intercept correction factor The negative value magnitude is used to offset the overestimation;
[0166] In response to the heavy rainfall level Increase the slope correction factor The positive magnitude is intended to mitigate underestimation;
[0167] : , ;
[0168] : , ;
[0169] : , ;
[0170] : , ;
[0171] S75: The piecewise linear deviation correction model is applied to the initial explicit precipitation estimation formula to obtain the corrected explicit precipitation estimation formula, which serves as the final quantitative precipitation estimation model.
[0172] .
[0173] in, Precipitation level The revised precipitation estimate is as follows.
[0174] S8: Generate quantitative precipitation estimation results:
[0175] S81: Obtain the radar CAPPI three-dimensional reflectivity data for the time period to be estimated, and perform analysis and quality control according to the method in step S2;
[0176] S82: Extract the three-dimensional radar reflectivity structural features of each grid point in the area to be estimated. ;
[0177] S83: Calculate the preliminary precipitation estimate based on the initial explicit precipitation estimation formula. ;according to Determine the precipitation level and select the corresponding correction factor;
[0178] S84: Input the three-dimensional radar reflectivity structure characteristics and corresponding correction coefficients into the final quantitative precipitation estimation model to obtain the estimated hourly cumulative precipitation value of the area to be estimated.
[0179] S85: The quantitative precipitation estimation results are validated using independent validation sets (event101, 124, 128) to assess their performance in terms of rainfall intensity. For example... Figure 4 As shown in the figure, the results indicate that the average deviation of each rainfall intensity level was reduced by more than 32% after correction. Specifically, the deviation of the 5-10mm rainfall intensity segment decreased from +18.56mm to +2.13mm, and the deviation of the rainfall intensity segment above 50mm decreased from -38.93mm to -7.25mm, effectively solving the structural error problem.
[0180] All comparison methods exhibit a consistent error trend: significant overestimation in weak and moderately weak precipitation segments, severe underestimation in heavy precipitation segments, and smaller deviation in moderately heavy precipitation segments. This demonstrates that this structural error is a common technical challenge in radar precipitation estimation. The proposed rainfall intensity-based adaptive correction method improves the consistency of estimation results across different rainfall intensity levels by setting correction coefficients for different rainfall intensity levels. Therefore, this invention not only fully utilizes the high-quality reflectivity data advantage of dual-polarization radar networks but also solves the long-standing structural error problem of overestimating weak precipitation and underestimating heavy precipitation in radar precipitation estimation through innovative rainfall intensity-based adaptive error correction technology, achieving unexpected technical results and providing a high-precision, interpretable solution for operational quantitative precipitation estimation.
[0181] Single-sample calculation example
[0182] In this embodiment, the mid-level average echo characteristic of a certain grid point is defined. dBZ is used to calculate the final precipitation estimate:
[0183] 1) Calculate the initial explicit formula output: ;
[0184] 2) Calculate preliminary precipitation estimates: mm;
[0185] 3) Determine the precipitation level: mm belongs to (Heavy rainfall exceeding 50mm), select correction factor , ;
[0186] 4) Calculate the corrected target variable: ;
[0187] 5) Calculate the final precipitation estimate: mm;
[0188] Example 2: As Figure 3 As shown, the present invention also provides a machine learning-based adaptive calibration system for radar precipitation estimation, used to implement the method described in Embodiment 1, comprising:
[0189] The data acquisition module is used to acquire the CAPPI three-dimensional reflectivity product of the dual-polarization weather radar network and the hourly precipitation data of the corresponding ground rain gauge station within a preset time period.
[0190] The radar analysis and quality control module is used to analyze and control the quality of the radar CAPPI three-dimensional reflectivity data to obtain effective radar reflectivity data.
[0191] The site matching module is used to establish the spatial matching relationship between ground rain gauges and radar grids, and extract multi-height radar reflectivity data for each rain gauge location.
[0192] The event sample construction module is used to identify typical precipitation events and construct an event-scale modeling dataset based on the effective radar reflectivity data and hourly precipitation data from ground rain gauges.
[0193] The machine learning error diagnosis module is used to extract three-dimensional radar reflectivity structural features from the event-scale modeling dataset, train multiple machine learning baseline models, select the optimal baseline model through leave-one-out cross-validation, and generate error distribution features for each rainfall intensity level.
[0194] The symbolic regression module is used to model the dataset based on the event scale and use the physical constraint symbolic regression method to mine interpretable explicit precipitation estimation formulas that satisfy monotonically increasing physical constraints.
[0195] The rainfall intensity adaptive correction module is used to perform rainfall intensity adaptive error correction on the explicit precipitation estimation formula according to the error distribution characteristics of each rainfall intensity level, so as to obtain the final quantitative precipitation estimation model.
[0196] The result generation module is used to input the radar CAPPI three-dimensional reflectivity data to be estimated into the final quantitative precipitation estimation model to generate quantitative precipitation estimation results for the corresponding region based on rainfall intensity.
[0197] The modules mentioned above communicate with each other via a data bus to achieve data transmission and sharing. This invention supports distributed deployment and adopts a three-tier architecture of "data preprocessing - model calculation - result output".
[0198] 1) Data preprocessing node: Responsible for radar data parsing, quality control and site matching. A single node can process real-time data from 10 radars.
[0199] 2) Model computation node: responsible for feature extraction, precipitation estimation and rainfall intensity correction. It uses GPU to accelerate the calculation. A single node can complete an hourly precipitation estimation within 1 minute across the whole country.
[0200] 3) Result output node: responsible for converting precipitation estimation results into the GRIB2 format commonly used in meteorological operations and pushing it to the meteorological operations system.
[0201] Finally, it should be noted that the above descriptions are merely preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. An adaptive calibration method for radar precipitation estimation based on machine learning, characterized in that, Includes the following steps: S1. Obtain the CAPPI three-dimensional reflectivity product of the dual-polarization weather radar network and the hourly cumulative precipitation data of the corresponding area from the ground rain gauge station within the preset time period; S2. Analyze and perform quality control on the radar CAPPI three-dimensional reflectivity data to obtain effective radar reflectivity data; S3. Establish the spatial matching relationship between ground rain gauges and radar grids, and extract multi-height radar reflectivity data for each rain gauge location; S4. Based on effective radar reflectivity data and hourly precipitation data from ground rain gauges, identify typical precipitation events and construct an event-scale modeling dataset. S5. Extract three-dimensional radar reflectivity structural features from the event-scale modeling dataset. Using the three-dimensional radar reflectivity structural features as input and the hourly cumulative precipitation at ground level as output, train multiple machine learning baseline models. Select the optimal baseline model through leave-one-out cross-validation and generate error distribution features for each rainfall intensity level based on the optimal baseline model. S6. Based on the event-scale modeling dataset, the physical constraint symbolic regression method is used to mine interpretable explicit precipitation estimation formulas that satisfy monotonically increasing physical constraints. S7. Based on the joint error distribution characteristics of the optimal baseline model and the initial explicit precipitation estimation formula, the explicit precipitation estimation formula is subjected to adaptive error correction based on rainfall intensity to obtain the final quantitative precipitation estimation model. S8. Input the radar CAPPI three-dimensional reflectivity data to be estimated into the final quantitative precipitation estimation model to generate the quantitative precipitation estimation results for the corresponding area.
2. The adaptive calibration method for radar precipitation estimation based on machine learning according to claim 1, characterized in that: The specific steps for parsing and quality control of the radar CAPPI three-dimensional reflectivity data described in step S2 include: S21: Parse the 256-byte fixed-length header of the radar CAPPI bin file and extract the header time, layer number, and height information; S22: Remove radar files whose header time is outside the preset time period, whose header year is abnormal, or whose reading errors occur; S23: Decompress the bz2 compressed data block, convert the original int16 little-endian data to the logarithmic reflectance factor dBZ, using the following formula: , Where r is the original integer value after decompression, and s is the preset scaling factor; S24: Convert the logarithmic reflectance factor dBZ to the linear reflectance factor Z using the following formula: , Where Z is the linear reflectance factor and dBZ is the logarithmic reflectance factor; S25: Aggregate and statistically analyze the linear reflectivity factor Z within the preset time window to obtain radar reflectivity characteristic data on an hourly scale. The aggregation method includes the average value, the maximum value, or the preset quantile. S26: Based on the quality control results of the reflectivity products of the dual-polarization weather radar network, and combined with the preset reflectivity threshold and effective scanning conditions, weak echoes and system noise are screened, invalid echoes below the preset noise threshold are eliminated, and weak precipitation echo signals that meet the effectiveness conditions are retained.
3. The adaptive calibration method for radar precipitation estimation based on machine learning according to claim 1, characterized in that: The specific steps for establishing the spatial matching relationship between ground rain gauges and radar grids in step S3 include: S31: Obtain the latitude, longitude, and altitude information of each ground rain gauge station; S32: Convert latitude and longitude coordinates into row and column numbers of the radar grid, and calculate the Euclidean distance from each rain gauge station to the center of the corresponding radar grid; S33: Select ground rain gauge stations that are within a preset distance range from the center of the radar grid as valid rain gauge stations; S34: Based on the row and column numbers of the effective rain gauges, extract the log reflectance factor data of the corresponding height layer.
4. The adaptive calibration method for radar precipitation estimation based on machine learning according to claim 1, characterized in that: The specific steps for constructing the event-scale modeling dataset described in step S4 include: S41: Based on hourly precipitation data from ground rain gauges, identify periods of continuous precipitation and classify typical precipitation events; S42: For each typical precipitation event, extract hourly precipitation data from all valid rain gauge stations within the event and multi-height radar reflectivity data for the corresponding time period; S43: Calculate the cumulative precipitation at each rain gauge station within the event and filter out samples with cumulative precipitation greater than a preset threshold; S44: Perform stratified balanced sampling on the selected samples according to precipitation events and rainfall intensity levels to construct an event-scale modeling dataset.
5. The adaptive calibration method for radar precipitation estimation based on machine learning according to claim 1, characterized in that: The specific steps in step S5 for training the machine learning baseline model and generating error distribution features include: S51: Extract 3D radar reflectivity structural features from the event-scale modeling dataset, including: Low-level multi-height neighborhood average echo characteristics The arithmetic mean of the logarithmic reflectance factor of the 3×3 neighborhood within the 500m, 1000m, and 1500m height layers; Mid-level average echo characteristics The arithmetic mean of the logarithmic reflectance factor for the 2000m, 3000m, and 4000m altitude layers is... ,in , , The average log reflectance factors for the 2000m, 3000m, and 4000m altitude layers are respectively. Vertical structural features The difference in log reflectivity factor between the 4000m and 2000m altitude layers; S52: Using the three-dimensional radar reflectivity structural features as input and the hourly cumulative precipitation on the ground as output, train various types of machine learning models, including but not limited to random forest models, linear regression models, and log-linear models. S53: The leave-one-out cross-validation method is used to evaluate the performance of each machine learning model. That is, one precipitation event is kept as the test set each time, and the remaining precipitation events are used as the training set. After repeating this process multiple times, the average performance index is calculated. The evaluation indexes include root mean square error, mean absolute error, correlation coefficient and Nash efficiency coefficient. S54: Select the machine learning model with the best overall performance as the optimal baseline model based on the evaluation results; S55: Input the event-scale modeling dataset into the optimal baseline model, generate the predicted value for each sample, and calculate the average deviation, relative deviation, and error distribution curve for each rainfall intensity level to obtain the error distribution characteristics for each rainfall intensity level.
6. The adaptive calibration method for radar precipitation estimation based on machine learning according to claim 1, characterized in that: The specific steps for mining explicit precipitation estimation formulas using the physical constraint symbolic regression method described in step S6 include: S61: Perform a logarithmic transformation on the target variable, hourly cumulative precipitation R. The transformation formula is as follows: , in, This is a preset minimum value used to avoid meaningless logarithmic operations; S62: Using the three-dimensional radar reflectivity structural features as input variables and the transformed target variable y as output, a physical constraint symbolic regression method is used to search and generate multiple candidate explicit formulas. The physical constraints include formulas that monotonically increase with the input features. S63: Perform performance evaluation and complexity analysis on multiple candidate explicit formulas, and select the formula that achieves the best balance between accuracy, complexity and physical interpretability as the initial explicit precipitation estimation formula. S64: The initial explicit precipitation estimation formula is a linear formula, in the form of: , Where a is the first preset coefficient; b is the second preset coefficient; S65: Convert the initial explicit precipitation estimation formula back into a precipitation estimation formula: 。 7. The adaptive calibration method for radar precipitation estimation based on machine learning according to claim 1, characterized in that: The specific steps of the adaptive error correction for rainfall intensity described in step S7 include: S71: Based on the error distribution characteristics generated by the optimal baseline model, the hourly cumulative precipitation at ground level is divided into four precipitation levels: weak precipitation, moderate to weak precipitation, moderate to heavy precipitation, and heavy precipitation, denoted as follows: , , and ; S72: Calculate the initial explicit precipitation estimation formula for each precipitation level. Average deviation The calculation formula is: , in, Precipitation level The number of samples; Let be the predicted precipitation for the i-th sample; Let be the observed precipitation for the i-th sample; S73: For each precipitation level A piecewise linear deviation correction model is constructed. The correction coefficients are obtained by least-squares fitting of the joint error of the optimal baseline model and the initial explicit formula at this level. The joint error is the arithmetic mean of the prediction error of the optimal baseline model and the prediction error of the initial explicit formula. The corrected target variable... for: , in, Precipitation level The slope correction factor, Precipitation level The intercept correction factor; S74: For light precipitation levels and Increase the intercept correction factor The negative value range; for heavy rainfall levels Increase the slope correction factor The positive range is intended to mitigate the underestimation of heavy rainfall; S75: The piecewise linear deviation correction model is applied to the initial explicit precipitation estimation formula to obtain the corrected explicit precipitation estimation formula, which serves as the final quantitative precipitation estimation model. , in, Precipitation level The revised precipitation estimate is as follows.
8. The adaptive calibration method for radar precipitation estimation based on machine learning according to claim 1, characterized in that: The specific steps for generating quantitative precipitation estimation results in step S8 include: S81: Obtain the radar CAPPI three-dimensional reflectivity data to be estimated, and perform analysis and quality control according to the method in step S2; S82: Extract the three-dimensional radar reflectivity structural features of the region to be estimated; S83: Based on the three-dimensional radar reflectivity structure characteristics, a preliminary judgment of the precipitation level is made, and a correction coefficient corresponding to the precipitation level is selected; S84: Input the three-dimensional radar reflectivity structure characteristics and corresponding correction coefficients into the final quantitative precipitation estimation model to obtain the estimated hourly cumulative precipitation value of the area to be estimated.
9. A machine learning-based adaptive calibration system for radar precipitation estimation, applied to the machine learning-based adaptive calibration method for radar precipitation estimation described in claim 1, characterized in that, include: The data acquisition module is used to acquire the CAPPI three-dimensional reflectivity product of the dual-polarization weather radar network and the hourly precipitation data of the corresponding ground rain gauge within a preset time period. The radar analysis and quality control module is used to analyze and control the quality of the CAPPI three-dimensional reflectivity data of the dual-polarization radar network to obtain effective radar reflectivity data. The site matching module is used to establish the spatial matching relationship between ground rain gauges and radar grids, and extract multi-height radar reflectivity data for each rain gauge location. The event sample construction module is used to identify typical precipitation events and construct an event-scale modeling dataset based on the effective radar reflectivity data and hourly cumulative precipitation data from ground rain gauges. The machine learning error diagnosis module is used to extract three-dimensional radar reflectivity structural features from the event-scale modeling dataset, train multiple machine learning baseline models, select the optimal baseline model through leave-one-out cross-validation, and generate error distribution features for each rainfall intensity level. The symbolic regression module is used to model datasets based on event scales and uses the physical constraint symbolic regression method to mine interpretable explicit precipitation estimation formulas that satisfy monotonically increasing physical constraints. The rainfall intensity adaptive correction module is used to perform rainfall intensity adaptive error correction on the explicit precipitation estimation formula according to the error distribution characteristics of each rainfall intensity level, so as to obtain the final quantitative precipitation estimation model.