A rainfall monitoring data anomaly identification and data fusion model optimization and parameter calibration method

By standardizing data processing and progressive anomaly identification, combined with optimal interpolation and global optimization algorithms, the problems of singleness and parameter adaptability in anomaly identification and data fusion in traditional rainfall monitoring are solved. This achieves high-precision fusion of multi-source data and parameter consistency, improves the accuracy and stability of rainfall data, and provides reliable data support for flash flood monitoring.

CN122346754APending Publication Date: 2026-07-07CHANGJIANG RIVER SCI RES INST CHANGJIANG WATER RESOURCES COMMISSION
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHANGJIANG RIVER SCI RES INST CHANGJIANG WATER RESOURCES COMMISSION
Filing Date
2026-04-10
Publication Date
2026-07-07

AI Technical Summary

Technical Problem

Traditional rainfall monitoring methods suffer from limitations such as limited anomaly identification techniques, lack of rigorous quality control before multi-source data fusion, and poor model parameter adaptability. These limitations make it difficult to adapt to the differences in rainfall characteristics across different regions, leading to misjudgments, missed judgments, and distorted fusion results, and thus failing to meet the high-precision requirements for flash flood disaster early warning.

Method used

A standardized data processing workflow is designed, employing progressive anomaly identification, quality control followed by fusion, and joint parameter optimization. Through unified resampling, temporal continuity verification, spatial neighborhood consistency verification, and optimal interpolation, a global optimization algorithm is constructed to achieve adaptive high-precision fusion of multi-source data and parameter coordination consistency.

Benefits of technology

It has achieved a unified standard for multi-source rainfall data, improved the coverage of anomaly identification and the reliability of judgment criteria, ensured that fusion calculations are not affected by erroneous data, and ensured that model parameters are consistent, thereby improving the accuracy and stability of rainfall data and providing reliable data support for flash flood monitoring and early warning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122346754A_ABST
    Figure CN122346754A_ABST
Patent Text Reader

Abstract

The application discloses a rainfall monitoring data anomaly identification and data fusion model optimization and parameter calibration method, belongs to the field of hydrological monitoring, collects multi-source heterogeneous data and standardizes processing to generate a unified input sequence; through double verification screening dynamic updating reference station; through four-layer progressive logic identification and marking abnormal data, forming quality control data after classification rejection or correction; taking radar data as the initial field and the quality control data as the calibration point, adaptively adjusting parameters to carry out multi-source data fusion; two types of models are included in the same framework, a comprehensive objective function is constructed to simultaneously optimize the partition parameters, and closed-loop optimization is completed; after independent sample verification, it is deployed, periodically iteratively updated. The application solves the problems of traditional methods, such as incomplete anomaly identification, contaminated fusion results and poor regional adaptability, improves the rainfall data quality and fusion accuracy, provides stable and reliable data support for mountain flood disaster monitoring and early warning, and has good practical value and application prospect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of hydrological monitoring, specifically relating to a method for identifying anomalies in rainfall monitoring data and optimizing data fusion models and calibrating parameters. Background Technology

[0002] With the rapid development of technology and the construction of digital twin watersheds, accurate, stable, and spatiotemporally continuous rainfall observation data has become a crucial foundation for monitoring and early warning of flash floods. Currently, ground-based automatic rain gauge observations and radar quantitative precipitation estimation are the main technical means of obtaining rainfall information. Together, they constitute a multi-source rainfall monitoring system, providing key data support for flood forecasting, risk warning, and emergency response. The quality of this data directly affects the scientific validity and timeliness of early warning decisions, and is of great significance to protecting the lives and property of the people.

[0003] Traditional data processing methods have significant shortcomings. Their anomaly identification methods are simplistic and lack comprehensive and systematic discrimination logic, making it difficult to cover different types of data anomalies and leading to frequent misjudgments and omissions. Furthermore, strict quality control is not carried out before multi-source data fusion, and anomalous data directly participates in the fusion calculation, causing distortion of the fusion results. Data processing-related models are mostly independently designed and calibrated without forming a collaborative optimization mechanism, resulting in poor parameter adaptability. In addition, there is a lack of targeted processing solutions for the differences in rainfall characteristics in different regions, making it difficult to adapt to the high-precision monitoring needs under complex underlying surface conditions. Overall, the quality and efficiency of data processing are insufficient to meet the actual requirements of flash flood disaster early warning. Summary of the Invention

[0004] To address one or more of the shortcomings or improvement needs of existing technologies, this invention provides a method for anomaly identification, data fusion model optimization, and parameter calibration in rainfall monitoring data. By designing standardized data processing, progressive anomaly identification, quality control before fusion, and joint parameter optimization schemes, it makes multi-source rainfall data more unified and standardized in terms of temporal sequence, spatial distribution, and format. This results in a more comprehensive coverage of anomaly identification and more rigorous and reliable judgment criteria. The fusion calculation process is not affected by erroneous data, and the model parameters maintain consistency.

[0005] To achieve the above objectives, this invention provides a method for anomaly identification in rainfall monitoring data, optimization of data fusion models, and parameter calibration, comprising the following steps: S100: Collects time-series rainfall observation data from ground automatic rain gauge stations, radar quantitative precipitation estimation grid data, latitude and longitude coordinate data of monitoring stations, and watershed topography and water system geographic data; uniformly resamples the raw data with non-uniform reporting, timestamp offset, and time discontinuity to generate rainfall sequences with consistent time intervals and continuous time sequence, which serve as the unified input for all subsequent processing steps; S200: Reads the standardized time-series rainfall sequence generated by S100, performs statistical analysis based on single-station long-time-series data to determine data stability, and constructs spatial consistency verification rules within the target area; selects benchmark stations and establishes a benchmark station directory based on the stability and consistency verification results; re-executes stability analysis and verification at a preset cycle and updates the benchmark station directory; S300: Reads the standardized time-series rainfall sequence from S100 and the reference station list from S200, and sequentially performs time continuity checks, spatial neighborhood consistency checks, extreme value rationality judgments, and radar-assisted checks on the rainfall sequence; among which, the spatial neighborhood check uses the reference station determined by S200 as a reference; after completing the discrimination, abnormal data are marked, and the rainfall dataset with abnormal markings is output. S400: Reads the rainfall dataset with anomaly markers output by S300, iterates through the data and performs processing according to the marker type; removes or corrects abnormal data, retains valid data, and forms a quality-controlled rainfall dataset; S500: Reads radar quantitative precipitation estimation data from S100 and quality-controlled precipitation dataset from S400; uses radar data as the initial estimation field and quality-controlled ground data as calibration points; adaptively adjusts the search radius and correlation function according to station density, and performs fusion calculation using the optimal interpolation method; during the fusion process, it masks abnormal stations marked by S300 and outputs a spatially continuous fused precipitation field. S600: Retrieve the anomaly identification results from S300 and the fused rainfall field results from S500, and incorporate the two models into the same optimization framework; construct an objective function that integrates the anomaly identification error and the fusion error; use a global optimization algorithm to simultaneously optimize the parameters of the anomaly identification and fusion models to obtain the optimal parameter set for each partition; substitute the optimal parameters back into the anomaly identification model and the data fusion model to complete the closed-loop optimization. S700: Select independent rainstorm and flood samples to perform accuracy verification on the model whose parameters have been optimized in S600; after successful verification, solidify the optimal parameters and deploy them online; during business operation, repeat steps S200 to S600 according to a preset cycle to achieve continuous model iteration.

[0006] As a further improvement of the present invention, step S100 specifically includes the following steps: S101: Collect time-by-time rainfall observation data from ground automatic rain gauge stations, radar quantitative precipitation estimation grid data, latitude and longitude coordinate data of monitoring stations, watershed topographic and geomorphological data, and river system geographic data to form a complete raw dataset; S102: Identify and organize records with non-uniform reporting, missing timestamps, timestamp offsets, and time discontinuities in the raw data, mark problematic records, and complete preliminary cleaning; S103: Resample the original rainfall data according to a unified time scale, fill in the missing time periods with time-series interpolation, take the valid values ​​of duplicate reports, and put the disordered data back into place according to the actual occurrence time to generate a standard time-series rainfall sequence. S104: The standard rainfall sequence is associated and matched with station coordinates and watershed geographic information to form a structured dataset.

[0007] As a further improvement of the present invention, in step S200, the time stability statistical analysis includes calculating the multi-year average rainfall, coefficient of variation, data integrity rate, missing data rate, online rate, and frequency of rainfall abrupt changes. Spatial consistency verification includes defining a spatial range centered on the station, extracting concurrent rainfall from the surrounding area, and comparing the consistency of rainfall processes, the synchronicity of rainfall magnitude, and the matching degree of rainfall start and end times.

[0008] As a further improvement to the present invention, the four-layer progressive anomaly identification in step S300 is specifically as follows: Continuous anomalies are identified by dividing the area into zones based on rainfall magnitude; a neighboring verification set is constructed by extracting a reference station centered on the target station, and the magnitude difference between the target station and the neighboring mean is compared; the upper limit of rainfall intensity is determined based on design rainstorms with different return periods, and extreme value anomalies are identified; the rainfall at ground stations is compared with that at radar stations at the same location, and suspicious anomalies are marked if the difference exceeds the limit and then verified.

[0009] As a further improvement of the present invention, in step S400, abnormal data such as equipment stagnation and data freeze are eliminated or interpolated and corrected; abnormal data such as numerical mutation and spatial outlier are directly eliminated; and unmarked abnormal data and data determined to be true extreme values ​​are retained.

[0010] As a further improvement of the present invention, in step S500, the search radius is expanded when the site density is lower than a preset threshold, and the search radius is reduced when the site density is higher than a preset threshold; the spatial correlation function form is adaptively selected according to the site density, and the weight coefficient and influence radius are calculated with the goal of minimizing the error variance.

[0011] As a further improvement of the present invention, in step S600, the key parameters for synchronous optimization include search radius, level threshold, continuous value condition, spatial correlation function parameters, and weight constraint range; the comprehensive objective function is composed of the false alarm rate and false negative rate of anomaly identification, as well as the weighted average error of fusion and field average error.

[0012] As a further improvement of the present invention, in step S600, parameter calibration is performed separately for mountainous areas, plains, arid areas and glacier areas to form a regionalized optimal parameter library.

[0013] As a further improvement of the present invention, in step S700, the verification indicators include anomaly identification accuracy, fusion error, and traffic process consistency; the preset period is a fixed time interval, and the reference station list, anomaly identification rules, fusion model and parameters are dynamically updated during the iteration process.

[0014] As a further improvement of the present invention, the watershed topography and water system geographic data include DEM, river network, small watershed boundaries, land use, and soil type data; the radar quantitative precipitation estimation data is radar QPE grid precipitation field data, which includes grid latitude and longitude and resolution information.

[0015] The aforementioned improved technical features can be combined with each other as long as they do not conflict with each other.

[0016] In summary, the beneficial effects of the above-described technical solutions conceived by this invention compared with the prior art include: The present invention discloses a method for anomaly identification, data fusion model optimization, and parameter calibration in rainfall monitoring data. Through standardized data processing, progressive anomaly identification, quality control followed by fusion, and joint parameter optimization, this method ensures that multi-source rainfall data are more uniform and standardized in terms of temporal sequence, spatial distribution, and format. The anomaly identification coverage is more comprehensive, the judgment criteria are more rigorous and reliable, the fusion calculation process is not affected by erroneous data, and the model parameters maintain synergistic consistency. This achieves accurate detection and classification of rainfall anomalies, adaptive high-precision fusion of multi-source data, and optimal matching of model parameters across different regions.

[0017] (2) The rainfall monitoring data anomaly identification, data fusion model optimization, and parameter calibration method of the present invention realizes fully automated closed-loop operation without frequent manual intervention. It significantly improves the accuracy, completeness, and stability of rainfall data, while also improving the estimation accuracy and spatial continuity of watershed rainfall. It can provide more reliable and efficient data support for flash flood disaster monitoring and early warning, and has significant engineering practical value and application prospects. Attached Figure Description

[0018] Figure 1 This is a flowchart of the method for identifying anomalies in rainfall monitoring data, optimizing the data fusion model, and calibrating parameters in an embodiment of the present invention; Figure 2 yes Figure 1 A detailed step diagram of step S100; Figure 3 yes Figure 1 The detailed steps of step S300 are shown in the diagram. Detailed Implementation

[0019] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments.

[0020] It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of the invention. Furthermore, the technical features involved in the various embodiments of the invention described below can be combined with each other as long as they do not conflict with each other.

[0021] Please see Figures 1-3 The preferred embodiment of the rainfall monitoring data anomaly identification, data fusion model optimization, and parameter calibration method of the present invention includes the following steps: S100: Collects time-series rainfall observation data from ground automatic rain gauge stations, radar quantitative precipitation estimation grid data, latitude and longitude coordinate data of monitoring stations, and watershed topography and water system geographic data; uniformly resamples the raw data with non-uniform reporting, timestamp offset, and time discontinuity to generate rainfall sequences with consistent time intervals and continuous time sequence, which serve as the unified input for all subsequent processing stages.

[0022] Specifically, step S100 includes the following steps: S101: Collects time-period rainfall observation data from ground automatic rain gauges, radar quantitative precipitation estimation grid data, latitude and longitude coordinate data of monitoring stations, watershed topography and geomorphology data, and river system geographic data. In this step, the rainfall, cumulative rainfall, and observation timestamps of each rain gauge station are obtained in real time from the monitoring platform; the radar QPE grid precipitation field, grid latitude and longitude, and resolution are obtained; the rain gauge station codes, latitude and longitude, elevation, and watershed are collected; and DEM, river network, small watershed boundaries, land use, and soil type data are collected to form a complete raw dataset.

[0023] S102: Identify and organize records in the raw data that are not uniformly reported, have missing timestamps, have offset timestamps, or have time discontinuities; During this process, time-series data is traversed station by station to check for time continuity and to mark issues such as duplicate timestamps, out-of-order data, missing data, inconsistent time periods, negative rainfall values, and sudden changes without a process. Missing time periods, offset times, and discontinuities are located to form a data quality list and complete preliminary cleaning.

[0024] S103: Resample the original rainfall data according to a uniform time scale to convert irregular time-series data into a standard rainfall sequence with consistent time intervals and continuous time sequence. When performing this step, a uniform time step is set, and the original rainfall is re-statistically analyzed at fixed intervals. Missing time periods are filled in by time-series interpolation, valid values ​​are taken for duplicate reports, and out-of-order data is put back into place according to the actual time of occurrence, so as to generate a standard time-series rainfall sequence with equal intervals, no missing data, and no out-of-order data.

[0025] S104: Associate and match standard rainfall sequences with station coordinates and watershed geographic information to form a structured dataset; Through the above operations, standard time-series rainfall at each station is bound to latitude, longitude, and elevation; stations are assigned to corresponding watersheds, rivers, and administrative regions through spatial mapping; and a unified index is established to integrate rainfall, stations, watersheds, and geographic information into structured data, providing a unified input for subsequent steps.

[0026] Through the above processing, multi-source heterogeneous data can be uniformly organized and spatiotemporally aligned, effectively eliminating problems such as timing chaos and inconsistent formats caused by equipment and transmission, thereby providing a data foundation with standardized format and continuous timing for subsequent anomaly identification, spatial verification and data fusion.

[0027] S200: Read the standardized time-series rainfall sequence generated by S1, conduct statistical analysis based on single-station long-series data to determine data stability, and construct spatial consistency verification rules within the target area; select benchmark stations and establish a benchmark station directory based on the stability and consistency verification results; re-execute stability analysis and verification according to a preset cycle, update the benchmark station directory, and provide a benchmark reference for spatial verification in subsequent anomaly identification.

[0028] S201: Read the standardized time-series rainfall sequence and basic station information output by S1; In this step, standardized rainfall sequences, station coordinates, elevations, and watershed information are loaded to construct a set of stations to be analyzed.

[0029] S202: Perform statistical analysis on multi-year long-term precipitation data of a single station to determine the temporal stability of the data at each station; In this process, the multi-year average rainfall, coefficient of variation, data integrity rate, missing data rate, online rate, and frequency of rainfall abrupt changes are calculated; unstable sites with long-term missing data, process breaks, systematic drift, and frequent equipment failures are eliminated.

[0030] S203: Construct spatial consistency verification rules within the target area and conduct spatial consistency verification on rainfall data at each station; When performing this step, the spatial range is defined with the station as the center, and the surrounding concurrent rainfall is extracted; the consistency of rainfall process, the synchronicity of rainfall magnitude, and the matching degree of rainfall start and end time are compared; and spatial isolated points, significantly deviated, and abnormal stations without meteorological evidence are identified.

[0031] S204: Designate stations that simultaneously meet the requirements of temporal stability and spatial consistency as reference stations and establish a list of reference stations; Based on the above criteria, stations that pass both time stability and spatial consistency verification are marked as reference stations, forming a list of reference stations with spatial indexes.

[0032] S205: Repeat the stability analysis and spatial consistency verification according to the preset cycle, and dynamically update the list of reference stations; In this step, S202–S204 are re-executed at a fixed cycle to add, delete, or modify new sites, restored sites, faulty sites, and relocated sites, ensuring the long-term validity of the base station.

[0033] In this step, the standardized rainfall sequence is first read, and statistical analysis is performed on the long-term rainfall of a single station. The multi-year average rainfall, coefficient of variation, data completeness, missing data rate, online rate, and frequency of rainfall abrupt changes are calculated to determine the station's temporal stability. A spatial range is delineated centered on the station, and concurrent rainfall from the surrounding area is extracted. The consistency of rainfall processes, the synchronicity of rainfall magnitude, and the matching degree of rainfall start and end times are compared to determine spatial consistency. Stations that pass both temporal stability and spatial consistency checks are designated as benchmark stations, and a benchmark station directory is established and re-checked and updated at fixed intervals. Through this dual screening of temporal stability and spatial consistency, a dynamically reliable benchmark station system can be constructed, providing high-quality reference stations for subsequent spatial anomaly identification and avoiding misjudgments and omissions due to reference failure.

[0034] S300: Reads the standardized time-series rainfall sequence from S1 and the list of reference stations from S2, and sequentially performs time continuity checks, spatial neighborhood consistency checks, extreme value rationality judgments, and radar-assisted checks on the rainfall sequence; among which, the spatial neighborhood check directly uses the reference stations determined by S2 as a reference; after completing all the discriminations, it adds a mark to the abnormal data and outputs the rainfall dataset with anomaly marks.

[0035] S301: Read S1 standard time-series rainfall and S2 reference station list; in this step, load standardized rainfall, station spatial information, and reference station list to construct identification input.

[0036] S302: Perform a time continuity test on the rainfall sequence and determine continuous equal-value anomalies by dividing the rainfall series into zones based on rainfall intensity; During this process, intervals are divided according to rainfall amount, and continuous equal values ​​are judged for different levels; equipment freeze, data stagnation, and abnormal constant values ​​are identified and marked.

[0037] S303: Using the target site as the center, extract the reference stations determined in S2 within the set spatial range and construct a neighborhood check set; when performing this step, delineate the spatial range with the site as the center, filter the reference stations within the range, and form a spatial reference set.

[0038] S304: Compare the statistical results of the target site with the neighborhood check set to perform spatial neighborhood consistency verification; By comparing the above data, the mean and variance of rainfall at the benchmark station are calculated to determine the magnitude difference between the target station and the mean of the neighboring area; if the difference exceeds the threshold, spatial anomalies are marked.

[0039] S305: Determine the reasonable upper limit of extreme rainfall based on the regional design rainstorm results, and make a judgment on the reasonableness of extreme rainfall; In this step, the upper limit of rainfall intensity is determined based on the design rainstorm with different return periods; if the observed value exceeds the upper limit, it is marked as an extreme value anomaly.

[0040] S306: Compare ground station data with radar quantitative precipitation estimation data within the same spatiotemporal range to complete radar-assisted verification; in this process, extract radar rainfall at the same location and time, and compare the magnitude difference between ground and radar; if the difference exceeds the limit, mark suspicious anomalies and review them.

[0041] S307: Add an anomaly marker to any data point that is determined to be an anomaly, and generate a rainfall dataset with anomaly markers; Through the above comprehensive discrimination, time anomalies, spatial anomalies, extreme value anomalies, and radar inconsistencies are uniformly marked as abnormal states, and labeled datasets are output. This multi-layered, progressive identification logic comprehensively covers various anomalies such as equipment failures, transmission errors, numerical drift, and spatiotemporal contradictions, improving the completeness of identification while effectively reducing false alarm and false negative rates.

[0042] S400: Reads the rainfall dataset with anomaly markers output by S3, traverses the data and performs processing according to the marker type; removes or corrects abnormal data, retains valid data, and forms a quality-controlled rainfall dataset, which serves as the sole ground input source for subsequent data fusion.

[0043] S401: Read the rainfall dataset with anomaly markers output by S3; in this step, load the rainfall data with time markers, spatial markers, extreme value markers, and radar verification markers.

[0044] S402: Traverse each rainfall record to identify the type and location of anomaly markers; during this process, analyze the marker type station by station and time period by time to locate the time, location, and anomaly category of the anomaly.

[0045] S403: Perform removal or interpolation correction for abnormal data such as equipment stagnation and data freeze; During this step, continuous constant anomalies are removed or repaired by time-series interpolation to ensure that the rainfall process is reasonable.

[0046] S404: Numerical mutations and spatial outliers are directly removed and not included in subsequent calculations; In this step, sudden changes without meteorological basis and spatial anomalies are directly removed and not included in subsequent calculations.

[0047] S405: Retain unlabeled outliers and data determined to be true extreme values ​​to form a quality-controlled rainfall dataset; Through the above screening process, normal data and real extreme rainfall are retained, while all erroneous data are filtered out, resulting in a high-quality dataset.

[0048] S406: Use the quality-controlled rainfall dataset as the sole ground input data source for subsequent data fusion.

[0049] After classification, cleaning, and repair, erroneous information can be removed while retaining the true rainfall signal, further providing clean and reliable observation input for subsequent multi-source data fusion.

[0050] S500: Reads radar quantitative precipitation estimation data from S1 and quality-controlled precipitation dataset from S4; uses radar data as the initial estimation field and quality-controlled ground data as calibration points; adaptively adjusts the search radius and correlation function according to station density, and performs fusion calculation using the optimal interpolation method; during the fusion process, abnormal stations marked by S3 are masked, and finally outputs a spatially continuous fused precipitation field.

[0051] S501: Read the radar quantitative precipitation estimation data in S1 as the initial estimation field for fusion calculation; In this step, radar QPE grid data is loaded to construct a spatially continuous initial rainfall estimation field.

[0052] S502: Read the quality-controlled ground precipitation dataset output by S4 and use it as a spatial calibration point; During this process, only sites that have passed quality control are loaded; abnormal sites are not included in the fusion process.

[0053] S503: Within the fused computing grid, count the number of valid calibration points and the site distribution density; During this step, each calculation grid point is traversed to count the number and density of effective surface stations in the surrounding area.

[0054] S504: Adaptively adjusts the spatial search radius based on site density; By employing the above adjustment strategy, the search radius is expanded when the sites are sparse and reduced when the sites are dense, thus ensuring the effectiveness of the calibration.

[0055] S505: Adaptively select the corresponding spatial correlation function form based on the density of the sites; In this step, the relevant functions are automatically switched according to the site density to ensure that the interpolation field is smooth, reasonable, and physically consistent.

[0056] S506: Calculate the weight coefficient and influence radius of each grid point based on the optimal interpolation method; In this process, with the goal of minimizing the error variance, grid point weights and influence radii are calculated to achieve optimal fusion of radar and ground.

[0057] S507: Abnormal sites marked with S3 are masked during the fusion calculation process and are not included in the calculation; During this step, abnormal sites are not involved in the weight calculation and fusion process, thus avoiding fusion distortion from the source.

[0058] S508: Weighted fusion of the initial estimated field with ground calibration points to generate and output a spatially continuous fused rainfall field.

[0059] This method of fusion not only preserves the spatial continuity advantage of radar data but also corrects deviations through quality-controlled ground observations, thereby significantly improving the accuracy of regional areal rainfall estimation.

[0060] S600: Retrieve the anomaly identification results from S3 and the fused rainfall field results from S5, and incorporate the two types of models into the same optimization framework; construct an objective function that integrates the anomaly identification error and the fusion error; use a global optimization algorithm to simultaneously optimize the parameters of the anomaly identification and fusion models to obtain the optimal parameter set for each partition; substitute the optimal parameters back into the anomaly identification model and the data fusion model to complete the closed-loop optimization of the model.

[0061] S601: Retrieve S3 anomaly identification results, anomaly marking information, and S5 fused precipitation field data; In this step, the anomaly identification log, tagging results, and fused rainfall field are loaded to construct the calibration input.

[0062] S602: Incorporate the parameters of the anomaly identification model and the parameters of the data fusion model into the same set of optimization variables; In this process, the search radius, level threshold, continuous value condition, relevant function parameters, weight range, etc. are uniformly encoded as optimization variables to achieve synchronous management and collaborative optimization of the two types of model parameters.

[0063] S603: Construct a comprehensive objective function composed of a weighted sum of anomaly identification error terms and data fusion error terms; When performing this step, a comprehensive objective is constructed using anomaly identification false alarm rate, false negative rate, fusion relative error, and field error weighting.

[0064] S604: Use a global optimization algorithm to generate multiple sets of parameter combinations and substitute them into the model for calculation; By using the above optimization method, multiple sets of parameter combinations are randomly generated and substituted into the model in sequence to obtain the objective function value.

[0065] S605: With the goal of optimizing the comprehensive objective function, the optimal parameter set is obtained through iterative optimization. In this step, the optimal parameter combination is preserved through evolutionary iteration until convergence, thus obtaining the global optimal solution.

[0066] S606: Determine the optimal parameter set for each region according to different underlying surface types and climate characteristics; In this process, independent rates are set for mountainous areas, plains, arid areas, and glacier areas to form a regional parameter database.

[0067] S607: Substitute the optimal parameter set back into the anomaly detection model and the data fusion model to complete the closed-loop optimization of the model.

[0068] By using joint calibration, the parameters of the anomaly identification and data fusion model can be matched with each other, avoiding the error accumulation and amplification problems caused by independent calibration, and enabling the overall model to maintain its optimal state in various regions.

[0069] S700: Select independent rainstorm and flood samples to perform accuracy verification on the model whose parameters have been optimized in S6; after successful verification, solidify the optimal parameters and deploy them online; during business operation, repeat steps S2 to S6 sequentially according to a preset cycle to achieve continuous model iteration.

[0070] S701: Select rainstorm and flood samples independent of the calibration samples as the validation set; in this step, select the sessions that did not participate in the calibration to ensure that the validation results are objective and independent.

[0071] S702: Input the validation set into the model that has completed parameter optimization in S6, perform calculations and output the results; during this process, run the entire process of anomaly identification, quality control and fusion, and output the result data.

[0072] S703: Compare the model output with the measured data to complete anomaly identification, fusion accuracy verification, and flood simulation verification; When performing this step, verify the identification accuracy, fusion error, and flow process consistency to determine whether the target is met.

[0073] S704: After successful verification, the optimal parameter set will be saved to the model configuration file; S705: Deploy the configured model to the business runtime environment; S706: Repeat S2–S6 according to a preset cycle to achieve continuous iterative updates of the model.

[0074] After verification and deployment, the model can be adapted to long-term changes in site conditions, underlying surface conditions and rainfall patterns through periodic iterative updates, ensuring long-term stable and high-precision operation of the system.

[0075] First, through multi-source data acquisition and standardization processing (S100), heterogeneous data such as ground rain gauges and radar QPE are time-series regularized, format unified, and spatiotemporally aligned to eliminate problems such as temporal disorder and format differences in the original data, providing standardized input for subsequent processing. Then, based on time stability statistics and spatial consistency verification (S200), a dynamically updated list of reference stations is selected to provide a reliable spatial reference for anomaly identification. Next, a four-layer progressive logic of "temporal continuity - spatial neighborhood consistency - extreme value rationality - radar-assisted verification" (S300) is adopted to comprehensively identify and mark various data anomalies. Finally, through classification, elimination, and correction (S400), pure ground observation data is formed, ensuring the quality of fusion input from the source.

[0076] In the data fusion stage (S500), radar data is used as the initial estimation field, and ground data after quality control is used as the calibration point. The search radius and correlation function are adaptively adjusted according to the station density. The optimal interpolation method is used to achieve weighted fusion of multi-source data, while abnormal stations are shielded to avoid result contamination, taking into account both data spatial continuity and estimation accuracy. Then, the anomaly identification and fusion model are incorporated into the same optimization framework (S600), a comprehensive error objective function is constructed, and the partition parameters are optimized synchronously through a global optimization algorithm to solve the parameter mismatch problem caused by independent calibration. Finally, after independent sample verification, the system is deployed online, and the base station and model parameters are updated through periodic iteration (S702) to ensure that the system adapts to long-term changes and continuously outputs high-quality rainfall data, providing stable and reliable support for flash flood disaster monitoring and early warning.

[0077] Those skilled in the art will readily understand that the above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A method for anomaly identification in rainfall monitoring data, optimization of data fusion models, and parameter calibration, characterized in that, The method includes the following steps: S100: Collects time-series rainfall observation data from ground automatic rain gauge stations, radar quantitative precipitation estimation grid data, latitude and longitude coordinate data of monitoring stations, and watershed topography and water system geographic data; uniformly resamples the raw data with non-uniform reporting, timestamp offset, and time discontinuity to generate rainfall sequences with consistent time intervals and continuous time sequence, which serve as the unified input for all subsequent processing steps; S200: Reads the standardized time-series rainfall sequence generated by S100, performs statistical analysis based on single-station long-time-series data to determine data stability, and constructs spatial consistency verification rules within the target area; selects benchmark stations and establishes a benchmark station directory based on the stability and consistency verification results; re-executes stability analysis and verification at a preset cycle and updates the benchmark station directory; S300: Reads the standardized time-series rainfall sequence from S100 and the reference station list from S200, and sequentially performs time continuity checks, spatial neighborhood consistency checks, extreme value rationality judgments, and radar-assisted checks on the rainfall sequence; among which, the spatial neighborhood check uses the reference station determined by S200 as a reference; after completing the discrimination, abnormal data are marked, and the rainfall dataset with abnormal markings is output. S400: Reads the rainfall dataset with anomaly markers output by S300, iterates through the data and performs processing according to the marker type; removes or corrects abnormal data, retains valid data, and forms a quality-controlled rainfall dataset; S500: Reads radar quantitative precipitation estimation data from S100 and quality-controlled precipitation dataset from S400; uses radar data as the initial estimation field and quality-controlled ground data as calibration points; adaptively adjusts the search radius and correlation function according to station density, and performs fusion calculation using the optimal interpolation method; during the fusion process, it masks abnormal stations marked by S300 and outputs a spatially continuous fused precipitation field. S600: Retrieve the anomaly identification results from S300 and the fused rainfall field results from S500, and incorporate the two models into the same optimization framework; construct an objective function that integrates the anomaly identification error and the fusion error; use a global optimization algorithm to simultaneously optimize the parameters of the anomaly identification and fusion models to obtain the optimal parameter set for each partition; substitute the optimal parameters back into the anomaly identification model and the data fusion model to complete the closed-loop optimization. S700: Select independent rainstorm and flood samples to perform accuracy verification on the model whose parameters have been optimized in S600; after successful verification, solidify the optimal parameters and deploy them online; during business operation, repeat steps S200 to S600 according to a preset cycle to achieve continuous model iteration.

2. The method for anomaly identification and data fusion model optimization and parameter calibration of rainfall monitoring data according to claim 1, characterized in that, Step S100 specifically includes the following steps: S101: Collect time-by-time rainfall observation data from ground automatic rain gauge stations, radar quantitative precipitation estimation grid data, latitude and longitude coordinate data of monitoring stations, watershed topographic and geomorphological data, and river system geographic data to form a complete raw dataset; S102: Identify and organize records with non-uniform reporting, missing timestamps, timestamp offsets, and time discontinuities in the raw data, mark problematic records, and complete preliminary cleaning; S103: Resample the original rainfall data according to a unified time scale, fill in the missing time periods with time-series interpolation, take the valid values ​​of duplicate reports, and put the disordered data back into place according to the actual occurrence time to generate a standard time-series rainfall sequence. S104: The standard rainfall sequence is associated and matched with station coordinates and watershed geographic information to form a structured dataset.

3. The method for anomaly identification and data fusion model optimization and parameter calibration of rainfall monitoring data according to claim 1, characterized in that, In step S200, the time stability statistical analysis includes calculating the multi-year average rainfall, coefficient of variation, data completeness rate, missing data rate, online rate, and frequency of rainfall abrupt changes; Spatial consistency verification includes defining a spatial range centered on the station, extracting concurrent rainfall from the surrounding area, and comparing the consistency of rainfall processes, the synchronicity of rainfall magnitude, and the matching degree of rainfall start and end times.

4. The method for anomaly identification and data fusion model optimization and parameter calibration of rainfall monitoring data according to claim 1, characterized in that, The four-layer progressive anomaly identification in step S300 is as follows: Continuous equal-value anomalies are identified by dividing the area into zones based on rainfall magnitude; a neighboring verification set is constructed by extracting a reference station centered on the target station, and the magnitude difference between the target station and the neighboring mean is compared; the upper limit of rainfall intensity is determined based on design rainstorms with different return periods, and extreme value anomalies are identified. Compare rainfall data from ground stations with those from radar at the same location. If the difference exceeds the limit, mark it as a suspicious anomaly and verify it.

5. The method for anomaly identification and data fusion model optimization and parameter calibration of rainfall monitoring data according to claim 1, characterized in that, In step S400, abnormal data such as equipment stagnation and data freeze are removed or interpolated for correction; abnormal data such as numerical mutations and spatial outliers are directly removed; and data that are unmarked abnormalities and determined to be true extreme values ​​are retained.

6. The method for anomaly identification and data fusion model optimization and parameter calibration of rainfall monitoring data according to claim 1, characterized in that, In step S500, the search radius is expanded when the site density is lower than the preset threshold, and the search radius is reduced when the site density is higher than the preset threshold. The spatial correlation function form is adaptively selected according to the site density, and the weight coefficient and influence radius are calculated with the goal of minimizing the error variance.

7. The method for anomaly identification and data fusion model optimization and parameter calibration of rainfall monitoring data according to claim 1, characterized in that, In step S600, the key parameters for synchronous optimization include search radius, level threshold, continuous value condition, spatial correlation function parameters, and weight constraint range; the comprehensive objective function is composed of the false alarm rate and false negative rate of anomaly identification, as well as the weighted average error of fusion and field average error.

8. The method for anomaly identification and data fusion model optimization and parameter calibration of rainfall monitoring data according to claim 7, characterized in that, In step S600, parameter calibration is performed separately for mountainous areas, plains, arid areas, and glacier areas to form a regionalized optimal parameter library.

9. The method for anomaly identification and data fusion model optimization and parameter calibration of rainfall monitoring data according to claim 1, characterized in that, In step S700, the verification indicators include anomaly identification accuracy, fusion error, and traffic process consistency. The preset period is a fixed time interval, and the reference station list, anomaly identification rules, fusion model, and parameters are dynamically updated during the iteration process.

10. The method for anomaly identification and data fusion model optimization and parameter calibration of rainfall monitoring data according to any one of claims 1 to 9, characterized in that, The watershed topographic and hydrological data include DEM, river network, small watershed boundaries, land use, and soil type data; the radar quantitative precipitation estimation data is radar QPE grid precipitation field data, which includes grid latitude, longitude, and resolution information.