Rainfall data analysis method based on multimodal data fusion and spatiotemporal correlation sensing

By employing a multimodal data fusion and spatiotemporal correlation sensing-based rainfall data analysis method, the challenges of identifying and processing rain gauge blockages in karst regions were solved. This enabled accurate identification of rain gauge blockages and analysis of calcium carbonate distribution, supporting research on karst forest community succession.

CN121579565BActive Publication Date: 2026-04-03GUIZHOU ACADEMY OF TESTING & ANALYSIS
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-01-26
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Funnel-shaped rain gauges in karst regions are prone to clogging due to the formation of calcium deposits from the decomposition of calcium bicarbonate. Furthermore, regular inspections are difficult in sparsely populated areas, and data processing on cloud platforms is challenging. Therefore, it is necessary to address the challenges of rain gauge clogging identification and data analysis.

Method used

By constructing a rainfall data analysis method based on multimodal data fusion and spatiotemporal correlation perception, using historical unblocked operation and maintenance data as a reference, and integrating real-time monitoring data, a discriminant function driven by multimodal data fusion is constructed to identify blocked rain gauges, and the distribution of calcium carbonate is analyzed through spatiotemporal correlation.

Benefits of technology

It enables accurate identification of rain gauge blockage, reduces the risk of misjudgment by traditional fixed threshold judgment, adapts to complex environments, and provides data support for the study of succession of native karst forest communities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121579565B_ABST
    Figure CN121579565B_ABST
Patent Text Reader

Abstract

This invention relates to a rainfall data analysis method based on multimodal data fusion and spatiotemporal correlation sensing. The method includes: the system collecting historical unblocked operation and maintenance data of rain gauges within a target area over a set time period, partitioning and storing this data in a computer system to build a historical baseline database; the system receiving real-time multimodal monitoring data from a preset number of rain gauges within the target area to build a real-time monitoring dataset; converting the time-series data of the historical unblocked operation and maintenance data and the real-time monitoring dataset into uniformly quantifiable scores according to the values ​​at each time step; determining analysis points and surrounding points in the target area; calculating the score residuals between the values ​​at each time step and the midpoint values ​​of the baselines in the historical baseline database; generating a two-dimensional baseline residual matrix; calculating dynamic deviation indices and state persistence parameters; and constructing a blockage discrimination function. The calculation results are then combined with a preset threshold to determine blockage. This invention enables full-process automation and efficient analysis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to data processing and analysis of rainfall monitoring equipment in karst areas, and in particular to a rainfall data analysis method based on multimodal data fusion and spatiotemporal correlation sensing. Background Technology

[0002] Rainfall is a crucial data collection indicator when studying karst landforms, especially the community succession patterns in primary karst forests, thus requiring the deployment of numerous monitoring devices such as funnel-shaped rain gauges. However, in primary karst landform areas, the soil is rich in calcium carbonate. Due to carbon transfer, the weakly acidic rainwater contains a large amount of dissolved calcium carbonate in the form of calcium bicarbonate. This calcium carbonate undergoes a transformation within the funnel-shaped rain gauge, becoming a significant cause of funnel blockage.

[0003] When rainwater containing calcium bicarbonate enters the rain gauge funnel, the narrow opening of the funnel slows the flow of rainwater, causing some rainwater to temporarily stagnate on the inner wall of the funnel. Simultaneously, diurnal temperature variations and localized temperature increases caused by sunlight in the field monitoring environment accelerate the evaporation of water from the stagnant rainwater, gradually increasing the concentration of calcium bicarbonate in the rainwater system. When the concentration exceeds the solubility threshold or environmental conditions (such as temperature and pressure) change, calcium bicarbonate rapidly decomposes into solid calcium carbonate, carbon dioxide, and water. The solid calcium carbonate forms fine crystals that uniformly adhere to the inner wall of the funnel, forming a dense calcium agglomerate layer. Although initially thin, this layer alters the smoothness of the funnel's inner wall, making it easier for subsequent rainwater carrying calcium bicarbonate to adsorb, decompose, and accumulate on the surface of the calcium agglomerate. The gradually thickening calcium agglomerate layer continuously reduces the effective pore size of the funnel, especially in narrow areas such as the funnel outlet, where blockage becomes more pronounced. Furthermore, the solid deposits formed from the conversion of dissolved calcium carbonate can create a synergistic clogging effect with the core working components of the rain gauge, further exacerbating the risk of blockage. The tipping bucket rain gauge has a tipping bucket receiving port directly connected to the bottom of the funnel. This receiving port is typically only a few millimeters in diameter to ensure accurate rainwater dripping into the tipping bucket to trigger counting. When solid calcium carbonate produced by the decomposition of calcium bicarbonate adheres to the edge of the funnel outlet, it gradually forms small protrusions. These protrusions not only obstruct the smooth flow of rainwater into the receiving port but also cause some rainwater to stagnate at the outlet, leading to further decomposition of calcium bicarbonate. This causes the deposits to continuously extend into the receiving port, eventually clogging the receiving channel.

[0004] More importantly, these sediments formed from dissolved calcium carbonate are fine-grained and have strong adhesion, making them difficult to detach by the scouring force of rainwater. In the early stages of attachment, they have little impact on rain gauge monitoring data (only manifesting as a slight error in rainfall intensity), making them difficult to detect by conventional inspection methods. By the time the sediments accumulate to a certain extent and cause significant blockage, the monitoring data has been distorted for a long time, greatly increasing the lag in operation and maintenance inspections.

[0005] Unlike rain gauges deployed in urban areas, it's impossible to rely on manual, regular inspections of the numerous rain gauges located in remote, pristine karst forests. Fortunately, the development of emerging technologies such as the Internet of Things (IoT) and artificial intelligence (AI) has provided a solution. By uploading all rain gauge monitoring data to a remote cloud platform or central control system via IoT technology, and analyzing changes in this data to determine rain gauge blockages, a feasible approach has become available.

[0006] However, processing and analyzing massive amounts of data also brings new challenges. First, the data uploaded from numerous rain gauges to the cloud platform or central control system may include various disorganized data packets, such as device numbers, rainfall information, maintenance records, dates, times, locations, and other multimodal data. The inconsistent dimensions between these data make identification difficult for the cloud platform or central control system. Especially since rainfall monitoring often needs to continue for several years or even decades, it's necessary to distinguish between historical and recent data, and to account for data errors and missing information. Early-deployed rain gauges, having not yet experienced congestion, provide relatively reliable historical, uncongested maintenance data. However, with the accumulation of monitoring tasks over time, many rain gauges will become congested, leading to significant errors in the data received by the cloud platform or central control system. Furthermore, the data from a single rain gauge needs to be compared with that of surrounding rain gauges deployed in the same area; otherwise, the failure of a single rain gauge may cause data analysis failure. Determining the analysis point and surrounding points is also crucial for system construction.

[0007] Therefore, linking the two dimensions of time and space, using historical unblocked operation and maintenance data as a reference, and conducting real-time analysis and processing of multimodal data monitored by the cloud platform or central control system through a preset scoring mechanism, identifying the numbers of blocked rain gauges, and analyzing the circulation patterns of calcium carbonate in the region are of great significance for studying community succession in primary karst forests. Summary of the Invention

[0008] The main objective of this invention is to provide a rainfall data analysis method based on multimodal data fusion and spatiotemporal correlation sensing, in order to solve the following two technical problems:

[0009] The first technical challenge is to use historical, uninterrupted operation and maintenance data as a reference to fuse the massive amounts of multimodal data from rain gauge monitoring uploaded by the cloud platform or central control system into a multi-table collaborative time-series dataset with a unified scoring metric. Furthermore, by using the spatial location of a specific rain gauge as an analysis point and associating it with surrounding points as the data analysis spatial range for integrated perception, a discriminant function driven by multimodal data fusion is constructed to identify rain gauges that are experiencing blockages.

[0010] The second technical problem is to make a rough estimate of the distribution and transfer of calcium carbonate in primary karst forest areas by statistically analyzing the spatial distribution and frequency of blocked rain gauges, thereby providing guidance for research on topics such as primary karst forest community succession, organic carbon stabilization mechanisms, carbon saturation, and carbon balance.

[0011] Based on a first key aspect of the present invention, a method for analyzing rainfall data based on multimodal data fusion and spatiotemporal correlation sensing is provided, the method comprising:

[0012] The system collects historical unblocked operation and maintenance data of rain gauges in the target area within a set time period and stores them in the computer system. The historical unblocked operation and maintenance data is then divided into categories by two indicators: daily cumulative rainfall and rainfall duration. The baseline indicator range for each scenario is calculated, and a historical baseline library for the equipment is constructed.

[0013] The system receives real-time multimodal monitoring data from a preset number of rain gauges within the target area, categorizes and writes the data into a real-time temporary table, verifies the validity of the data, and then constructs a real-time monitoring dataset which is stored in the computer system.

[0014] The system uses ETL tools to convert the time-series data of historical non-blocking operation and maintenance data and real-time monitoring datasets into scores that can be uniformly quantified according to the column values ​​of each time step;

[0015] In the target area, the analysis point and surrounding points are determined according to the spatial distribution of rain gauges. The fractional residuals of each time step column value of the analysis point and surrounding points and the midpoint value of the baseline in the historical baseline database are calculated one by one to generate a two-dimensional baseline residual matrix. Based on the two-dimensional baseline residual matrix, two dynamic deviation indices and state persistence parameters are calculated.

[0016] Based on the two dynamic deviation indices and the state persistence parameter, a multimodal data fusion-driven discrimination function is constructed. The computer system determines whether the device is blocked based on the calculation results and a preset threshold.

[0017] In this invention, a historical baseline library for equipment is constructed by clustering historical unblocked operation and maintenance data. After verifying and quantifying real-time multimodal monitoring data, a two-dimensional baseline residual matrix is ​​generated by combining spatiotemporal correlation to determine the analysis point and surrounding points. Dynamic deviation indicators and state persistence parameters are calculated, and data analysis and judgment are achieved through a multimodal fusion discriminant function and preset thresholds. In essence, using historical unblocked operation and maintenance data as a reference, and considering the abnormal impact factors caused by extreme weather (such as sudden rainstorms or prolonged rainfall), quantifiable indicators are provided for the dynamic changes of subsequent monitoring data, thereby discovering equipment anomalies from the differences between real-time and historical states.

[0018] By constructing a historical baseline library from historical, uninterrupted operation and maintenance data, a quantitative standard can be provided for the normal operation status of rain gauges in corresponding scenarios, avoiding the shortcomings of traditional single baselines that cannot adapt to different rainfall intensities and durations. The subsequent operational status of rain gauges can be evaluated by comparing the difference between the continuously monitored data and the baseline.

[0019] As a further preferred embodiment, in the aforementioned method, the rain gauge is a tipping bucket rain gauge;

[0020] The validity of the verification data includes confirming whether the sampling frequency meets the standard, confirming whether the equipment is under maintenance, and confirming whether the monitored value exceeds the equipment standard.

[0021] For data whose sampling frequency does not meet the standard, add time judgment and re-transmit the data for recalculation; or classify the equipment as having a substandard sampling frequency after rain.

[0022] When determining whether the equipment is under maintenance, the system checks whether there is a maintenance schedule between the time the manufacturer was assigned and the time the schedule was terminated. If the schedule has not been terminated, the system indicates that the table is empty and stores the result in the equipment maintenance time period table.

[0023] When confirming whether the monitored value exceeds the equipment standard, first calculate whether the rainfall intensity per minute exceeds the preset value, and then determine whether the equipment should be classified as having monitored values ​​exceeding the equipment standard based on whether the rainfall value exceeds the preset value.

[0024] By verifying the compliance of sampling frequency, checking the equipment maintenance period, and determining whether the monitoring values ​​exceed the equipment standards and corresponding processing logic, the effectiveness of real-time multimodal monitoring data is ensured.

[0025] As a further preferred embodiment, in the aforementioned method, the device historical baseline library is constructed according to the following steps:

[0026] The computer system extracts the daily cumulative rainfall and rainfall duration from pre-stored historical non-blocking operation and maintenance data as core parameters for scene classification, and uses the K-means clustering algorithm to divide the categories;

[0027] For each rainfall scenario, the computer system automatically extracts key indicators from the historical non-blocking operation and maintenance data for that scenario, including average rainfall intensity, rainfall intensity fluctuation coefficient, sampling frequency compliance rate, average operating voltage, average wireless signal strength, and rain stoppage and zeroing time, and calculates the baseline range of each indicator through statistical analysis.

[0028] The computer system stores all rainfall scenarios, including scenario identifier codes, corresponding cumulative rainfall and duration intervals, baseline indicator groups, and historical data sample sizes for each scenario, in the form of structured data tables to construct a historical baseline database for the equipment.

[0029] Using daily cumulative rainfall and duration of rainfall as clustering features, the K-means algorithm is used to divide rainfall scenarios, extract key indicators for each scenario and calculate baseline ranges, and structure and store relevant information to build a historical baseline library for equipment, thereby eliminating the abnormal impact of extreme weather conditions in historical unblocked operation and maintenance data.

[0030] As a further preferred embodiment, in the aforementioned method, determining the analysis point and surrounding points based on the spatial distribution of the rain gauges includes:

[0031] Each tipping bucket rain gauge in the target area is taken as an analysis point to be analyzed. The system extracts the latitude and longitude coordinates of all rain gauges from the equipment information table and calculates the straight-line distance between the analysis point and other rain gauges based on the Haversine formula.

[0032] Using the analysis point as the center, rain gauges within a preset straight-line distance range are selected as candidate surrounding points;

[0033] For the selected surrounding points, spatial weights are calculated using the inverse square distance rule. The real-time data timestamps of the analysis point and surrounding points are simultaneously verified. Surrounding point data with timestamp errors less than a preset value are retained to form spatiotemporal correlation pairs between the analysis point and surrounding points.

[0034] The calculation of distance and the determination of spatial weights play a very important role in determining the sample range of surrounding points. The spatial linkage effect helps to eliminate data errors that may exist in a single device.

[0035] As a further preferred embodiment, in the aforementioned method, the process of constructing the two-dimensional baseline residual matrix includes:

[0036] The computer system, based on a real-time monitoring dataset, sets each element of the time series data to correspond to a sampling time step, including rainfall value, sampling frequency, operating voltage, and wireless signal strength in the time series data, and each element is associated with a unique timestamp identifier;

[0037] The computer system extracts the baseline index range for the current rainfall scenario from the device's historical baseline database. The baseline index range includes the normal fluctuation range of rainfall value, the compliant range of sampling frequency, the normal range of operating voltage, and the normal range of wireless signal strength.

[0038] The computer system calculates the residuals between the scores of each field column value at each time step in the time series data and the corresponding midpoint values ​​of the baseline index range, and constructs a two-dimensional baseline residual matrix with the time step number as the row index and the residuals of each dimension column value as the column index.

[0039] The value of each cell in the two-dimensional baseline residual matrix is ​​the residual of the corresponding time step and the corresponding parameter.

[0040] The two-dimensional baseline residual matrix can comprehensively quantify the deviation of each parameter in the real-time monitoring dataset from the corresponding parameters in the historical non-blocking operation and maintenance data, thereby discovering whether there is blockage in the equipment through changes in the equipment status.

[0041] As a further preferred embodiment, in the aforementioned method, the two dynamic deviation indices include the average relative deviation and the mutation consistency coefficient, and their calculation steps are as follows:

[0042] The computer system automatically traverses the residual value of each cell in the two-dimensional baseline residual matrix, calculates the relative deviation for each residual value, and then takes the arithmetic mean of the relative deviations for all time steps and all column values ​​to obtain the average relative deviation.

[0043] The computer system automatically calculates the abrupt change direction of the residual sequence of each column value in the two-dimensional baseline residual matrix according to the time step order, and counts the proportion of the number of times the abrupt change direction of adjacent time steps is consistent to the total number of adjacent time steps, and obtains the abrupt change consistency coefficient.

[0044] The calculation steps for the state persistence parameter are as follows: The computer system automatically sets the total number of continuous time steps within the normal range of the residuals, divides it by the total number of monitoring time steps, and obtains the state persistence parameter.

[0045] In the above scheme, by using three-dimensional quantitative analysis of the deviation characteristics between real-time data and the baseline, random interference and systematic anomalies can be accurately distinguished. Specifically, the average relative deviation quantifies the overall deviation level of all time steps and all parameters, the mutation consistency coefficient identifies the systematic pattern of residual mutations, thereby filtering out random fluctuations, and the state persistence parameter captures the cumulative effect of anomalies to eliminate instantaneous deviations. The three parameters comprehensively characterize the essence of data deviation from the overall degree, mutation pattern, and persistence characteristics, providing accurate and interference-resistant quantitative feature inputs for subsequent fusion and discrimination.

[0046] As a further preferred embodiment, in the aforementioned method, the multimodal data fusion discriminant function includes at least one of the following polynomials or a combination thereof:

[0047] The first polynomial includes the average relative deviation, mutation consistency coefficient, and state persistence parameter. Based on the synergistic effect of the overall deviation characteristics of the data and the abnormal persistence correlation characteristics, the weight is adjusted in combination with the systematic mutation characteristics to amplify the contribution of systematic anomalies and filter random fluctuations.

[0048] The second polynomial introduces the meteorological matching correlation coefficient, the terrain correction correlation coefficient, and the baseline fluctuation characteristics of historical unblocked data. Through logarithmic processing, it achieves differential calibration for different meteorological and terrain scenarios, while suppressing extreme value interference.

[0049] The third polynomial performs a smooth probabilistic transformation of the systematic mutation characteristics through an error function, supplementing the basic contribution to avoid excessively low discriminant values ​​when there are no systematic mutations.

[0050] The fourth polynomial screening method identifies deviations from historical normal levels and, combined with the secondary reinforcement effect of persistent abnormalities, emphasizes the weight of long-term significant abnormalities.

[0051] Based on a second key aspect of the present invention, a rainfall data analysis system for implementing the aforementioned method is provided, the system comprising at least one or a combination of the following modules:

[0052] The data acquisition module is used to collect historical unblocked operation and maintenance data of tipping bucket rain gauges in the target area within a set time period, as well as real-time multimodal monitoring data of a preset number of tipping bucket rain gauges in the target area.

[0053] A storage module, connected to the data acquisition module, is used to store data;

[0054] A baseline library construction module is connected to the data acquisition module and the storage module respectively, and is used to construct a device historical baseline library and store it in the storage module;

[0055] The data preprocessing module is connected to the data acquisition module and the storage module respectively. It is used to classify and write the real-time multimodal monitoring data into a real-time temporary table and perform validity verification to obtain the real-time monitoring dataset and store it in the storage module.

[0056] The quantization conversion module is connected to the storage module and the baseline library construction module respectively. It is used to call the historical non-blocking operation and maintenance data and real-time monitoring dataset in the storage module through ETL tools, and convert the two types of data into unified quantized scores according to the time step column values.

[0057] The spatiotemporal correlation module is connected to the storage module and the quantization conversion module respectively, and is used to call parameters from the storage module to determine the analysis point and surrounding points, and generate spatiotemporal correlation pairs;

[0058] The residual matrix generation module is connected to the quantization conversion module, the baseline library construction module and the spatiotemporal correlation module respectively, and is used to generate a two-dimensional baseline residual matrix and store it in the storage module.

[0059] The parameter calculation module, connected to the residual matrix generation module, is used to calculate two dynamic deviation indices and a state persistence parameter.

[0060] The discriminant analysis module is connected to the parameter calculation module and the storage module respectively. It is used to construct a discriminant function driven by multimodal data fusion based on the two dynamic deviation indices and the state persistence parameter, calculate the judgment result, and store the judgment result in the storage module.

[0061] Based on the third main aspect of the present invention, an application of the aforementioned rainfall data analysis system in a method for extrapolating organic carbon transfer in primary karst forests is provided, characterized in that the method for extrapolating organic carbon transfer in primary karst forests includes the following steps:

[0062] The tipping bucket rain gauges, which are spatially distributed within a predetermined target area, are artificially induced to become clogged, and the clog data is collected through the rainfall data analysis system.

[0063] Spatial differences in clogging initiation time, clogging accumulation rate, and clogging degree of artificially induced rain gauges were extracted and combined with natural clogging data from blank control rain gauges to remove background interference, thus obtaining effective clogging feature data.

[0064] Based on the timing of blockage occurrence, spatial distribution of cumulative rate, and spatial interpolation results of rainfall from artificially induced rain gauges, the migration path of calcium carbonate from soil via rainwater and the thermal map of carbon transfer intensity are inversely deduced.

[0065] As a further preferred embodiment, in the aforementioned system, the blockage of the tipping bucket rain gauges in the predetermined spatial distribution within the artificially induced target area is achieved by the following method: selecting the tipping bucket rain gauge as a carrier, and applying a chemically inert induction coating to the narrow opening area in the lower part of the inner wall of its funnel, the edge of the funnel outlet, and the inner wall of the tipping bucket receiving interface.

[0066] The induced coating is a modified polytetrafluoroethylene coating, a quartz coating, or an inert ceramic coating;

[0067] A blank rain gauge without an induction coating was deployed simultaneously as a control group. The blank rain gauge had the same model, latitude and longitude of deployment location, and altitude as the rain gauge under artificial induction.

[0068] Compared with existing technologies, this invention achieves accurate identification of rain gauge anomalies by constructing a scenario-based historical baseline library for devices and fusing multimodal data. Specifically, this invention provides personalized normal reference standards for different rainfall scenarios using a baseline library with daily cumulative rainfall and rainfall duration as clustering features. Combined with three-dimensional quantitative analysis of average relative deviation, mutation consistency coefficient, and state persistence parameters, it effectively distinguishes between random fluctuations and systematic anomalies, reducing the risk of misjudgment and missed judgment caused by traditional fixed threshold judgments. It is particularly suitable for the early identification of gradual anomalies such as blockage.

[0069] Secondly, this invention enhances adaptability and anti-interference capabilities in complex environments through spatiotemporal correlation sensing and full-process data quality control design. It uses the Haversine formula to filter surrounding related devices, utilizes spatial correlation to filter single-point sensor drift and local environmental interference, and combines a multi-layered data validity verification mechanism—including sampling frequency verification, maintenance period checks, and monitoring value range determination—to ensure the reliability of input data. Simultaneously, the introduction of meteorological matching coefficients and terrain correction coefficients allows the discrimination logic to adapt to different meteorological conditions and terrain features, expanding the application scenarios of the solution in karst landforms and complex rainfall areas.

[0070] This invention, after enabling remote monitoring of rain gauge blockage, can conversely induce blockage in rain gauges to analyze the overall picture of carbon transfer rate and transfer path under the influence of calcium carbonate in the target area. By deploying rain gauges with chemical inert induction coatings in karst forest areas and simultaneously setting up blank control groups, the blockage process can be controlled and the background interference can be eliminated. Combined with the rainfall data analysis system, the blockage time sequence, accumulation rate and spatial difference data are accurately collected. This invention is not only adapted to the carbon transfer characteristics of the original karst landform, but also provides an in-situ, controllable and visualized technical path for organic carbon transfer research. Attached Figure Description

[0071] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, obtaining other drawings based on these drawings without creative effort still falls within the scope of the present invention.

[0072] Figure 1 The following is a flowchart illustrating the execution of a rainfall data analysis method based on multimodal data fusion and spatiotemporal correlation sensing in one embodiment of the present invention.

[0073] Figure 2 This illustration shows a schematic diagram of the range of surrounding points considered when calculating the second congestion score in one embodiment of the present invention.

[0074] Figure 3 This paper illustrates a visualization example of reverse-engineering carbon transfer pathways based on rain gauge clogging data, according to one embodiment of the present invention. Detailed Implementation

[0075] The preferred embodiments of the present invention will be described in detail below to provide a clearer understanding of the purpose, features, and advantages of the invention. It should be understood that the following embodiments are not intended to limit the scope of the invention, but are merely illustrative of the essential spirit of the technical solution of the invention.

[0076] In the following description, certain specific details are set forth for the purpose of illustrating various disclosed embodiments in order to provide a thorough understanding of the various disclosed embodiments. However, those skilled in the art will recognize that embodiments may be practiced without one or more of these specific details. In other instances, well-known techniques associated with the invention may not have been shown or described in detail to avoid unnecessarily obscuring the description of the embodiments.

[0077] Throughout this specification, references to "an embodiment" or "an embodiment" indicate that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment. Therefore, the appearance of "in an embodiment" or "an embodiment" in various places throughout the specification does not necessarily refer to the same embodiment. Furthermore, a particular feature, structure, or characteristic may be combined in any manner in one or more embodiments.

[0078] like Figure 1 As shown, in one embodiment, a rainfall data analysis method based on multimodal data fusion and spatiotemporal correlation sensing includes the following steps 100-500 executed by a computer system:

[0079] Step 100: The system collects historical unblocked operation and maintenance data of rain gauges in the target area within a set time period and stores it in the computer system. The historical unblocked operation and maintenance data is then divided into categories by two indicators: daily cumulative rainfall and rainfall duration. The baseline indicator range for each scenario is calculated, and a historical baseline library for the equipment is constructed.

[0080] Step 200: The system receives real-time multimodal monitoring data from a preset number of rain gauges within the target area, categorizes and writes the data into a real-time temporary table, verifies the validity of the data, and then constructs a real-time monitoring dataset and stores it in the computer system.

[0081] Step 300: The system uses an ETL tool to convert the time-series data of historical non-blocking operation and maintenance data and real-time monitoring dataset into scores that can be uniformly quantified according to the column values ​​of each time step.

[0082] Step 400: In the target area, determine the analysis point and surrounding points according to the spatial distribution of rain gauges, calculate the fractional residuals between the time step values ​​of the analysis point and surrounding points and the midpoint values ​​of the baselines in the historical baseline database, and generate a two-dimensional baseline residual matrix; based on the two-dimensional baseline residual matrix, calculate two dynamic deviation indices and state persistence parameters.

[0083] Step 500: Construct a multimodal data fusion-driven discrimination function based on the two dynamic deviation indices and the state persistence parameter. The computer system determines whether the device is blocked based on the calculation results and a preset threshold.

[0084] In this invention, multimodal data refers to a collection of various types of data acquired from multiple sensing channels, which have different forms of expression but are interconnected. The idea behind this invention is to identify potential equipment anomalies by analyzing subtle changes in existing monitoring data and historical data (measured in this invention as residuals). Multimodal data is used to avoid invalid calculations generated by single data points. For example, rainfall might be used as a single data point, but rainfall monitoring data identical to historical data could be the result of different rainfall durations or instantaneous changes caused by occasional abnormal weather conditions, making it difficult to comprehensively represent the true state of the equipment. Using multimodal data, however, leverages the powerful statistical and computational capabilities of computers to quickly analyze and organize massive and complex data, thereby comprehensively measuring subtle changes in monitoring equipment.

[0085] The spatiotemporal correlation in this invention refers to the temporal correlation based on historical data and real-time monitoring data, and the spatial correlation based on the analysis point and surrounding points. In karst forest regions, rain gauge clogging is often caused by the migration and deposition of calcium carbonate. Since calcium carbonate is unevenly distributed in different spaces, the frequency and severity of clogging can vary. A single analysis point cannot comprehensively characterize the geological and meteorological conditions of the area. However, by incorporating surrounding points with similar meteorological conditions but potentially different geological conditions into the spatial correlation, information relevant to the study of calcium carbonate migration patterns can be found.

[0086] In most embodiments, the present invention is adapted to rain gauges with a tipping bucket structure, data output capability (providing monitoring data such as rainfall value and collection time), and support for real-time data transmission or storage. Typical models include the NRG Systems 2226 tipping bucket rain gauge and the Geonor T-200B series.

[0087] In step 100, the computer system can use a server equipped with a MySQL 8.0 database. First, a historical data storage area needs to be established, and historical non-blocking operation and maintenance data storage needs to be implemented. The aim is to build a structured historical data foundation to provide reliable data support for the subsequent construction of the baseline database. The specific implementation is as follows:

[0088] The database and storage area use a MySQL relational database to create a historical data storage area. It is horizontally partitioned by device number (DEVICE) and date (YYYYMMDD). Each partition corresponds to the historical data of a single device on a single day, which improves data retrieval efficiency. For example, when retrieving historical data of a single device, partition filtering can reduce the scanning range by more than 90%.

[0089] In one embodiment, the platform houses approximately 2,900 rain gauge devices. During daily monitoring, these rain gauges transmit monitoring results at set frequencies, with approximately 100,000 to 300,000 data points collected daily depending on rainfall. Historical maintenance data from rain gauges confirmed to be free of blockages through manual inspection are selected, while data with blockage records or abnormal maintenance records (such as maintenance duration > 72 hours) are excluded.

[0090] Multimodal data generally includes at least the following three core data categories:

[0091] ①Rainfall monitoring records: including equipment number, collection timestamp, rainfall value converted from tipping bucket count, sampling frequency, etc.;

[0092] ② Equipment operation log: including operating voltage, wireless signal strength, tipping bucket trigger frequency, etc.;

[0093] ③ Historical maintenance records: including maintenance start time (KSSJ), maintenance end time (JSSJ), maintenance type, etc.

[0094] The aforementioned data is written to the "Rainfall Information Table Historical Partition (rainfall_history)," "Device Run Log Table (device_run_log_history)," and "Device Maintenance Record Table Historical Partition (device_maintenance_history)," respectively, and a primary key association is established using the device number and timestamp. Note that the form names such as "rainfall_history" are merely naming conventions for ease of code input in computer systems and do not impose any limitations or practical impact on the technical solution of this invention; the same applies hereinafter.

[0095] In one embodiment, in step 200, the computer system, via an IoT gateway (e.g., using the MQTTv3.1.1 protocol), with the Broker deployed on the Alibaba Cloud IoT platform, receives multimodal monitoring data from over 2900 rain gauges in real time. The receiving frequency matches the device sampling frequency. The monitoring data typically includes at least: device number, collection timestamp, rainfall value, bucket count, operating voltage, signal strength, daily cumulative rainfall (real-time cumulative value), and rainfall duration (cumulative minutes from the first rainfall of the day to the present).

[0096] Then, a real-time temporary table is designed and data is written. Specifically, a "real-time temporary table (tmp_yl_realtime)" is created in MySQL. The fields include device number, timestamp, rainfall value, bucket count, voltage, signal strength, cumulative rainfall, duration, and data status (to be verified / valid / invalid). The InnoDB engine is used to support transactions to ensure the atomicity of data writing.

[0097] During data validity verification, the system automatically performs three levels of verification. Data that fails the verification is marked as "invalid" and written to the exception log table.

[0098] Data validity checks generally include format checks, which aim to check whether the timestamp is in the format "YYYY-MM-DD HH:MM:SS" and whether the rainfall value is a non-negative value (≥0); field integrity checks, which ensure that the four core fields of equipment number, timestamp, rainfall value, and cumulative rainfall are not empty; and data rationality checks, such as ensuring that the tipping bucket count and rainfall value meet the conversion relationship, whether the operating voltage is within a reasonable range, and whether the signal strength is reasonable, to avoid data distortion caused by weak signals.

[0099] In the following possible implementation methods, data validity verification typically includes the following situations: confirming whether the sampling frequency meets the standard, confirming whether the equipment is under maintenance, and confirming whether the monitored value exceeds the equipment standard.

[0100] For data whose sampling frequency does not meet the standard, add time judgment and re-transmit the data for recalculation; or classify the equipment as having a substandard sampling frequency after rain.

[0101] When determining whether the equipment is under maintenance, the system checks from the time the equipment was assigned by the manufacturer to the time the scheduling was terminated. If the scheduling has not been terminated, the system indicates that the equipment is under maintenance with a null value and stores the result in the equipment maintenance time period table.

[0102] When confirming whether the monitored value exceeds the equipment standard, first calculate whether the rainfall intensity per minute exceeds the preset value, and then determine whether the equipment should be classified as having monitored values ​​exceeding the equipment standard based on whether the rainfall value exceeds the preset value.

[0103] The following are some examples of specific processing steps for data validity verification in some embodiments:

[0104] 1. Verification of sampling frequencies that do not meet standards

[0105] (1) Recalculate the supplementary data

[0106] Add a time check to the system, specifically to determine if the data being calculated contains data from other dates. If so, calculate the data from those dates in chronological order. The specific logic is as follows:

[0107] First, retrieve the data to be calculated from the rainfall information table `rainfall_information` according to `storage_time`. Then, check if there is data from other dates based on `time` (collection time). If so, recalculate each day in chronological order; otherwise, only calculate the retrieved data.

[0108] (2) The sampling frequency of the equipment after rain is not equal to 6 minutes / time.

[0109] From the data in the rainfall information table (rainfall_information), the frequency of each transmission is calculated based on the device number and the time (collection time). If the current value (rainfall value) is greater than 0 and the next transmission interval is not 6 minutes, then this device is classified as "the device sampling frequency after rain is not equal to 6 minutes".

[0110] 2. Verification of whether the equipment is under maintenance.

[0111] A query of the device maintenance history table `device_maintenance_history` reveals the `eventid` field, which represents the event ID and distinguishes each maintenance event, and the `eventtype` field, which represents the event type. The complete maintenance process should be: vendor assignment, vendor confirmation and approval, assignment of maintenance personnel, maintenance personnel confirmation, maintenance personnel check-in, maintenance personnel submission, and scheduling termination. In this embodiment, based on communication with the business database technical personnel, it was learned that maintenance data prior to 2022 is test data, and the data flow is not standardized; the actual data should be based on data from 2022 onwards.

[0112] Therefore, when determining whether a device is under maintenance, the time should be from the "assigned vendor" time to the "scheduling termination" time. If the scheduling has not yet terminated, a null value should be used.

[0113] The obtained devicecode is then associated with the guiid in the deviceinfo table to obtain information such as the device's no and device information, which is then stored in the SBWHSJD table.

[0114] 3. Verification of monitored values ​​exceeding equipment standards

[0115] (1) Rainfall intensity exceeding 13 mm per minute

[0116] After completing the calculation that "the sampling frequency of the equipment is not equal to 6 minutes after the rain", the remaining rain gauges are calculated according to the equipment number and the time (collection time) for each transmission. Then, the value (rainfall value) is divided by the time difference to obtain the rainfall intensity per minute. If the rainfall intensity exceeds 12mm, the equipment is classified as "rainfall intensity exceeds 12mm per minute".

[0117] (2) Extreme values ​​were observed during monitoring.

[0118] After calculating whether the rainfall intensity per minute exceeds 12 mm, the remaining rain gauges are then judged based on the value (rainfall value). When the value (rainfall value) is greater than 80, the device is classified as "monitoring value exceeds equipment standard".

[0119] According to the method of the above embodiments, the constructed real-time monitoring dataset includes at least the following types of forms: basic data source table, intermediate association table, calculation process table, score accumulation table, and result storage table; wherein, the basic data source table includes at least a rainfall information table, an equipment maintenance record table, and an equipment information table; the intermediate association table includes at least an equipment maintenance time period table; the calculation process table includes at least a daily rainfall intensity table and a detailed score table; the score accumulation table includes at least a daily score table; and the result storage table includes at least a final score table.

[0120] In most embodiments, the process of constructing the real-time monitoring dataset includes at least the following steps: extracting valid data from a real-time temporary table and writing it into a rainfall information table; synchronizing daily maintenance records from the equipment management system to an equipment maintenance record table; generating an equipment maintenance time period table by associating the equipment information table with the equipment code; calculating the rainfall intensity per minute, the duration of the same rainfall intensity, and the total score based on the rainfall information table data and writing them into a daily rainfall intensity table; synchronously calculating the distance, weight, and score of surrounding points and writing them into a detailed score table; and associating all tables with the equipment number and timestamp.

[0121] The following embodiments will illustrate the process of building a historical baseline library for devices.

[0122] First, the computer system completes the screening of historical non-blocking operation and maintenance data and the extraction of scenario classification parameters. The computer system retrieves historical non-blocking operation and maintenance data from January 2022 to December 2024 from historical data storage areas (e.g., from the "rainfall_history" and "device_run_log_history" tables in a MySQL database), and filters valid samples using the following rules:

[0123] Data from devices with records of manual blockage, maintenance durations exceeding 72 hours, or data loss rates >5% were removed. Daily maintenance records for several rain gauges confirmed to be blocked during inspections were retained. Then, the following two core parameters for scenario classification were extracted from the filtered data:

[0124] (1) Daily cumulative rainfall (R): The rainfall value of a single device is accumulated over a calendar day (00:00-24:00), in mm, and rounded to one decimal place. (2) Duration of rainfall (T): The cumulative number of minutes from the first rainfall (rainfall value > 0.1 mm) to the last rainfall of the day is calculated and converted to hours (rounded to the nearest integer). The above parameters are associated with "device number and date" to form a sample set. ,in This is the effective sample size.

[0125] Then, the computer system segmented the rainfall scenarios based on K-means clustering. The computer system used the K-means clustering algorithm from the Python sklearn library to analyze the sample set. The specific steps for scene segmentation are as follows:

[0126] Min-max normalization (mapping to the [0,1] interval) is performed on R and T to eliminate dimensional differences, as shown in the formula.

[0127]

[0128]

[0129] and These represent the standardized values ​​of R and T, respectively.

[0130] Then, the number of clusters K is determined by the silhouette coefficient. In this embodiment, the silhouette coefficient is highest (0.78) when K=8, resulting in the best clustering effect; therefore, K=8 is set. After initializing the cluster centers, the similarity between the samples and the centers is calculated using Euclidean distance, and the centers are iteratively updated until convergence. For each clustering result, the minimum and maximum values ​​of R and T are extracted to determine the scene interval, and a unique identifier code is assigned to each scene.

[0131] For each clustering scenario, the computer system extracts six key indicators from the corresponding historical non-blocking data and determines the baseline range through statistical analysis: average rainfall intensity ( ): Calculate the arithmetic mean of all 5-minute rainfall intensities in this scenario, in mm / h; rainfall intensity fluctuation coefficient ( The ratio of the standard deviation to the mean of rainfall intensity reflects the stability of rainfall intensity. Sampling frequency compliance rate ( ): Percentage of valid records with a sampling interval of 5 minutes; Average operating voltage ( ): The arithmetic mean of the device's operating voltage; the mean of the wireless signal strength ( ): Arithmetic mean of wireless signal strength; duration of time for signal strength to return to zero after rain ( The longest duration during which rainfall values ​​return to zero after the rainfall ends. For each indicator, a 95% confidence interval is calculated as the baseline range, using the formula: Baseline lower limit = ,in, The mean of the indicators. (Standard deviation).

[0132] Finally, the computer system creates a "device historical baseline database table (baseline_lib)" in the MySQL database. During storage, a primary key index is created by scene_id to ensure unique mapping for a single scene; at the same time, a composite index of R_min and T_min is created to support quick matching of baseline indicators for the corresponding scene based on real-time cumulative rainfall and duration, providing efficient query support for subsequent residual calculations.

[0133] To adapt to factors such as equipment aging and environmental changes, the system performs baseline library updates monthly or quarterly, retrieving historical non-blocking data from the past 3-6 months, repeating the clustering and baseline calculation process, fine-tuning the scenario intervals and baseline ranges (if the difference between the old and new baselines is <10%, the original baseline is maintained; otherwise, the new baseline is updated), and retaining historical versions (distinguished by the version field) to ensure the timeliness and compatibility of the baseline library.

[0134] The construction process of the two-dimensional baseline residual matrix is ​​described in one of the following possible implementations.

[0135] First, the computer system extracts the real-time monitoring data and score data of the target rain gauge from the real-time partition table and score accumulation table of the real-time monitoring dataset, and organizes them into a time-step sequence according to a fixed sampling frequency. This sequence includes the original four types of parameters: rainfall value, sampling frequency, operating voltage, and wireless signal strength, and also incorporates a real-time congestion score. All parameters are associated with a unique timestamp. Simultaneously, the data is preprocessed to remove invalid values, and missing values ​​are filled according to scenario adaptation rules to ensure sequence continuity.

[0136] Secondly, the system converts all parameters into standardized scores of 0-100. For example, rainfall values ​​within the baseline range are scored proportionally, while those exceeding the baseline are penalized. By standardizing the units of measurement, different parameters (including blockage characteristic scores) can be directly used in subsequent calculations.

[0137] Next, based on the current rainfall scenario, the system extracts the corresponding baseline index range from the equipment's historical baseline database, including the baseline range of the original parameters and the baseline range of the congestion score, and calculates the midpoint of the score for each parameter. Subsequently, for the standardized score at each time step, the residuals with the corresponding parameter score midpoints are calculated, with residuals outside the scenario recorded as 0.

[0138] Finally, a two-dimensional baseline residual matrix is ​​constructed using the time step number as the row index and the fractional residuals of six parameters (rainfall value, sampling frequency, operating voltage, wireless signal strength, and congestion score) as the column index. The value of each cell in the matrix intuitively reflects the degree to which the corresponding time step and the corresponding parameter (including the congestion feature score) deviate from the baseline midpoint. The congestion score residual strengthens the quantification of congestion features such as persistent low rainfall intensity in rainy scenarios, and also highlights the quantification of congestion features such as surrounding rain but no rain in dry scenarios, improving the matrix's sensitivity to congestion status and providing more accurate feature input for subsequent multi-parameter fusion and discrimination.

[0139] The following examples will illustrate the algorithm for the congestion score and its processing in a real-time monitoring dataset.

[0140] For the first type of blockage score, first identify devices with a cumulative rainfall value greater than 10mm for the day (considering that a certain amount of rainfall is required to determine the blockage status of the funnel, less than 10mm is considered light rain and is not reasonable for judging whether the funnel is blocked; therefore, the total rainfall for the day must be at least moderate rain, i.e., greater than 10mm, to determine the blockage status of the funnel). Then, remove devices whose sampling frequency does not meet the standard, are under maintenance, whose monitoring values ​​exceed the equipment standard, or whose rainfall fluctuates, retaining only devices in normal condition. These devices are used as target analysis points. Calculate the rainfall intensity per minute for each target analysis point for the day, and score them based on the duration of equal rainfall intensity values. The scoring formula is shown below:

[0141] The formula for calculating the score is based on the continuous duration of constant rainfall intensity:

[0142] S=0 (t<1)

[0143] S=(t×100) / 4 (t ϵ[1~4])

[0144] S=100 (t>4)

[0145] As can be seen from the above formula, no points are awarded when the continuous duration of constant rainfall intensity is less than 1 hour, 100 points are awarded when it is greater than 4 hours, and the score is between 0 and 100 depending on the duration when it is between 1 and 4 hours.

[0146] The algorithm for rainfall intensity per minute works as follows: Subtract the previous collection time from the current data collection time to obtain the time interval, and divide the collected rainfall value by the interval time to obtain the rainfall intensity.

[0147] The calculation method for the duration of the same rainfall intensity is as follows: After obtaining the rainfall intensity record per minute in the above steps, the adjacent data with the same rainfall intensity are accumulated at time intervals. After accumulation, the data with zero rainfall intensity are removed. Then, the continuous duration (minutes) is converted to hours and compared with the score level table in this step to finally obtain the total score corresponding to the duration of each rainfall intensity.

[0148] Store this result in the daily rainfall intensity score table FRACTION_DAY_YQ.

[0149] Based on the daily rainfall intensity score table, the next step is to calculate the effective score for each target analysis point each day. It's important to note that a device may change from a blocked state to a fully open state; in this case, the score needs to be reset to zero. The method for determining this is to take the highest rainfall intensity for that device on that day and compare it with the highest-scoring rainfall intensity (if scores are the same, take the highest rainfall intensity within that score range). Two scenarios may occur:

[0150] ① The highest rainfall intensity of the day is greater than the highest rainfall intensity with the highest score: There are two scenarios: First, the highest rainfall intensity of the day occurs before the highest rainfall intensity with the highest score, meaning the end time of the highest rainfall intensity of the day is less than the start time of the highest rainfall intensity with the highest score. In this case, the historical scores are reset to zero, but the score calculated in this instance must be retained. Second, the highest rainfall intensity of the day occurs after the highest rainfall intensity with the highest score, meaning the start time of the highest rainfall intensity of the day is greater than the end time of the highest rainfall intensity with the highest score. In this case, all historical scores, including the score calculated in this instance, should be reset to zero.

[0151] ②The highest rainfall intensity of the day is equal to the highest rainfall intensity of the score: In this case, the historical calculation score should be retained, as well as the score of this calculation.

[0152] The final valid score is stored in the table VALID_SCORE_DAY_SOURE.

[0153] To reset historical scores to zero: Match the device (device), no, and date, and then delete the corresponding record from the VALID_SCORE_DAY_SOURE table.

[0154] Finally, after obtaining the daily scores, the scores are summed according to the device and no. The final total score is then stored in the table FINAL_SCORE_SOURE.

[0155] The blockage detection algorithm in this embodiment is based on the rainfall intensity of the equipment per minute. The longer the funnel maintains a constant rainfall intensity, the more likely it is to be blocked. The algorithm then accumulates the scores calculated each day and ranks the final scores to improve the accuracy of the detection.

[0156] For the second type of congestion score, the rain gauge with a cumulative rainfall of 0 for the day is first selected as the target analysis point. Centered on this target analysis point, the probability of rain at the target point is analyzed based on the rainfall data of nearby rain gauges. In one embodiment, the range of selected surrounding points is as follows: Figure 2 As shown, all rain gauges within a radius d centered on the analysis point are included in the statistical scope of surrounding points. For example, surrounding point A is included in the statistics because it falls within the radius d centered on the analysis point. The influence of surrounding points on the analysis point is as follows: the closer to the target point and the greater the rainfall value, the greater the influence on the judgment of the target point; conversely, the farther away from the target point, the smaller the influence on the judgment of the analysis point.

[0157] In this embodiment, the congestion score is calculated by accumulating scores. The influence score of surrounding points on the analysis point is calculated each day, and the scores are accumulated. If the rainfall of the device is not zero during the accumulation process, the score is reset to zero. Finally, the total scores are ranked. The device with the higher score is more likely to be in a congested state.

[0158] In practice, first identify the devices with a cumulative rainfall of 0 for the day and use them as analysis points. Then, using the analysis points as the center, calculate which rain gauges are within 1 km, 3 km, and 5 km of them. These rain gauges need to be removed if the situation is "the collection frequency does not meet the standard" or "extreme values ​​are detected". Only normal rain gauges are used for analysis.

[0159] The Haversine formula for calculating distance between two points using their latitude and longitude:

[0160] S = 2 × R × arcsin( + ×COS( .LA / 180)×COS( .LA / 180))

[0161] in, .LO is The longitude of the location of the rain gauge station; .LA The latitude of the location of the rain gauge; .LO is The longitude of the location of the rain gauge station; .LA is The latitude of the rain gauge location; Earth's radius R is taken as R=6367.138km.

[0162] Next, calculate the distance weights of surrounding points: the distance weight value is between 0 and 1, with the closer the distance, the closer the value is to 1, and the farther the distance, the closer the value is to 0; the formula for calculating the distance weights of surrounding points is:

[0163]

[0164] in, Distance weights; Indicates the nth point; This indicates the total number of surrounding points.

[0165] Then set the base score based on the rainfall value from the rain gauge;

[0166] In one embodiment, the base scores corresponding to different rainfall values ​​are as follows:

[0167]

[0168] After obtaining the actual distance S, distance weight C, and base score B for each surrounding rain gauge through the above steps, the impact score of a single rain gauge on the target rain gauge is calculated using the following formula:

[0169] The scoring formula for a single rain gauge in the surrounding area is as follows:

[0170]

[0171] After obtaining the score for each rain gauge, the location of the rain gauge is determined and divided into four quadrants. Based on the points in quadrants 1 to 4, the average score for each quadrant is calculated:

[0172]

[0173] After obtaining the average score G for each location, the scores of the four quadrants are added together, and then multiplied by the corresponding range coefficient according to the analysis range to obtain the final score for that range. In most embodiments, an analysis range coefficient T is also added: 15 for 1 km, 10 for 3 km, 6 for 5 km, 3 for 8 km, and 1 for 10 km. Finally, the final scores for 1 km, 3 km, and 5 km are added together to obtain the final score for the analysis point on that day.

[0174] For rain gauges with a daily rainfall value greater than 0, all their historical scores are cleared. After obtaining the daily scores, the historical scores of the rain gauges are accumulated, and the total scores are ranked from highest to lowest.

[0175] In this embodiment, the closer the rain gauge is to the target analysis point, the higher the rainfall value, and the higher the score. By calculating the average score of the location, the problem of high scores caused by a relatively dense concentration of rain gauges can be reduced. The calculation method of accumulating scores by scoring daily can improve the accuracy of the judgment.

[0176] Finally, the results are stored in the final score table QDS_SOURE, and the daily score table QDS_DAY_SOURE and the detailed score table SOURE_LIST_HIS are also stored.

[0177] The record types include: the equipment to be analyzed, date, rainfall, equipment number within one kilometer, distance, rainfall, orientation, the sum of the squares of the reciprocals of the distances from all points within one kilometer to the analysis point, distance weight and score between the analysis point and the target rain gauge station; equipment number within three kilometers, distance, rainfall, orientation, the sum of the squares of the reciprocals of the distances from all points within three kilometers to the analysis point, distance weight and score between the analysis point and the target rain gauge station; equipment number within five kilometers, distance, rainfall, orientation, the sum of the squares of the reciprocals of the distances from all points within five kilometers to the analysis point, distance weight and score between the analysis point and the target rain gauge station; equipment number within eight kilometers, distance, rainfall, orientation, the sum of the squares of the reciprocals of the distances from all points within eight kilometers to the analysis point, distance weight and score between the analysis point and the target rain gauge station; and equipment number within ten kilometers, distance, rainfall, orientation, the sum of the squares of the reciprocals of the distances from all points within ten kilometers to the analysis point, distance weight and score between the analysis point and the target rain gauge station.

[0178] In some embodiments, the analysis focuses on two scenarios: 40-0 and 40-40-0. Assuming an error of 10mm, a change in rainfall from 40-50mm to 0-10mm is considered an abnormal device and is included in the list of abnormal devices. A 40-40-40-0 change is excluded from this calculation because it has already triggered a separate warning process. The algorithm for analyzing rainfall changes first filters out abnormal devices. Then, it iterates through the data, recording the rainfall for each device for four consecutive days, identifying the 40-0 and 40-40-0 scenarios, and categorizing them as "rainfall changes" to include them in the list of abnormal devices.

[0179] In one possible implementation, the two dynamic deviation metrics include the average relative deviation and the mutation consistency coefficient, which are calculated as follows:

[0180] The computer system first iterates through the residual values ​​of all cells in the two-dimensional baseline residual matrix, where the matrix dimension is n×m, n is the total number of time steps, and m is the parameter dimension (e.g., m=6 when there are 6 types of parameters such as rainfall and sampling frequency). For each residual value... (No. Time step, first The relative deviation of a parameter is calculated by taking the baseline range of the corresponding parameter as the benchmark, using half the width of the range as the denominator, and the absolute value of the residual as the numerator, i.e., the relative deviation. If the baseline range half-width is 0 (e.g., the sampling frequency baseline range is fixed at [100, 100] minutes), then the relative deviation is directly taken as the absolute value of the residual (because when the denominator is 0, the residual itself already reflects the absolute deviation).

[0181] After the traversal is complete, the system processes all n×m relative deviations. Take the arithmetic mean, i.e., the average relative deviation:

[0182]

[0183] This value quantifies the overall deviation of all time steps and all parameters, and its range is [0, +∞). The larger the value, the more significant the overall deviation from the baseline.

[0184] Then, the computer system extracts the residual sequence for each class of parameters (m classes in total) in the two-dimensional baseline residual matrix. (j=1 to m). For each sequence, calculate the direction of residual mutation between two adjacent time steps in time step order: if - This is denoted as a "positive mutation" (symbol +1); if - A negative mutation is denoted as -1; if the difference is 0, it is denoted as no mutation (symbol 0).

[0185] Subsequently, the system statistically analyzes the consistency of mutation directions for all parameters within the same adjacent time step (i and i-1). Specifically, for the i-th group of adjacent time steps (i=2 to n), if more than 70% of the mutation direction signs in the m types of parameters are the same (both +1 or both -1), it is considered "consistent in direction," and the count is incremented by 1. The total number of adjacent time steps is n-1. Since there are n-1 adjacent pairs in n time steps, the mutation consistency coefficient... This represents the proportion of "consistent" occurrences among all adjacent time steps to the total number of adjacent time steps. The value ranges from [0,1]. A higher value indicates that the residual mutations are systematic (non-random fluctuations).

[0186]

[0187] Wherein, n-1 represents the total number of adjacent time steps (n time steps contain n-1 pairs of adjacent pairs). It is an indicator function for the consistency of mutation directions between adjacent time steps, used to mark whether the residual mutation directions of the i-th group of adjacent time steps (i.e., time steps i and i-1) satisfy the consistency condition.

[0188] For calculating the state persistence parameters, the computer system first sets the normal range for the residuals: for each type of parameter j, the normal range is... ,in The threshold is 20% of the half-width of the baseline j range, and this threshold is determined based on the residual fluctuation characteristics of historical non-blocking data.

[0189] The system checks each time step whether the residuals of all parameters are within the normal range. For the i-th time step, if... If the condition holds true for all j=1 to m, it is marked as a "normal time step"; otherwise, it is marked as an "abnormal time step". The total duration of consecutive normal time steps is calculated and divided by the total monitoring duration to obtain the state persistence parameter. :

[0190]

[0191] The value ranges from [0,1]. A higher value indicates a more stable device state (a shorter duration of deviation from the baseline).

[0192] After completing the statistical analysis of the above parameters, the computer system performs calculations using a multimodal data fusion discriminant function. In most embodiments of the present invention, the multimodal data fusion discriminant function includes at least one or a combination of the following polynomials:

[0193] The first polynomial includes the average relative deviation, mutation consistency coefficient, and state persistence parameter. Based on the synergistic effect of the overall data deviation characteristics and the anomaly persistence correlation characteristics, it combines systematic mutation characteristics for weight adjustment, amplifying the contribution of systematic anomalies and filtering random fluctuations. In one possible implementation, the first polynomial is expressed as follows:

[0194]

[0195] in, The average relative deviation. The mutation consistency coefficient, This is the state persistence parameter.

[0196] The design goal of this polynomial is to accurately capture the core anomaly characteristics of overall deviation and persistent anomalies, and to filter out random fluctuations through systematic mutation features, providing a basic quantitative basis for the discriminant function. From the perspective of solving technical problems, the core manifestation of rain gauge anomalies (such as calcium deposition blockage in karst areas) is that the data deviates from the normal baseline and persists, rather than random instantaneous fluctuations. Therefore, it is necessary to prioritize the fusion of average relative deviation. With state persistence parameter To avoid ignoring persistent low-level anomalies, a combination of linear terms and square root terms is used. This approach preserves the strong contribution of significant anomalies while also taking into account the fundamental impact of weak anomalies. And through... The product of this combination achieves a collaborative logic where the greater the deviation and the longer it lasts, the stronger the abnormal signal.

[0197] Meanwhile, to amplify systematic mutations such as regular data shifts caused by congestion and suppress random fluctuations, a 1- As the denominator—when When the polynomial approaches 1 (a complete systemic mutation), the denominator approaches 0, the polynomial value is significantly amplified, and the core anomaly is quickly identified. When the value is 0 (random fluctuation), the denominator is 1, and the polynomial value is determined only by the basic characteristics, thus avoiding misjudgment.

[0198] The second polynomial incorporates meteorological matching correlation coefficients, terrain correction correlation coefficients, and baseline fluctuation characteristics of historical unblocked data. It uses logarithmic processing to achieve differential calibration for different meteorological and terrain scenarios while suppressing extreme value interference. In one possible implementation, the second polynomial is expressed as follows:

[0199]

[0200] in, The baseline residual standard deviation of historical non-blocking data. For meteorological matching coefficients, This is the terrain correction factor.

[0201] The core objective of this polynomial is to address the inconsistency in discrimination criteria under different meteorological and topographical scenarios, thereby improving the function's adaptability in complex environments. This invention can adapt to different terrains such as karst forests and plains, as well as different meteorological conditions such as heavy rain and light rain. These scenarios can lead to differences in the baseline of rain gauge data fluctuations—for example, concentrated rainfall in karst depressions increases the risk of calcium deposition and blockage, requiring more sensitive discrimination criteria. Therefore, a meteorological matching coefficient is introduced. With terrain correction factor Used to dynamically adjust scene adaptation weights; combined with the baseline residual standard deviation of historical non-blocking data. ,pass Normalize the scene coefficients to eliminate calibration deviations caused by different historical normal fluctuation ranges.

[0202] To avoid excessive interference from extreme values ​​of scene coefficients on the overall discrimination result, logarithmic operations are employed. The slow growth of the logarithmic function suppresses the impact of extreme values ​​while preserving the monotonic trend that the stronger the scene adaptation requirement, the greater the correction contribution. Furthermore, a constant of 1 is added to ensure that the correction term remains positive, without weakening the fundamental contribution of the core features. This ultimately forms a scene adaptation correction structure multiplied by a polynomial fusion with the core features, achieving synergy between core abnormal signals and scene-differentiated calibration.

[0203] The third polynomial performs a smooth probabilistic transformation of the systematic mutation characteristics through an error function, supplementing the fundamental contribution to avoid excessively low discriminant values ​​when there are no systematic mutations. In one possible implementation, the third polynomial is expressed as follows:

[0204]

[0205] in, Let be the error function. The mutation consistency coefficient.

[0206] This polynomial aims to represent discrete systematic mutation characteristics. This is transformed into a smooth, probabilistic output, avoiding misjudgments caused by step-like decisions, while also supplementing the discriminant function with a basic contribution. Because... The range of values ​​is [0,1]. Direct substitution can easily lead to step errors with similar degrees of abrupt change but huge differences in discriminant values. Therefore, the Gaussian error function is chosen. This function can smoothly map input values ​​to the [-1, 1] interval, meeting the requirement of gradual variable transformation of mutation features. To make... (Complete systemic mutation) When the function output is close to its peak value, The output is 0 when there are no systematic mutations. Multiply (Approximately 0.886), achieving input range and Precise matching of the effective response range of a function.

[0207] Considering that in the absence of systematic mutations, the discriminant function still needs to retain its basic contribution to maintain a reasonable range of values, in The function output is superimposed with 0.5, which locks the polynomial's value range in the range of [0.5, 1.3]. This avoids the discriminant value being too low when there is no mutation, and also avoids overshadowing the dominant role of the core features.

[0208] The fourth polynomial filters out deviations exceeding historical normal levels, and, combined with the secondary reinforcement effect of persistent anomalies, highlights the weight of long-term significant anomalies. In one possible implementation, the fourth polynomial is expressed as follows:

[0209]

[0210] in, For historical non-blocking data The 75th percentile, The average relative deviation. This is the state persistence parameter.

[0211] The purpose of this polynomial design is to distinguish between slight deviations within the normal fluctuation range and significant anomalies exceeding normal levels, thereby strengthening the contribution weight of long-term severe anomalies and thus conforming to the characteristics of gradual anomalies such as blockage. Only when Exceeding the 75th percentile of historical non-congestion data Only when this is determined is it considered a significant deviation, therefore the following method is adopted. .when When the value exceeds a certain threshold, this parameter is reset to zero, filtering out minor deviations. Conversely, it retains the contribution of the excess portion, focusing on significant anomalies. Furthermore, anomalies such as blockages have a cumulative effect that becomes more severe with longer duration; therefore, the state persistence parameter is considered... Take the square This causes the contribution of long-term anomalies to grow non-linearly, highlighting the core weight of persistent and significant anomalies.

[0212] In one possible implementation, the complete multimodal data fusion discriminant function can be expressed as follows:

[0213]

[0214] in, For multimodal data fusion discriminant function, The average relative deviation. The mutation consistency coefficient, For state persistence parameters, The baseline residual standard deviation of historical non-blocking data. For historical non-blocking data The 75th percentile, Let be the error function. For meteorological matching coefficients, This is the terrain correction factor.

[0215] In the expression, the third polynomial As the denominator, it achieves a linkage between the magnitude of enhancement and the degree of mutation, that is, when When the denominator is high (systematic anomaly), the reinforcement amplitude is moderate, avoiding overjudgment. When Low but When the threshold is significantly exceeded (such as a sudden severe blockage), the denominator is small, and the amplification amplitude is increased to ensure that no emergency anomaly is missed.

[0216] The above function expressions are not unique. In actual system settings, those skilled in the art can make adaptive adjustments to the system according to different research objectives. In practical applications, the selection is based on various comprehensive factors of the target area during the monitoring period. For example, if the target area is flat and the influence of terrain on the monitoring results is negligible, the terrain correction coefficient can be discarded. If no abnormal weather conditions occurred during the monitoring period, the weather matching coefficient can be cancelled.

[0217] In most embodiments, the computation of the multimodal data fusion discriminant function includes the following steps:

[0218] The computer system first retrieves the static data set of the current scene from the device's historical baseline library. , Call the preset Value and Functional model; simultaneously, based on the time-series data of the current sliding window, dynamic data sets are calculated in real time. , , , and The dynamic parameters are normalized to the [0-1] interval; the static and dynamic data sets are substituted into the multimodal data fusion discrimination function; the computer system automatically executes the function operation, wherein the static data set remains constant during the operation, and the dynamic data set is updated once every set time sliding window and re-substituted into the function to generate real-time discrimination values. .

[0219] Finally, the computer system makes a judgment based on the calculation results and preset thresholds, including the following steps: The computer system extracts the judgment thresholds corresponding to the rain gauge and rainfall scenario from the preset threshold library, verifies the rationality of the thresholds, and triggers a self-calibration mechanism if it is unreasonable; it generates a judgment value sequence according to a set time sliding window based on the judgment values ​​output by the multimodal data fusion discriminant function, smooths the judgment value sequence, removes isolated outliers and replaces them; it sets the number of continuous verification windows, traverses the processed judgment value sequence, and marks blocked and normal candidate states; it calls real-time meteorological data, combines it with the actual rainfall status to confirm or exclude the candidate states, and marks suspected states that need to be manually reviewed in combination with the equipment operating parameters; it associates the final judgment result with the equipment number, judgment time period and other information to generate a report, triggers hierarchical alarms and records them to the computer system.

[0220] After the rainfall data analysis system based on multimodal data fusion and spatiotemporal correlation sensing of this invention is built, run, and verified, as a possible implementation, this system can be used in the method of extrapolating organic carbon transfer in primary karst forests. Specifically, a possible method for extrapolating organic carbon transfer in primary karst forests includes the following steps:

[0221] Clogging of tipping bucket rain gauges within a predetermined spatial distribution within a target area is artificially induced, and clogging data is collected using the rainfall data analysis system. Spatial differences in clogging initiation time, clogging accumulation rate, and clogging severity of the artificially induced rain gauges are extracted. These are then combined with natural clogging data from a control rain gauge to subtract background interference, yielding effective clogging characteristic data. Based on the clogging sequence, spatial distribution of accumulation rate, and spatial interpolation results of rainfall, the migration path of calcium carbonate from soil via rainwater and the carbon transfer intensity thermogram are inversely extrapolated. Figure 3 As shown, the rain gauges in the target area are displayed in a spatial plane based on the order and degree of blockage of the rain gauges determined by the system (compared to the set threshold), so that the induced blockage result in the target area can be intuitively estimated.

[0222] The blockage of the tipping bucket rain gauges within the predetermined spatial distribution area in the artificially induced target area is achieved through the following method: A tipping bucket rain gauge is selected as the carrier, and a chemically inert induction coating is applied to the narrow opening area in the lower part of its funnel, the edge of the funnel outlet, and the inner wall of the tipping bucket receiving interface; the induction coating is a modified polytetrafluoroethylene coating, a quartz coating, or an inert ceramic coating; a blank rain gauge without an induction coating is simultaneously deployed as a control group, and the blank rain gauge has the same model, latitude and longitude, and altitude as the artificially induced rain gauge.

[0223] The essence of induced blockage is to regulate the calcium bicarbonate in rainwater in karst areas through physical coatings. The depositional environment accelerates its transformation into solid calcium carbonate. The calcium deposits are transformed and adhered to key parts of the rain gauge funnel, forming a controllable calcium layer blockage, which is done in two steps:

[0224] Firstly, physical induction is considered, utilizing coatings to regulate rainwater retention and evaporation, creating conditions for concentration. For example, surface roughness and hydrophilicity are designed, with the coating surface roughness controlled at Ra=0.5-2μm and a contact angle of 60-90° (moderate hydrophilicity). This prevents rainwater from sliding off quickly and forms a stable thin layer of retained water at key locations such as the narrow opening of the funnel and the receiving interface, rather than dripping directly. These areas are where calcium deposits preferentially form during natural blockage, and the coating further enhances the retention effect.

[0225] Secondly, we consider adapting to the microenvironment of the wild to accelerate evaporation. Due to the diurnal temperature difference and localized warming caused by daytime sun exposure in karst forests, the coating, as an inert material, does not absorb moisture but only acts as a retention carrier to accelerate the evaporation of retained rainwater, so that the concentration of calcium bicarbonate in the rainwater quickly exceeds the solubility threshold, creating a prerequisite for subsequent chemical decomposition.

[0226] The calcium bicarbonate in rainwater in karst regions originates from the initial reaction between calcium carbonate in the soil and weakly acidic rainwater (soil carbon transfer process). When rainwater enters the funnel coated with an induction layer, the concentration effect caused by this physical induction disrupts the dissolution equilibrium of calcium bicarbonate, causing it to rapidly decompose into solid calcium carbonate. The calcium carbonate crystals produced by decomposition are uniformly deposited on the inner wall of the funnel due to the adhesion sites provided by the surface roughness of the coating, initially forming a thin calcium agglomerate layer. Subsequently, the calcium bicarbonate carried by rainwater will be further adsorbed, concentrated, and decomposed on the surface of the calcium agglomerate layer, gradually thickening the calcium agglomerate layer, eventually reducing the funnel pore size and blocking the receiving interface, achieving controllable induced clogging.

[0227] In this process, it is necessary to ensure that the inducing coating is a chemically inert material that does not react chemically with calcium bicarbonate or rainwater. It should only accelerate the conversion of dissolved to solid states during the natural carbon transfer process through physical environmental regulation, so as to ensure that the blockage data can accurately reflect the intensity of regional carbon transfer.

[0228] The induced clogging process involves two key chemical reactions in the carbon cycle of karst regions, corresponding to soil carbon dissolution and migration of carbon to rainwater, and rainwater carbon deposition and clogging formation, as detailed below:

[0229] Karst forest soils are rich in calcium carbonate ( Rainwater absorbs airborne pollutants during its fall. It forms weakly acidic carbonic acid (carbon dioxide). Carbonic acid reacts with calcium carbonate in the soil to form water-soluble calcium bicarbonate, which seeps into the surface runoff with rainwater and eventually enters rain gauges. This is the core step in the transfer of carbon from soil to water bodies, and the reaction formula is as follows: The reaction conditions are a weakly acidic environment and room temperature.

[0230] After entering the funnel coated with the induction coating, the rainwater loses moisture and increases calcium bicarbonate concentration due to retention and evaporation. Simultaneously, the localized temperature rise caused by outdoor sun exposure decreases. The solubility in water, combined with the above conditions, drives the reversible reaction in the reverse direction, causing calcium bicarbonate to decompose into solid calcium carbonate. When water is dissolved, the calcium carbonate crystals produced adhere to the coating surface, forming a calcium agglomerate layer, which directly causes the funnel to become clogged.

[0231] After the rain gauge is artificially clogged, the rainfall data analysis system analyzes the time and spatial distribution of the clog to... Figure 3The displayed planar graphics allow for a preliminary exploration of the activity patterns of calcium carbonate within the target area, providing valuable data for the study of the stabilization mechanism of organic carbon in karst forest regions affected by calcium carbonate.

[0232] The technical terms, principles, or means related to the technical solutions of the present invention mentioned in the above embodiments, which are not described in detail above, are all well-known technologies or common practices that are known to those skilled in the art.

[0233] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely illustrative of the principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the present invention as claimed. The scope of protection of this invention is defined by the appended claims and their equivalents.

Claims

1. A rainfall data analysis method based on multimodal data fusion and spatiotemporal correlation sensing, characterized in that, The method includes: The system collects historical unblocked operation and maintenance data of rain gauges in the target area within a set time period, stores it in the computer system in partitions, and divides the historical unblocked operation and maintenance data into categories according to two indicators: daily cumulative rainfall and rainfall duration. It calculates the baseline indicator range for each scenario and builds a historical baseline library for the equipment. The system receives real-time multimodal monitoring data from a preset number of rain gauges within the target area, categorizes and writes the data into a real-time temporary table, verifies the validity of the data, and then constructs a real-time monitoring dataset which is stored in the computer system. The system uses ETL tools to convert the time-series data of historical non-blocking operation and maintenance data and real-time monitoring datasets into scores that can be uniformly quantified according to the column values ​​of each time step; In the target area, the analysis point and surrounding points are determined according to the spatial distribution of rain gauges. The fractional residuals of each time step column value of the analysis point and surrounding points and the midpoint value of the baseline in the historical baseline database are calculated one by one to generate a two-dimensional baseline residual matrix. Based on the two-dimensional baseline residual matrix, two dynamic deviation indices and state persistence parameters are calculated. Based on the two dynamic deviation indices and the state persistence parameter, a multimodal data fusion-driven discriminant function is constructed. The computer system determines whether the device is blocked based on the calculation results and a preset threshold. The process of constructing the two-dimensional baseline residual matrix includes: The computer system, based on a real-time monitoring dataset, sets each element of the time series data to correspond to a sampling time step, including rainfall value, sampling frequency, operating voltage, and wireless signal strength in the time series data, and each element is associated with a unique timestamp identifier; The computer system extracts the baseline index range for the current rainfall scenario from the device's historical baseline database. The baseline index range includes the normal fluctuation range of rainfall value, the compliant range of sampling frequency, the normal range of operating voltage, and the normal range of wireless signal strength. The computer system calculates the scores of each field column value at each time step in the time series data and the residuals between them and the midpoint values ​​of the corresponding baseline index range. A two-dimensional baseline residual matrix is ​​constructed using the time step number as the row index and the residuals of each dimension column value as the column index. The value of each cell in the two-dimensional baseline residual matrix is ​​the residual of the corresponding time step and the corresponding parameter; The two dynamic deviation indices include the average relative deviation and the mutation consistency coefficient, and their calculation steps are as follows: The computer system automatically traverses the residual value of each cell in the two-dimensional baseline residual matrix, calculates the relative deviation for each residual value, and then takes the arithmetic mean of the relative deviations for all time steps and all column values ​​to obtain the average relative deviation. The computer system automatically calculates the abrupt change direction of the residual sequence of each column value in the two-dimensional baseline residual matrix according to the time step order, and counts the proportion of the number of times the abrupt change direction of adjacent time steps is consistent to the total number of adjacent time steps, and obtains the abrupt change consistency coefficient. The calculation steps for the state persistence parameter are as follows: The computer system automatically sets the total number of continuous time steps within the normal range of the residuals, divides it by the total number of monitoring time steps, and obtains the state persistence parameter.

2. The rainfall data analysis method based on multimodal data fusion and spatiotemporal correlation sensing according to claim 1, characterized in that, The rain gauge is a tipping bucket rain gauge; The validity of the verification data includes confirming whether the sampling frequency meets the standard, confirming whether the equipment is under maintenance, and confirming whether the monitored value exceeds the equipment standard. For data whose sampling frequency does not meet the standard, add time judgment and re-transmit the data for recalculation; or classify the equipment as having a substandard sampling frequency after rain. When determining whether the equipment is under maintenance, the system checks whether there is a maintenance schedule between the time the manufacturer was assigned and the time the schedule was terminated. If the schedule has not been terminated, the system indicates that the table is empty and stores the result in the equipment maintenance time period table. When confirming whether the monitored value exceeds the equipment standard, first calculate whether the rainfall intensity per minute exceeds the preset value, and then determine whether the equipment should be classified as having monitored values ​​exceeding the equipment standard based on whether the rainfall value exceeds the preset value.

3. The rainfall data analysis method based on multimodal data fusion and spatiotemporal correlation sensing according to claim 1, characterized in that, The device historical baseline library is constructed according to the following steps: The computer system extracts the daily cumulative rainfall and rainfall duration from pre-stored historical non-blocking operation and maintenance data as core parameters for scene classification, and uses the K-means clustering algorithm to divide the categories; For each rainfall scenario, the computer system automatically extracts key indicators from the historical non-blocking operation and maintenance data for that scenario, including average rainfall intensity, rainfall intensity fluctuation coefficient, sampling frequency compliance rate, average operating voltage, average wireless signal strength, and rain stoppage and zeroing time, and calculates the baseline range of each indicator through statistical analysis. The computer system stores all rainfall scenarios, their corresponding cumulative rainfall and duration intervals, baseline indicator groups, and historical data sample sizes in a structured data table format, thus constructing a historical baseline database for the equipment.

4. The rainfall data analysis method based on multimodal data fusion and spatiotemporal correlation sensing according to claim 1, characterized in that, The process of determining the analysis point and surrounding points based on the spatial distribution of rain gauges includes: Each tipping bucket rain gauge in the target area is taken as an analysis point to be analyzed. The system extracts the latitude and longitude coordinates of all rain gauges from the equipment information table and calculates the straight-line distance between the analysis point and other rain gauges based on the Haversine formula. Using the analysis point as the center, rain gauges within a preset straight-line distance range are selected as candidate surrounding points; For the selected surrounding points, spatial weights are calculated using the inverse square distance rule. The real-time data timestamps of the analysis point and surrounding points are simultaneously verified. Surrounding point data with timestamp errors less than a preset value are retained to form spatiotemporal correlation pairs between the analysis point and surrounding points.

5. The rainfall data analysis method based on multimodal data fusion and spatiotemporal correlation sensing according to claim 1, characterized in that, The discriminant function driven by the multimodal data fusion includes at least one of the following polynomials or a combination thereof: The first polynomial includes the average relative deviation, mutation consistency coefficient, and state persistence parameter. Based on the synergistic effect of the overall deviation characteristics of the data and the abnormal persistence correlation characteristics, the weight is adjusted in combination with the systematic mutation characteristics to amplify the contribution of systematic anomalies and filter random fluctuations. The second polynomial introduces the meteorological matching correlation coefficient, the terrain correction correlation coefficient, and the baseline fluctuation characteristics of historical unblocked data. Through logarithmic processing, it achieves differential calibration for different meteorological and terrain scenarios, while suppressing extreme value interference. The third polynomial performs a smooth probabilistic transformation of the systematic mutation characteristics through an error function, supplementing the basic contribution to avoid excessively low discriminant values ​​when there are no systematic mutations. The fourth polynomial screening method identifies deviations from historical normal levels and, combined with the secondary reinforcement effect of persistent abnormalities, emphasizes the weight of long-term significant abnormalities.

6. A rainfall data analysis system for implementing the method of any one of claims 1-5, characterized in that, The system includes at least one or a combination of the following modules: The data acquisition module is used to collect historical unblocked operation and maintenance data of tipping bucket rain gauges in the target area within a set time period, as well as real-time multimodal monitoring data of a preset number of tipping bucket rain gauges in the target area. A storage module, connected to the data acquisition module, is used to store data; A baseline library construction module is connected to the data acquisition module and the storage module respectively, and is used to construct a device historical baseline library and store it in the storage module; The data preprocessing module is connected to the data acquisition module and the storage module respectively. It is used to classify and write the real-time multimodal monitoring data into a real-time temporary table and perform validity verification to obtain the real-time monitoring dataset and store it in the storage module. The quantization conversion module is connected to the storage module and the baseline library construction module respectively. It is used to call the historical non-blocking operation and maintenance data and real-time monitoring dataset in the storage module through ETL tools, and convert the two types of data into unified quantized scores according to the time step column values. The spatiotemporal correlation module is connected to the storage module and the quantization conversion module respectively, and is used to call parameters from the storage module to determine the analysis point and surrounding points, and generate spatiotemporal correlation pairs; The residual matrix generation module is connected to the quantization conversion module, the baseline library construction module and the spatiotemporal correlation module respectively, and is used to generate a two-dimensional baseline residual matrix and store it in the storage module. The parameter calculation module, connected to the residual matrix generation module, is used to calculate two dynamic deviation indices and a state persistence parameter. The discriminant analysis module is connected to the parameter calculation module and the storage module respectively. It is used to construct a discriminant function driven by multimodal data fusion based on the two dynamic deviation indices and the state persistence parameter, calculate the judgment result, and store the judgment result in the storage module.

Citation Information

Patent Citations

  • Rain condition monitoring data intelligent diagnosis method based on spatial relationship

    CN119535644A

  • Blockage detection method applied to rain gauge

    CN119758488A