High-density built-up area automatic rainfall sampling method and device

By constructing a pollution risk prediction model and dynamic sampling strategy in densely built-up areas, and combining geographic information and water quality sensors, the problems of high sampling cost, low efficiency and insufficient data representativeness in existing technologies are solved. This enables accurate capture and efficient monitoring of pollution peaks, providing an efficient and scientific pollution monitoring method.

CN121276015BActive Publication Date: 2026-02-10TIANJIN WATER RESOURCES RES INST +2
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511843061.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-09
Publication Date
2026-02-10
Estimated Expiration
2045-12-09

AI Technical Summary

Technical Problem

Existing technologies for monitoring rainfall pollution in densely built-up areas suffer from problems such as high cost, low efficiency, and significant safety hazards associated with manual sampling, while automated sampling devices struggle to capture pollution peaks and lack data representativeness.

Method used

By collecting geographic information and rainfall dynamic parameters of high-density built-up area plots, a pollution risk prediction model is constructed, a composite dynamic sampling strategy is generated, and real-time analysis is performed using integrated water quality sensors. Multi-factor spatial weighted compensation is also performed to calculate the equivalent pollution concentration.

Benefits of technology

It enables precise capture of pollution peaks, improves data representativeness and sampling efficiency, provides scientific pollution source tracing and load assessment capabilities, and enhances the accuracy and comparability of water sample analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121276015B_ABST
    Figure CN121276015B_ABST
Patent Text Reader

Abstract

The application relates to the field of rainfall pollution detection, in particular to an automatic rainfall sampling method and device for high-density built-up areas. Geographic information parameters and rainfall dynamic parameters of the high-density built-up area are collected. Multi-source data obtained are subjected to format unification, space-time alignment and normalization processing to construct a standardized multi-dimensional feature data set. The standardized multi-dimensional feature data set is input into a pollution risk prediction model to dynamically calculate and output a pollution risk grade of a rainfall event and a time window of a key pollution stage. A composite dynamic sampling strategy is generated based on the pollution risk grade and the time window. According to the composite dynamic sampling strategy, water sample collection and storage are performed, and the collected water sample is preliminarily analyzed by using an integrated water quality sensor to generate measured concentration data. The measured concentration data is subjected to weighted compensation by using the geographic information parameters of the high-density built-up area to calculate equivalent pollution concentration representing the comprehensive pollution level of the high-density built-up area.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the technical field of rainfall pollution detection, and in particular to an automated rainfall sampling method and device for high-density built-up areas. Background Technology

[0002] Within the framework of urban environmental governance and water resource protection, precise monitoring and analysis of non-point source pollution caused by rainfall events are crucial for assessing the health of urban water environments, tracing pollution sources, and developing effective prevention and control measures. Rainfall runoff, as it flows through highly impermeable surfaces in densely built-up areas, washes away and carries various pollutants accumulated during sunny days, forming water bodies with complex compositions and drastically changing concentrations. Therefore, systematically collecting and analyzing such water samples to determine their pollutant components, concentrations, and other chemical and physical properties is of paramount importance for environmental science research and municipal management.

[0003] Currently, the existing technologies for obtaining rainfall runoff samples mainly include two types: manual on-site sampling and conventional automated sampling. Manual sampling relies entirely on staff going to the site during rainfall to manually collect samples; the drawbacks of this method are obvious: firstly, it is costly and inefficient; secondly, rainfall events are random and sudden, and manual responses are often delayed, making it difficult to capture the complete dynamic changes of the pollution process; thirdly, outdoor operations in inclement weather pose safety hazards, and the subjectivity of manual operation can easily lead to insufficient sample representativeness, affecting the accuracy and comparability of subsequent analytical results.

[0004] Conventional automated sampling devices have overcome some of the shortcomings of manual sampling to a certain extent. They typically use fixed time intervals or simple flow thresholds to trigger sampling. However, when applied to the complex scenario of high-density built-up areas, their technical limitations remain significant. Urban rainfall runoff pollution exhibits a significant "first-strike effect," meaning that within the first few minutes to tens of minutes of rainfall, the pollutant concentration in the runoff reaches its peak, carrying the majority of the pollution load. Existing automated devices, based on passive and fixed sampling strategies, are highly susceptible to missing this brief but crucial pollution peak phase, leading to a severe underestimation of the overall pollution event. Furthermore, the underlying surface properties of high-density built-up areas are extremely complex, with different land use types (such as main roads, commercial rooftops, and residential green spaces) contributing vastly different levels of pollution to runoff. Current technologies, when analyzing water samples obtained from a single outfall, only reflect the mixed water quality at that cross-section and cannot characterize the spatial heterogeneity of the entire catchment area, resulting in severely insufficient data representativeness.

[0005] Therefore, there is an urgent need for an automated rainfall sampling method and device for high-density built-up areas. Summary of the Invention

[0006] To address the aforementioned technical problems, this invention provides an automated rainfall sampling method for high-density built-up areas, comprising the following steps:

[0007] S1. Collect geographic information parameters that characterize the static spatial heterogeneity of various plot units within high-density built-up areas, as well as rainfall dynamic parameters that characterize environmental dynamic changes;

[0008] S2. Perform format unification, spatiotemporal alignment and normalization on the acquired multi-source data to construct a standardized multidimensional feature dataset;

[0009] S3. Input the standardized multidimensional feature dataset into the pollution risk prediction model pre-trained based on historical monitoring data, dynamically calculate and output the pollution risk level of the rainfall event and the time window of the key pollution stage; and generate a composite dynamic sampling strategy based on the pollution risk level and time window.

[0010] S4. Based on the composite dynamic sampling strategy, water samples are collected and stored from each plot unit, and the collected water samples are preliminarily analyzed using integrated water quality sensors to generate measured concentration data.

[0011] S5. Using the geographic information parameters of various plot units within the high-density built-up area, the generated measured concentration data is weighted and compensated to calculate the equivalent pollution concentration representing the comprehensive pollution of the high-density built-up area.

[0012] The process of acquiring the geographic information parameters includes: using a geographic information system platform and applying remote sensing image interpretation and spatial analysis technology to accurately divide the high-density built-up area into multiple plot units with clear functional attributes; and collecting data for each plot unit to obtain the corresponding geographic information parameters; the geographic information parameters include at least: land use type, surface runoff data of each plot unit, and underground rainwater pipe network topology.

[0013] The land use types include commercial areas, main traffic artery areas, high-density residential areas, industrial and warehousing areas, and urban green spaces; the surface runoff data includes the total rainfall flow within the plot unit area and the rainwater inflow into the underground rainwater pipe network, which is identified and quantified through rainwater runoff tracking technology; the underground rainwater pipe network topology includes the length, slope, runoff relationship, and estimated transmission delay time of the underground rainwater pipe network.

[0014] The rainfall dynamic parameters include: forecasts of future short-term rainfall intensity and duration based on meteorological radar, cumulative duration of sunny days before rainfall, and traffic flow data of main roads in the region reflecting the intensity of activity.

[0015] The pollution risk prediction model uses historical monitoring data, which includes multi-source data of historical rainfall events and corresponding actual pollutant concentration monitoring results, as training samples for offline training.

[0016] In practical applications, the underground rainwater pipe network topology, the cumulative duration of sunny days before rainfall, the predicted peak rainfall intensity, the predicted rainfall duration, and the average traffic flow within a preset time window before rainfall are used as input features. Through an internal nonlinear mapping relationship, a quantitatively classified pollution risk level is output. At the same time, based on the estimated transmission delay time of the underground rainwater pipe network, the time window for the occurrence of the first-impact pollution peak is predicted as the time window for the critical pollution stage.

[0017] The weighted compensation of the measured concentration data is specifically as follows: based on the multi-factor spatial weighted compensation formula, the measured concentration value of the plot unit is transformed into an equivalent pollution concentration that can represent the entire high-density built-up area;

[0018] The compensation formula based on multi-factor spatial weighting is as follows:

[0019] ;

[0020] in, The equivalent pollution concentration obtained after compensation is an indicator used to evaluate the overall pollution load; This represents the measured concentration data of pollutants for the i-th plot unit; Runoff contribution weighting factor for the i-th plot unit; This represents the total number of land parcels within the catchment area.

[0021] The runoff contribution weighting factor is calculated based on surface runoff data and is used to measure the proportion of rainwater runoff that flows into the underground rainwater pipe network in the total rainfall generated by the plot unit.

[0022] An automated rainfall sampling device for high-density built-up areas includes:

[0023] The multi-source data acquisition module collects geographic information parameters that characterize the static spatial heterogeneity of various plot units within a high-density built-up area, as well as rainfall dynamic parameters that characterize dynamic environmental changes.

[0024] The data preprocessing module performs format unification, spatiotemporal alignment, and normalization on the acquired multi-source data to construct a standardized multidimensional feature dataset.

[0025] The sampling strategy generation module inputs the standardized multidimensional feature dataset into a pollution risk prediction model pre-trained based on historical monitoring data, dynamically calculates and outputs the pollution risk level of rainfall events and the time window of key pollution stages; and generates a composite dynamic sampling strategy based on the pollution risk level and time window.

[0026] The water quality measurement module, based on a composite dynamic sampling strategy, performs water sample collection and storage for each plot unit, and uses integrated water quality sensors to perform preliminary analysis on the collected water samples to generate measured concentration data.

[0027] The weighted compensation module uses geographic information parameters of high-density built-up areas to perform weighted compensation on measured concentration data, and calculates the equivalent pollution concentration that represents the comprehensive pollution level of high-density built-up areas.

[0028] The present invention has the following technical effects:

[0029] 1. This invention provides a solid spatial data foundation for the entire technical solution by refining and quantifying the static characteristics of high-density built-up areas; it transforms the vague concept of high-density built-up areas into a digital model defined by specific parameters and capable of scientific calculation by computers; it greatly improves the accuracy and reliability of subsequent risk prediction and data compensation, making the entire method no longer based on rough estimation, but on the precise characterization of the physical entities of the catchment area; it creates the prerequisites for achieving accurate and scientific pollution source tracing and load assessment, fundamentally enhancing the scientific nature and operability of the invention.

[0030] 2. This invention, through a pollution risk prediction model, can anticipate pollution risks and accurately locate key monitoring periods, which gives the subsequently generated dynamic sampling strategy unprecedented accuracy and foresight. Its most significant benefit is that it can ensure the capture of the fleeting but extremely high pollution load "first-strike effect" with the highest probability and time accuracy, solving the problem of key sample loss that has long existed in traditional sampling methods and greatly improving the temporal representativeness of water samples.

[0031] 3. This invention creatively solves the most challenging problem of "data representativeness" in urban non-point source pollution monitoring; it provides a complete, operable, and scientific post-processing method that can upgrade the "point" data of physical sampling points into macroscopic indicators that can scientifically characterize the "area" state of the entire catchment area; the equivalent pollution concentration it generates has far higher decision-making value than the original measured concentration data, and can be used more accurately to assess the overall water environment quality of the region, identify key pollution contributing areas, and measure the macroscopic effectiveness of pollution control measures, so that the final output of this invention achieves a qualitative leap in scientificity and practicality. Attached Figure Description

[0032] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0033] Figure 1 This is a flowchart of an automated rainfall sampling method for high-density built-up areas according to the present invention;

[0034] Figure 2 This is a schematic diagram of the logic for obtaining the equivalent pollution concentration of the present invention;

[0035] Figure 3 This is a structural diagram of an automated rainfall sampling device for high-density built-up areas according to the present invention. Detailed Implementation

[0036] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below. Obviously, the described embodiments are only a part of the embodiments of this invention, and not all of them. Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without creative effort are part of this invention.

[0037] Example 1:

[0038] This invention proposes an automated rainfall sampling method for high-density built-up areas, the process of which is as follows: Figure 1 As shown, it includes:

[0039] S1. Collect geographic information parameters that characterize the static spatial heterogeneity of various plot units within high-density built-up areas, as well as rainfall dynamic parameters that characterize environmental dynamic changes;

[0040] S2. Perform format unification, spatiotemporal alignment and normalization on the acquired multi-source data to construct a standardized multidimensional feature dataset;

[0041] S3. Input the standardized multidimensional feature dataset into the pollution risk prediction model pre-trained based on historical monitoring data, dynamically calculate and output the pollution risk level of the rainfall event and the time window of the key pollution stage; and generate a composite dynamic sampling strategy based on the pollution risk level and time window.

[0042] S4. Based on the composite dynamic sampling strategy, water samples are collected and stored from each plot unit, and the collected water samples are preliminarily analyzed using integrated water quality sensors to generate measured concentration data.

[0043] S5. Using the geographic information parameters of various plot units within the high-density built-up area, the generated measured concentration data is weighted and compensated to calculate the equivalent pollution concentration representing the comprehensive pollution of the high-density built-up area.

[0044] Furthermore, the process of obtaining the geographic information parameters includes:

[0045] Using a geographic information system platform and employing remote sensing image interpretation and spatial analysis techniques, high-density built-up areas are precisely divided into multiple plot units with distinct functional attributes. Data is collected for each plot unit to obtain corresponding geographic information parameters. These geographic information parameters include at least: land use type, surface runoff data for each plot unit, and underground rainwater pipe network topology.

[0046] The land use types include commercial areas, main traffic artery areas, high-density residential areas, industrial and warehousing areas, and urban green spaces; the surface runoff data includes the total rainfall flow within the plot unit area and the rainwater inflow into the underground rainwater pipe network, which is identified and quantified through rainwater runoff tracking technology; the underground rainwater pipe network topology includes the length, slope, runoff relationship, and estimated transmission delay time of the underground rainwater pipe network.

[0047] Specifically, the underground pipe network map, high-resolution satellite remote sensing imagery, and land use status map of the catchment area or pump outlet to be monitored are analyzed. Based on a digital elevation model, the catchment area boundary upstream of the sampling point is automatically or semi-automatically and accurately defined. Subsequently, through remote sensing image processing techniques such as supervised classification or visual interpretation, the surface cover within the catchment area is finely divided into multiple polygonal plot units, and each unit is assigned a clear land use type attribute, such as "commercial area," "traffic arterial area," or "high-density residential area." For each plot, rainwater runoff tracking yields quantified surface runoff data. Finally, the underground pipe network map is converted into topology network data, and information such as the material and slope of each pipe segment is input to construct an underground pipe network model capable of hydraulic calculations.

[0048] This invention provides a solid spatial data foundation for technical solutions by refining and quantifying the static characteristics of high-density built-up areas. It transforms the vague concept of high-density built-up areas into a digital model defined by specific parameters that can be scientifically calculated by computers. This greatly improves the accuracy and reliability of subsequent risk prediction and data compensation, making the entire method no longer based on rough estimations but on the precise characterization of regional physical entities. This creates the prerequisites for achieving accurate and scientific pollution source tracing and load assessment, fundamentally enhancing the scientific nature and operability of the invention.

[0049] The rainfall dynamic parameters include: forecasts of future short-term rainfall intensity and duration based on meteorological radar, cumulative duration of sunny days before rainfall, and traffic flow data of main roads in the region reflecting the intensity of activity.

[0050] Specifically, the system periodically calls the city meteorological bureau's nowcasting service via an application programming interface (API) to obtain radar reflectivity data for the area where the sampling point is located for the next 1-3 hours, and then parses it into minute-level rainfall intensity forecasts. Simultaneously, it queries data released by the city traffic information center in real time through another API to obtain real-time traffic flow data from traffic flow monitoring loops on several pre-set main roads within the area. Furthermore, it continuously records the cumulative dry season since the last effective rainfall (e.g., cumulative rainfall exceeding 2 mm) as a parameter for cumulative sunny day duration.

[0051] The pollution level of rainfall runoff is not static but closely related to three core dynamic factors: 1. Pollutant accumulation: The longer the duration of dry weather, the more dry deposition, tire wear, and oil accumulate on the surface (especially along major traffic arteries). 2. Washing energy: The greater the rainfall intensity, the stronger the washing capacity of accumulated pollutants on the surface, and the higher the pollutant concentration in the initial runoff (i.e., the "first flush effect"). 3. Real-time intensity of pollution sources: Traffic flow is directly related to the real-time emission and deposition rate of pollutants such as vehicle exhaust and brake pad wear.

[0052] This solution captures these dynamic parameters in real time, providing the most timely input variables for subsequent pollution risk prediction, thus giving the invention both foresight and real-time capability. It elevates the basis for sampling decisions from a passive signal of "rainfall has occurred" to an active and intelligent level of "based on the predicted rainfall process and the current state of pollution sources."

[0053] By acquiring these dynamic parameters in real time, we can gain insight into the unique pollution potential of each rainfall event, such as distinguishing the different environmental risks posed by "strong thunderstorms after a prolonged drought" and "continuous light rain." This makes subsequent sampling strategies highly targeted, maximizing the optimization of sampling resources, ensuring accurate capture of key pollution events, and significantly improving the efficiency and scientific rigor of sampling.

[0054] The pollution risk prediction model uses historical monitoring data, which includes multi-source data of historical rainfall events and corresponding actual pollutant concentration monitoring results, as training samples for offline training.

[0055] In practical applications, the underground rainwater pipe network topology, the cumulative duration of sunny days before rainfall, the predicted peak rainfall intensity, the predicted rainfall duration, and the average traffic flow within a preset time window before rainfall are used as input features. Through an internal nonlinear mapping relationship, a quantitatively classified pollution risk level is output. At the same time, based on the estimated transmission delay time of the underground rainwater pipe network, the time window for the occurrence of the first-impact pollution peak is predicted as the time window for the critical pollution stage.

[0056] In the pollution risk prediction model construction phase, historical data from the past few years are collected. Each data sample corresponds to a complete rainfall event, including dynamic parameters of the event (underground stormwater drainage network topology, duration of sunny days, peak rainfall intensity, and average traffic flow) as input features to the model, and the risk level corresponding to the measured peak pollutant concentration of the event as the output label (e.g., the top 20% of concentrations are defined as "high risk"). Using this data, a gradient boosting decision tree model is trained through a machine learning platform (such as TensorFlow or Scikit-learn). After training, this lightweight model is deployed in the embedded controller of the sampling device. During actual operation, the controller inputs the dynamically acquired parameters into the model in real time, and the model outputs a specific risk level (e.g., "high," "medium," or "low") and a predicted "first-strike effect" time window (e.g., "5-15 minutes after the start of rainfall") within milliseconds.

[0057] When the pollution risk level output by the model is "high", the generated composite dynamic sampling strategy is specifically a high-frequency dynamic first-rush capture strategy: the remote terminal sends a preparatory instruction to the on-site sampling device 30 to 60 minutes in advance based on the predicted rainfall start time, enabling the device to complete system self-check and enter standby mode; once the device's pump start sensor detects the formation of initial runoff, the sampling system immediately starts and enters high-frequency sampling mode, continuously collecting water samples at extremely short time intervals (e.g., every 1 to 3 minutes); at the same time, the device's built-in high-precision turbidity or conductivity sensor monitors the rapid changes in water quality in real time at a second-level frequency and feeds the data back to the controller; the controller continuously calculates the rate of change of water quality parameters, and when the rate of change changes from a peak positive value to a negative value and remains below a preset threshold for a period of time, it determines the end of the "first-rush effect" pollution peak stage based on the time window of the critical pollution stage, and then automatically reduces the sampling frequency to a medium-risk level, so as to achieve accurate capture of critical pollution processes and optimize sample storage resources.

[0058] If the risk level is "medium", it applies to regular showers or continuous rainfall processes. After runoff is detected, sampling will be performed at a fixed, relatively short time interval (e.g., every 15 to 30 minutes) to record water quality fluctuations evenly throughout the rainfall process and ensure the continuity and integrity of the data.

[0059] If the risk level is "low", it is suitable for continuous light rain or the end of the rainfall. In order to save equipment energy consumption and sample storage space, a significantly extended sampling time interval (e.g., every 60 to 120 minutes) will be adopted. Its main purpose is to monitor the water quality background value after the rainfall and confirm the dissipation of the pollution process, so as to achieve optimal resource allocation and scientific sampling under different risk scenarios.

[0060] Rainfall pollution processes are influenced by multiple factors and are difficult to describe with simple linear formulas. Machine learning algorithms can learn from historical data and uncover hidden, complex nonlinear patterns. Specifically, advanced machine learning models such as GBDT can automatically learn the complex interactions between these factors; for example, the risk only increases sharply when both sunny duration and rainfall intensity are high simultaneously. This allows for the construction of a high-precision predictor that can make intelligent judgments about future pollution risks based on data-driven scientific evidence, rather than simple thresholds or rules.

[0061] This invention, through a pollution risk prediction model, can anticipate pollution risks and accurately pinpoint key monitoring periods. This enables the subsequent dynamic sampling strategy to possess unprecedented accuracy and foresight. Its most significant benefit is ensuring the capture of the fleeting but highly polluting "first-strike effect" with the highest probability and temporal precision, solving the long-standing problem of key sample loss in traditional sampling methods and greatly improving the temporal representativeness of water samples.

[0062] The weighted compensation of the measured concentration data is specifically as follows: based on the multi-factor spatial weighted compensation formula, the measured concentration value of the plot unit is transformed into an equivalent pollution concentration that can represent the entire high-density built-up area;

[0063] The compensation formula based on multi-factor spatial weighting is as follows:

[0064] ;

[0065] in, The equivalent pollution concentration obtained after compensation is an indicator used to evaluate the overall pollution load; This represents the measured concentration data of pollutants for the i-th plot unit; Runoff contribution weighting factor for the i-th plot unit; This represents the total number of land parcels within the catchment area.

[0066] The runoff contribution weighting factor is calculated based on surface runoff data and is used to measure the proportion of rainwater runoff that flows into the underground rainwater pipe network in the total rainfall generated by the plot unit.

[0067] The role of the runoff contribution weighting factor is as follows: for a plot of land, when all the rainwater runoff generated by rainfall flows into the underground rainwater pipe network, the measured concentration data at this time can most accurately reflect the pollution level of the plot of land. However, in reality, due to evaporation, surface infiltration, and other factors, the rainwater flow into the underground rainwater pipe network is inevitably less than the rainwater runoff generated by rainfall. In this case, the measured concentration data deviates from the true pollution level and cannot accurately reflect the pollution level of the plot of land, thus reducing the reliability of the data. Furthermore, the larger the proportion of rainwater inflow, the closer the rainwater flow into the underground rainwater pipe network is to the rainwater runoff generated by rainfall, and the higher the reliability of the measured concentration data.

[0068] Therefore, by weighting the measured concentration data of each plot unit according to the runoff contribution weighting factor, the overall pollution situation of high-density built-up areas can be accurately obtained.

[0069] This invention creatively solves the most challenging problem of data representativeness in urban non-point source pollution monitoring, providing a complete, operable, and scientific post-processing method that can elevate the "point" data of a physical sampling point into a macroscopic indicator that can scientifically characterize the "area" state of the entire catchment area. The "equivalent pollution concentration" it generates has far greater decision-making value than the original measured concentration data, and can be used more accurately to assess the overall water environment quality of the region, identify key pollution contributing areas, and measure the macroscopic effectiveness of pollution control measures. This results in a qualitative leap in the scientific rigor and practicality of the final output of this invention.

[0070] Example 2:

[0071] This invention also proposes an automated rainfall sampling device for high-density built-up areas, the structure of which is as follows: Figure 3 As shown, it includes:

[0072] The multi-source data acquisition module collects geographic information parameters that characterize the static spatial heterogeneity of various plot units within a high-density built-up area, as well as rainfall dynamic parameters that characterize dynamic environmental changes.

[0073] This module is used to sense the external environment and obtain decision-making basis. Its function is to collect geographic information parameters that characterize the static spatial heterogeneity of high-density built-up areas in real time, as well as rainfall dynamic parameters that characterize environmental dynamic changes. In terms of hardware implementation, this module consists of a highly integrated multi-functional communication and positioning board, specifically including:

[0074] Remote Dynamic Data Acquisition Unit: At its core is a full-network compatible cellular communication module supporting multiple network standards including 4G, 5G, and NB-IoT. Equipped with a high-gain external antenna, it ensures a stable and reliable network connection even in complex urban canyon environments. This module proactively accesses the city's public meteorological data server periodically (e.g., polling every 5 minutes) via a pre-defined application programming interface (API) to obtain radar echo extrapolation rainfall intensity prediction data for the next 1 to 3 hours, covering the geographical grid of the sampling point. Simultaneously, it queries the city's intelligent traffic management platform in real time via another API to obtain real-time traffic flow data from traffic monitoring loops on pre-defined key traffic arteries within the catchment area.

[0075] Local dynamic data generation unit: The device's CPU integrates a high-precision real-time clock, continuously powered by a backup battery. This unit is responsible for continuously recording the cumulative rainless time since the last effective rainfall (the criteria for which is determined, for example, cumulative rainfall exceeding 2 mm, which can be preset by the user), and updating it in real time as a parameter of the cumulative sunny time before rainfall.

[0076] Static Parameter Storage and Positioning Unit: This module contains a large-capacity local Flash memory. Before device deployment, a geographic information parameter database of the sampling point's catchment area (including land use types, surface runoff data, and underground stormwater pipe network topology of each discrete plot unit, generated through a GIS platform) is pre-stored here. The board also integrates a high-precision GPS / BeiDou dual-mode positioning receiver to accurately acquire its geographic coordinates during equipment installation and maintenance, and to provide accurate and unified timestamps for all collected data, ensuring subsequent spatiotemporal alignment of the data.

[0077] The data preprocessing module performs format unification, spatiotemporal alignment, and normalization on the acquired multi-source data to construct a standardized multidimensional feature dataset.

[0078] This module, as a core software algorithm, is embedded and runs in the central processing unit. Its function is to perform format unification, spatiotemporal alignment, and normalization on the acquired multi-source data to construct a standardized multidimensional feature dataset. Its specific workflow is as follows:

[0079] Unified format: Data from different API interfaces (such as meteorological data in JSON format and traffic data in XML format) are parsed into a unified internal data structure.

[0080] Spatiotemporal alignment: Using high-precision timestamps provided by GPS, data from different sources and with different sampling frequencies are aligned to a unified timeline (e.g., unified into minute-level time series), and interpolation or resampling is performed to fill in missing values.

[0081] Normalization: Feature data with different dimensions and numerical ranges (such as rainfall intensity mm / h, traffic flow veh / h, and sunny day duration h) are uniformly scaled to the interval [0,1] using mathematical methods such as min-max scaling. This step is crucial for eliminating numerical differences between different features and improving the training efficiency and prediction accuracy of subsequent machine learning models.

[0082] The sampling strategy generation module inputs the standardized multidimensional feature dataset into a pollution risk prediction model pre-trained based on historical monitoring data, dynamically calculates and outputs the pollution risk level of rainfall events and the time window of key pollution stages; and generates a composite dynamic sampling strategy based on the pollution risk level and time window.

[0083] This module serves as the device's decision-making brain, also running as software within the central processing unit. Its core task is to receive preprocessed data and intelligently generate the optimal sampling strategy.

[0084] Pollution Risk Prediction Model: The core of this module is a Gradient Boosting Decision Tree (GBDT) model that has been trained offline using historical monitoring data. This model takes standardized cumulative sunny days, predicted peak rainfall intensity, predicted rainfall duration, and average traffic flow as input features. It can perform forward inference calculations within milliseconds and output two key results: a quantified, categorical pollution risk level (e.g., high, medium, and low), and a predicted time window for the most likely occurrence of the "first-strike effect" pollution peak (e.g., "5-15 minutes after the start of rainfall").

[0085] The logic for generating a composite dynamic sampling strategy is as follows: The module pre-defines a detailed rule base to transform the model's prediction results into specific, executable sequences of hardware control instructions. This rule base is crucial for realizing the dynamic sampling concept of this invention, and its logic is as follows:

[0086] If the predicted risk level is "high": generate a "high-frequency dynamic first-flow capture strategy" instruction set. This instruction set includes: (1) immediately sending a preparatory instruction to the water quality measurement module and completing the self-check 30 minutes in advance; (2) after detecting the initial runoff, controlling the peristaltic pump and distribution valve to sample at an extremely high frequency of 2 minutes / time; (3) simultaneously reading the turbidity / conductivity sensor values ​​at a second-level frequency and calculating their rate of change in real time; (4) when the rate of change turns from positive to negative and the peak value is determined in combination with the prediction time window, automatically modifying the sampling frequency control instruction to 30 minutes / time.

[0087] If the predicted risk level is "medium": generate a "routine monitoring strategy" instruction set, which specifies that sampling should be performed at a fixed frequency of 15 minutes / time after the start of runoff.

[0088] If the predicted risk level is "low": Generate an "energy-saving monitoring strategy" instruction set, which specifies that sampling should be performed at long intervals of 60 minutes to save energy consumption and sample storage space.

[0089] The water quality measurement module, based on a composite dynamic sampling strategy, performs water sample collection and storage for each plot unit, and uses integrated water quality sensors to perform preliminary analysis on the collected water samples to generate measured concentration data.

[0090] This module is the core actuator of the device, responsible for the physical collection, storage, and preliminary analysis of water samples. Its internal structure is sophisticated, consisting of the following cooperating units:

[0091] Self-cleaning water inlet unit: The sampling tube's inlet employs a dual-layer filtration structure (outer coarse screen + inner fine screen) to block debris. A miniature pressure sensor is connected in series in the pipeline. The central processing unit monitors the pipeline pressure in real time. Once the pressure abnormally increases (indicating blockage), it automatically triggers the self-cleaning program: pausing sampling and precisely controlling the peristaltic pump to reverse for 3-5 seconds, forcefully flushing the filter screen with the clean water already drawn into the tube, thus achieving proactive anti-clogging.

[0092] Precision Sampling and Sensing Unit: A high-precision, reversible industrial-grade peristaltic pump serves as the power source, with its flow rate precisely calibrated to ensure consistent volume for each sample. After being drawn in, the water sample first flows through a specially designed "flow-through multi-parameter water quality analysis cell." This cell integrates a high-precision turbidity sensor, a four-electrode conductivity sensor, and an ultraviolet spectral sensor for rapid COD (Chemical Oxygen Demand) determination. The sensors complete the measurement the instant the water sample flows through, generating measured concentration data.

[0093] Sample dispensing and storage unit: After being measured in the analytical cell, the water sample flows into a multi-channel rotary dispensing valve driven by a high-precision stepper motor. Below the valve is a sample tray that can hold 24 500mL polyethylene standard sample vials. The central processing unit, according to the sampling strategy, synchronously controls the port of the rotary valve and the position of the sample tray by sending precise pulse signals, ensuring that each water sample is accurately dispensed into a designated, uniquely timestamped sample vial.

[0094] The weighted compensation module uses geographic information parameters of high-density built-up areas to perform weighted compensation on measured concentration data, and calculates the equivalent pollution concentration that represents the comprehensive pollution level of high-density built-up areas.

[0095] Data Retrieval and Calculation: This module is automatically triggered after a complete rainfall sampling event. It retrieves all measured concentration data sequences recorded for this event from local storage and retrieves the geographic information parameters of the plot units associated with the sampling point from the static database. Then, using a multi-factor spatial weighted compensation formula, it performs calculations to ultimately determine the equivalent pollution concentration that represents the overall pollution level of the entire catchment area.

[0096] This final, scientifically enhanced analysis result will be packaged together with the original measured data sequence, sampling strategy, meteorological and traffic data to form a complete event report, which will then be reported to the user's data management cloud platform via the communication module. Through the specific and close collaboration of the aforementioned modules, the reliability, advancement, and feasibility of the entire invention are ensured, enabling it to efficiently and accurately complete the fully automated rainfall sampling task proposed in this invention, from intelligent prediction to scientific assessment.

[0097] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the technical solutions of the embodiments of the present invention.

Claims

1. An automated rainfall sampling method for high-density built-up areas, characterized in that, Includes the following steps: S1. Collect geographic information parameters that characterize the static spatial heterogeneity of various plot units within high-density built-up areas, as well as rainfall dynamic parameters that characterize environmental dynamic changes; S2. Perform format unification, spatiotemporal alignment and normalization on the acquired multi-source data to construct a standardized multidimensional feature dataset; S3. Input the standardized multidimensional feature dataset into the pollution risk prediction model pre-trained based on historical monitoring data, dynamically calculate and output the pollution risk level of the rainfall event and the time window of the key pollution stage; A composite dynamic sampling strategy is generated based on the pollution risk level and time window. The construction process of the pollution risk prediction model includes: using historical monitoring data containing multi-source data of historical rainfall events and corresponding actual pollutant concentration monitoring results as training samples for offline training; In practical applications, the underground rainwater pipe network topology, the cumulative duration of sunny days before rainfall, the predicted peak rainfall intensity, the predicted rainfall duration, and the average traffic flow within a preset time window before rainfall are used as input features. Through an internal nonlinear mapping relationship, a quantitatively classified pollution risk level is output. At the same time, based on the estimated transmission delay time of the underground rainwater pipe network, the time window for the occurrence of the first-impact pollution peak is predicted as the time window for the critical pollution stage. In actual operation, the controller inputs the dynamic parameters acquired in real time into the model, and the model outputs a specific risk level, including high, medium, and low, as well as a predicted first-impact effect time window within milliseconds. When the pollution risk level output by the model is high, the generated composite dynamic sampling strategy is specifically a high-frequency dynamic first-rush capture strategy: the remote terminal sends a preparatory command to the on-site sampling device 30 to 60 minutes in advance according to the predicted rainfall start time, so that the device can complete the system self-check and enter the standby state; once the pump sensor of the device detects the formation of the initial runoff, the sampling system starts immediately and enters the high-frequency sampling mode to continuously collect water samples at extremely short time intervals; at the same time, the device's built-in high-precision turbidity or conductivity sensor monitors the rapid changes in water quality in real time at a frequency of seconds and feeds the data back to the controller; the controller continuously calculates the rate of change of water quality parameters, and when the rate of change changes from a peak positive value to a negative value and remains below the preset threshold for a period of time, the first-rush effect pollution peak stage is determined to have ended in combination with the time window of the key pollution stage, and the sampling frequency is automatically reduced to a medium-risk level, so as to achieve accurate capture of key pollution processes and optimize sample storage resources; S4. Based on the composite dynamic sampling strategy, water samples are collected and stored from each plot unit, and the collected water samples are preliminarily analyzed using integrated water quality sensors to generate measured concentration data. S5. Using the geographic information parameters of various plot units within the high-density built-up area, the generated measured concentration data is weighted and compensated to calculate the equivalent pollution concentration representing the comprehensive pollution of the high-density built-up area.

2. The automated rainfall sampling method for high-density built-up areas according to claim 1, characterized in that, The process of obtaining the geographic information parameters includes: Using a geographic information system platform and employing remote sensing image interpretation and spatial analysis techniques, high-density built-up areas are precisely divided into multiple plot units with distinct functional attributes. Data is collected for each plot unit to obtain corresponding geographic information parameters. These geographic information parameters include at least: land use type, surface runoff data for each plot unit, and underground rainwater pipe network topology.

3. The automated rainfall sampling method for high-density built-up areas according to claim 2, characterized in that: The land use types include commercial areas, main traffic artery areas, high-density residential areas, industrial and warehousing areas, and urban green spaces; the surface runoff data includes the total rainfall flow within the plot unit area and the rainwater inflow into the underground rainwater pipe network, which is identified and quantified through rainwater runoff tracking technology; the underground rainwater pipe network topology includes the length, slope, runoff relationship, and estimated transmission delay time of the underground rainwater pipe network.

4. The automated rainfall sampling method for high-density built-up areas according to claim 1, characterized in that, The rainfall dynamic parameters include: forecasts of future short-term rainfall intensity and duration based on meteorological radar, cumulative duration of sunny days before rainfall, and traffic flow data of main roads in the region reflecting the intensity of activity.

5. The automated rainfall sampling method for high-density built-up areas according to claim 1, characterized in that, The measured concentration data is weighted and compensated, specifically by converting the measured concentration values ​​of the plot unit into equivalent pollution concentrations that can represent the entire high-density built-up area based on a multi-factor spatial weighting compensation formula. The compensation formula based on multi-factor spatial weighting is as follows: ; in, The equivalent pollution concentration obtained after compensation is an indicator used to evaluate the overall pollution load; This represents the measured concentration data of pollutants for the i-th plot unit; Runoff contribution weighting factor for the i-th plot unit; This represents the total number of land parcels within the catchment area.

6. The automated rainfall sampling method for high-density built-up areas according to claim 5, characterized in that, The runoff contribution weighting factor is calculated based on surface runoff data and is used to measure the proportion of rainwater runoff that flows into the underground rainwater pipe network in the total rainfall generated by the plot unit.

7. An automated rainfall sampling device for high-density built-up areas, characterized in that, An automated rainfall sampling method for high-density built-up areas according to any one of claims 1 to 6 includes: The multi-source data acquisition module collects geographic information parameters that characterize the static spatial heterogeneity of various plot units within a high-density built-up area, as well as rainfall dynamic parameters that characterize dynamic environmental changes. The data preprocessing module performs format unification, spatiotemporal alignment, and normalization on the acquired multi-source data to construct a standardized multidimensional feature dataset. The sampling strategy generation module inputs the standardized multidimensional feature dataset into a pollution risk prediction model pre-trained based on historical monitoring data, dynamically calculates and outputs the pollution risk level of rainfall events and the time window of key pollution stages; and generates a composite dynamic sampling strategy based on the pollution risk level and time window. The water quality measurement module, based on a composite dynamic sampling strategy, performs water sample collection and storage for each plot unit, and uses integrated water quality sensors to perform preliminary analysis on the collected water samples to generate measured concentration data. The weighted compensation module uses geographic information parameters of high-density built-up areas to perform weighted compensation on measured concentration data, and calculates the equivalent pollution concentration that represents the comprehensive pollution level of high-density built-up areas.

Citation Information

Patent Citations

  • Method for evaluating dynamic output characteristics of urban plot scale non-point source pollution

    CN112163347A

  • Water quality comprehensive regulation and control intelligent model based on water quality hydrodynamic force

    CN119538740A