Attribution of in-stream water quality via monitoring reporting and verification sensor-geospatial fusion networks
The integration of sensor-geospatial fusion networks with machine learning models improves water quality monitoring and treatment by correlating in-field and lab data, optimizing water treatment processes, and promoting green alternatives, addressing inefficiencies in current systems.
Patent Information
- Application Number
- PCT/US2025/014340
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-02-02
- Filing Date
- 2025-02-03
- Publication Date
- 2025-08-07
AI Technical Summary
Current water quality monitoring systems rely on infrequent data collection, limited spatial interpolation, and are poor at incorporating external data collected at inconsistent rates or from environments with sparse sensor deployment, leading to inefficiencies in managing water quality and treatment processes.
Implementing monitoring reporting and verification (MRV) via sensor-geospatial fusion networks that utilize machine learning models to correlate in-field sensor data with lab-based analyses, enabling continuous and real-time water quality monitoring and attribution of water quality changes to various actors, and optimizing water treatment processes through sensor-remote sensing fusion.
Enhances water quality monitoring and treatment efficiency by reducing the need for costly infrastructure upgrades, offering nearly-immediate and long-term emission reductions, and enabling cost-effective, less energy-intensive green alternatives to manmade solutions.
Smart Images

Figure 00000027_0000 
Figure 00000028_0000 
Figure 00000029_0000
Abstract
Description
TITLEATTRIBUTION OF IN-STREAM WATER QUALITY VIA MONITORING REPORTING AND VERIFICATION SENSOR-GEOSPATIAL FUSION NETWORKSCROSS-REFERENCES TO RELATED APPLICATIONS
[0001] The present disclosure claims priority to U.S. Provisional Patent Application 63 / 549, 111 titled “ATTRIBUTION OF IN-STREAM WATER QUALITY VIA MONITORING REPORTING AND VERIFICATION SENSOR-GEOSPATIAL FUSION NETWORKS”, which was fded on 2024-02-02, and which is incorporated herein in its entirety.BACKGROUND
[0002] In recent years, telemetry-connected electronic sensors have been developed and applied within water service programs to perform objective and continuous site-level monitoring for various usages and functionalities. These sensors can be used for the monitoring, reporting, and verification (MRV) of various environmental management goals, such as the generation of carbon credits, management of land and water resources, and control of water treatment processes. A digital MRV system may facilitate project design, automated monitoring, control, and data assimilation, robust verification, and data visualization; however, current systems rely on infrequent collection of data, limited spatial interpolation, and are generally poor at incorporating external data collected at inconsistent rates or from environments with sparse sensor deployment.SUMMARY
[0003] The present disclosure provides drinking water treatment via monitoring reporting and verification sensor-geospatial fusion networks.
[0004] Additional features and advantages of the disclosed method and apparatus are described in, and will be apparent from, the following Detailed Description and the Figures. The features and advantages described herein are not all-inclusive and, in particular, many additional features and advantages will be apparent to one of ordinary skill in the art in view of the figures and description. Moreover, it should be noted that the language used in the specification has been principally selected for readability and instructional purposes, and not to limit the scope of the inventive subject matter.BRIEF DESCRIPTION OF THE DRAWINGS
[0005] Figure 1 illustrates an example environment, such as a watershed, in which embodiments of the present disclosure may be practiced.
[0006] Figure 2 illustrates an example functional model for tracking water quality and attributing changes in water quality to various actors, according to embodiments of the present disclosure.
[0007] Figure 3 illustrates an example system, as may use the described models for monitoring, reporting, and validation with sensor-remote sensing fusion, according to embodiments of the present disclosure.
[0008] Figure 4 is a flowchart of an example method for attributing contamination of a water source to a land-based contamination source, according to embodiments of the present disclosure.
[0009] Figure 5 illustrates a block diagram of an example system for attributing contamination of a water source to a land-based contamination source, according to embodiments of the present disclosure.
[0010] Figure 6 is a flowchart of an example method for analyzing data in a contamination source detection system, according to embodiments of the present disclosure.
[0011] Figure 7 illustrates an example computing device, according to embodiments of the present disclosure.DETAILED DESCRIPTION
[0012] The present disclosure provides for monitoring, reporting, and validation (MRV) with sensor-remote sensing fusion for drinking water treatment via monitoring reporting and verification sensor-geospatial fusion networks, in which various remote sensors report
[0013] The climate impacts of conventional water and wastewater treatment (which, in general, are proportional to influent water quality, in-stream water quality, water demand, and carbon intensity of the local electric grid) are generally approached as manmade problems requiring manmade solutions, with nature based solutions poised as a lesser (if even considered) alternative as the effect of the manmade solutions are more easily monitored, despite havingpotential lower impact and requiring more costly inputs than managing the natural environment. By improving influent and in-body water quality, there can be both nearly-immediate and longterm (over 30 years) avoidances of emissions by reducing the need for further upgrades to gray infrastructure. By using the improved MRV techniques describe herein, operators are given tools to find and quantify the effects of green alternatives to manmade solutions, which may be less expensive, less energy intensive, and less carbon intensive, among various other incentives (such as carbon credits).
[0014] As discussed herein, water quality may refer to various parameters that are judged individually, in aggregate, or in direct correlation with one another. The present disclosure contemplates that various localities and professional organizations shall be understood to define various goals or levels of water quality that one of ordinary skill in the art is expected to be familiar with. The individual parameters may include chemical measurements, such as the presence, absence, or concentration of various chemicals or classes of chemicals in a given amount of water, which may directly indicate the presence of a chemical in question (e.g., higher concentration of Calcium ions to indicate a higher concentration of Calcium in the water) or provide indirect evidence of a chemical in question (e.g., higher concentration of Calcium ions to indicate a higher concentration acids in the water that dissolve rocks carrying Calcium). The individual parameters may include physical measurements, such as the amount of water flowing through a given area, a speed of flow of the water, a number of influx and efflux directions of the water, a temperature, a turbulence, or the like. The individual parameters may include various microbial or biological measurements, such as the presence, absences, or concentration of various microbes, fishes, crustaceans, amphibians, plants, algae, fungi, or other water-dwelling aquatic or semi-aquatic life and the markers thereof in and around the water. Additionally, the various microbial or biological measurements and chemical measurements may identify the presence, absence, or concentration of fecal matter or other makers of the effects of non-aquatic life in and around the water (e.g., due to farm or habitation runoff into a waterway) on the water. For avoidance of doubt, these various parameters may collectively be referred to as physio-bio-chemical parameters.
[0015] Figure 1 illustrates an example environment 100, such as a watershed, in which embodiments of the present disclosure may be practiced. As illustrated, various bodies HOa-g (generally or collectively, bodies 110) are illustrated, which may include standing bodies 110 of water (e.g., lakes, reservoirs, dammed streams, aquafers), moving bodies 110 of water (e.g., rivers,streams, aqueducts, canals), and temporary bodies 1 10 (e.g., arroyos, seasonally dry / flooded creeks, retaining or catchment ponds, storm sewers), and include natural bodies 110, purely human-made bodies 110, and enhanced or human-engineered natural bodies 110. Surrounding lands 120a-c (generally or collectively, lands 120) associated with various owners and users may drain into theses bodies 110 due to collected precipitation (e g., rain runoff, snowmelt, etc.). Additionally, various users 130a-c (generally or collectively, users 130) may output water and other effluents to the bodies 110, draw water and other inputs from the bodies 110, or use the water in the bodies 110 for motive force (e.g., hydroelectric power generation), navigation, or as a cooling source. The owners or users of lands 120 that drain into the bodies 110 may also be users 130 of the bodies 110, but are not required to do so. Accordingly, the entities that own or use the lands 120 and the users 130 of the bodies 110 may collectively be referred to herein as actors.
[0016] The quality and quantity of the water in the various bodies 110 may affect various nonhuman organisms living in the bodies 110 (e.g., invertebrates, fish, amphibians, water plants, algae) or using the water therein for habitat (e.g., waterfowl, beavers) or drinking purposes, affect the lands 120 bordering those bodies, and affect the ability of the users 130 to apply the water for various uses. Accordingly, various sensors can be deployed and monitored throughout the environment 100 to identify water quality, and take actions to address various issues related to water quality.
[0017] In various embodiments, the sensors may include computing devices (such as those discussed in greater detail in regard to Figure 7), that collect various data related to the quantity and characteristics of the water in the bodies 110, the land 120 as use thereof, and how the actors previously, currently, and expectedly used / use / will use the water and the surrounding lands. In addition to quantitative values for various features of the water and the land 120 (e.g., rainfall in a given time period, particulate counts (e.g., total organic carbon (TOC)) in a given time period, temperature at a given time, presence of a given biomarker or chemical, locations / thicknesses of vegetative cover, various fluorescence measures, etc.) the sensors may collect qualitative data and survey-reported data (e.g., from the actors). The sensors may, therefore, be deployed to specific portions of the environment 100 for longitudinal data collection, be intermittently present in the environment 100 (e.g., satellites collecting images of the environment 100, research teams deploying or collecting sensors or data at various intervals) to collect snapshots of data.
[0018] As will be appreciated, saturation of the environment 100 with sensors to measure every possible variable affecting water quality continuously and in real-time is not feasible. Accordingly, the environment 100 is modeled by one or more functional models that use the collected data to extrapolate various data that are not directly measured.
[0019] Figure 2 illustrates an example functional model 200 for tracking water quality and attributing changes in water quality to various actors, according to embodiments of the present disclosure. A transfer model 210 receives in-site sensor data 220 from various sensors deployed throughout an environment 100 being monitored for water quality measures, and lab training data 230. The in-site sensor data 220 includes various details collected directly from sensors in the environment 100, while the lab training data 230 include data generated in a laboratory setting from data or sample collected from the environment 100. As will be appreciated, in-field sensors may ordinarily lack certain capabilities that laboratory analysis tools provide, which requires extracting a sample from the environment and performing an “offline” or non-real-time analysis in a different setting to provide those values. For example, conventional in-field sensors may provide real-time data on water and air temperature at various points in the environment 100 using easy to deploy temperature probes. In contrast, performing a conventional population survey of different microbes in a body 110 may require extracting a water sample and performing a statistical analysis for the different strains of microbes visually identified therein, which is either infeasible or impossible to perform in real-time or with in-field systems alone. The present disclosure therefore augments the functionalities of the in-field systems via a transfer model 210 that is trained to correlate in-field sensor data with lab-based analyses (e.g., set as a ground truth or labeled output for a training date set) to model what values the lab-based analyses would produce using as-of-yet uncollected in-field data to thereby avoid or reduce the amount of lab-based analyses needed.
[0020] The transfer model 210 may be one of various types of machine learning (ML) models, which identifies correlations between the various values for the in-site sensor data 220 and the lab training data 230, and extrapolates modelled lab data 235 (i.e., the predictions of quantified water quality parameters modelled from the in-field data that are conventionally determined by lab instruments) from given inputs of in-site sensor data 220. Accordingly, based on training the transfer model 210 using previously collected in-site sensor data 220 and lab training data 230, the transfer model 210 may develop a function that generates a value for modelled lab data 235 basedon one or more values of newly collected in-site sensor data 220. For example, the level of a given pollutant in a body 110 in parts per million (PPM) may require the use of a centrifuge and various tests that are impractical to carry out in real time in the field, and may therefore be collected and calculated in a laboratory. Several collected values of these lab-generated data are used along with data collected from the environment to train the model transfer model 210 so that in-site sensor data 220 can later be used to generate values that approximate the training data within a given confidence threshold to thereby be used to extrapolate, with confidence, values for the modelled lab data 235 as though those values were based on lab analysis.
[0021] In various embodiments, the sensors may provide in-site sensor data 220 that include one or more of turbidity, conductivity, fluorescent dissolved organic matter (fDOM) measurements, chlorophyll a (Chl-A) measurements, temperature, flow rate / water speed, water level, and metadata related to z-score, N-day averages for various values, sensor percentile, days since deployment, days since last cleaning / maintenance, presence indicators for various chemical compounds, etc. In various embodiments, the lab data include turbidity, conductivity, Total Organic Carbon (TOC), Total Nitrogen (TN ), Kjedldahl Nitrogen content, weighted combinations of N and P lab data, presence or quantity indicators for various chemical compounds / biomarkers / species / strains, etc.
[0022] As will be appreciated, some of the developed transfer functions in the transfer model 210 may be specific to one environment 100 or portion of a given environment 100 and not applicable to other environments 100 or other portions of the given environment 100. For example, a first environment 100 located in a cold climate and a second environment 100 located in a warm climate may each produce a transfer function for microbe content based on in-site data 220 for ambient temperature, the presence and type of farms, and precipitation levels, but have different outputs due to whether ambient temperatures preclude the presence of various strains of microbe in the given environment 100, the different crops grown or livestock raised on those farms, and whether the precipitation is snow or rain, among other factors. In another example, a standing body 110 may have different values calculated than a moving body 110 in the same environment 100 based on the same inputs of the in-site data 220 from that environment 100 based on the water flow properties in those different bodies 110 as different portions of the same environment 100.
[0023] A calibration model 250 receives the modelled lab data 235 and (optionally) various in-site sensor data 220 to produce interpolated data 260 matched in space in time to theenvironment 100. The calibration model 250 identifies the time and space where modelled lab data 235 are assigned in the environment 100, and generates modelled sensor data 225 for values measured by “virtual” sensors at locations where the physical sensors are not deployed. For example, using in-site sensor data 220 collected at time tO-tn from various physical sensors, the calibration model 250 can place the modelled lab data 235 for various extrapolated (but otherwise lab-calculated values) at specific coordinates or zones in the environment 100 at various times.
[0024] When determining values for modelled sensor data 225 for virtual sensors, the calibration model 250 may extrapolate a value based on reported values from two or more physical sensors in the environment 100. For example, a virtual temperature sensor “placed” between two physical temperature sensors to spatially interpolate a temperate at a third location in the modeled environment where no physical temperature sensor is located may be expected to report a modelled temperature value between the two physically measured temperature values, or a different value if another environmental feature that affect temperature is identified in the environment (e.g., a heat exchanger from a power plant). Similarly, the calibration model 250 can use the in-site sensor data 220 collected from times to-tnfrom various physical sensors to calculate values measured at times before data collection (e.g., to-x), extrapolated / forecasted after data collection (e.g., tn+x), or temporally interpolated between two or more times of data collected (e.g., at time ti when readings are taken at time to and time t2, but not time ti). When generating forecasted values, the calibration model 250 lags the data features by an equivalent number of time intervals and cross-validates the predictions against observed data (when eventually collected) for retraining and improving the calibration model.
[0025] In various embodiments, the calibration model 250 uses a multi-fold (e.g., / / -fold) stratified cross-validation structure to improve the accuracy of the predicted values over time. In cross-validation, the observed data are sequentially partitioned into independent training and testing subsets. Multiple equally-sized subsamples are generated randomly with the time series observations from one physical sensor being grouped in the same partition. The calibration model 250 is trained with at least one of the subsets and is tested on one remaining, held-out subsample as a testing dataset. This training process repeated a total of n times (“ / / -fold”) with each of the subsamples being used once as the testing dataset. Cross-validation allows for the calculation of performance statistics and the ability to generalize the model to new data as part of the n-fold stratified retraining process.
[0026] The calibration model 250 may also receive additional features for analysis including: rainfall data, stream flow data, location data (e.g., latitude, longitude, altitude, and combinations thereof), hydrologic unit code (HUC12) land cover classification, topographic models, temporal livestock density data, temporal nutrient / insecticide / herbicide application data.
[0027] Using these data, the calibration model 250 can identify attribution data 270, which identify from the various actors and environmental factors the causes of water quality variability due to predictors including land-management practices, water-management practices, and weather events (e.g., storms, wildfires, droughts).
[0028] Figure 3 is an example system 300, as may use the described models for MRV with sensor-remote sensing fusion, according to embodiments of the present disclosure. The system 300 includes one or both of water treatment sites 310 that extract water from a watershed, and treated water site 320 that output water to the watershed, which the system 300 may signal or control via a machine learning model 330 (running on one or more computing devices) that receives data from a plurality of sensors 340. In various embodiments, the plurality of sensors 340 includes sensors 340 that are disposed in bodies 110 of water in the watershed that are configured to continuously collect and transmit data to the machine learning model 330. In various embodiments, the plurality of sensors 340 include data collection devices that provide qualitative data collected via survey of actors, inputs of topographical and uses of land 120 in the watershed, soil data for the watershed, precipitation data, temperature data, forecasted weather data, and other data related to the watershed that is not reported directly from the environment.
[0029] In various embodiments, the sensors 340 report various data related to water quality and water volume in the watershed. These sensors 340 may include an optical sensor to identify transmittance, reflectance and / or fluorescence of the water. In various embodiments the sensors 340 are configured to collect remote sensing data such as rainfall, biomass cover and land surface properties, and / or quantitative and qualitative survey data. The data collected by the sensors 340 may be collected via wireless transmissions (e.g., using cellular communication or satellite uplinks) so that the sensors may remain deployed in the field and not require an operator to go out to where the sensor 340 is deployed collect the data from the sensors 340.
[0030] In various embodiments, the sensors 340 may be deployed to various bodies 110 of water that include constantly moving water (e.g., rivers), bodies of intermittent moving water (e.g.,seasonally dry creeks), natural bodies of standing water (e.g., lakes), manmade bodies of standing water (e.g., reservoirs), manmade bodies of moving water (e g., aqueducts).
[0031] In some embodiments, the sensors 340 may include (or be supplemented with data from) devices used to collect farm survey details, such as the types and quantities of crops / livestock present on a parcel of land; the types and quantities of fertilizers, pesticides, and herbicides used; harvest and planting timings, and other operation details of the farm. Additionally or alternatively, the sensors 340 may include (or be supplemented with data from) devices used to collect land survey data details, such as soil type, demarcations between properties, topologies, plant cover, seasonal precipitation data, or the like.
[0032] Although illustrated with respect to a natural watershed, the present disclosure contemplates that the sensors 340 and machine learning model 330 may similarly be deployed to and used with respect to various water systems, including natural man-made lakes, canals, and seas / oceans, and various closed (or semi-closed) systems. Accordingly, the presently described systems can be used in non-facility based water treatment applications, including waters or chlorinated piped systems that are nominally self-contained (e.g., not continuously pulling from or discharging to rivers / streams / reservoirs), such as in seagoing vessels, space vessels, secure facilities, wherein the “watershed” refers to a collection area or outflow area that a defined environment may collect from or discharge to during normal operations or intermittently. For example, a vessel may include water shipments, water recyclers, showers / sinks / toilets, etc., in a first artificial watershed for potable water, and may include bilges in a second artificial watershed for buoyancy / balance systems in the vessel that are periodically (but not continuously) opened to the natural environment to dump or intake water.
[0033] Accordingly, the treated water sites 320 may include water treatment sites 310 that output potable water, but may also include other grades of treated water. For example, a water treatment site 310 may treat water to remove a given microbe, a given living organism, a given chemical, or fecal matter may yield higher-quality, but still not potable (for human consumption) water. Water treatment sites 310 may include human-controlled treatment plants, biological filters (e.g., mangrove forests), managed wetlands, septic fields, stocked bodies of water (e.g., to introduce a given microbe, animal, plant, algae, or the like), and gated bodies of water (e.g., to remove, kill, or deter entry of various microbes, animals, plants, algae, or the like) and the like where one or more water quality parameters are intended to be altered.
[0034] Using the machine learning model 330, the collected data from the plurality of sensors 340 are used to generate a time series of estimates for water quality in the watershed, which in turn is used to activate at least one water treatment site 310 to extract or forego extraction of water from the watershed or at least one treated water site 320 to discharge or forego discharge of water into the watershed based on the time series of estimated of water quality. As will be appreciated, foregoing extraction may include a total pause in water extraction for a predefined length of time or a reduction in water extraction of at least 5% of the volume normally extracted during a similar time period of nominal extraction. In some embodiments, discharge includes diverting potable water from a water treatment site 310 into the watershed (rather than a municipal water network) after treatment or processing, opening a reservoir, or outputting water from a treated water site 320 to the watershed. As will be appreciated, foregoing discharge to the watershed may include discharging water to a retaining pond or other body that can be separated or blocked from bodies 110 that are part of the watershed or a reduction in water output of at least 5% of the volume normally output during a similar time period of nominal output. Additionally, discharge can include water (treated or collected) and one or more treatment solution for affecting water quality downstream from the treated water site 320 within the watershed.
[0035] For example, when the water quality in the watershed is impaired by wildfire in the lands within the watershed based on a first data series from a first sensor and a second data series from a second sensor, the machine learning model 330 may reduce extraction from the bodies 110 in the watershed to improve downstream water quality (e.g., by diluting the effects of the wildfire on the water) and thereby reduce strain on downstream treatment facilities or actors. Additionally or alternatively, the machine learning model 330 may increase extraction from the bodies in the watershed to reduce the effect of runoff from the land affecting the flow in the bodies 110 (e.g., due to lack of vegetation increasing water input to the bodies 110).
[0036] For example, when the water quality in the watershed is impaired by human development in the watershed based on a data series from the sensors, such as farming, building, diverting streams, or the like, the machine learning model 330 may time the extraction from or input to the bodies 110 based on human activities to reduce a strain on water treatment sites 310 and treated water sites 320 (e.g., by timing extraction to reduce intake of runoff fertilizers, pesticides, waste, or debris, by timing output to dilute the effect of runoff fertilizers, pesticides, waste or debris).
[0037] In various embodiments, the system may seek to optimize water usage according to various targets. These targets may include goals set by an operator of a water treatment site 310 or treated water site 320, such as reduced power usage, timed power usage to generation capacity of renewable generation systems, reduced reagent usage, improved flowrates, increases facility uptime / reduced maintenance expenses, or the like. In some embodiments, these targets may include regulatory set mandates (e.g., a maximum content in a body 110 of water for a given chemical) or green initiative goals, such as the conditions to receive (or avoid forfeiting) carbon credits.
[0038] Because not all water treatment sites 310 in a given watershed may be configured to affect all water quality parameters of interest, or that a first water treatment site 310 may be more efficient or effective at affecting a given water quality parameter than a second water treatment site 310, the machine learning model 330 is able to engage in water quality trading throughout the environment. This water quality trading may be between multiple water treatment sites 310 or treated water sites 320, but may also be between one or more water treatment sites 310 and surrounding users, or between two or more surrounding users (and no water treatment sites 310 or treated water sites 320).
[0039] For example, if a managed septic field is used as a treatment site 310 for multiple users, the model 330 can allocate usage (e.g., in total amount, flow within a given time period, etc.) between the multiple users to avoid or reduce runoff from the treatment site 310 into a local waterway. Similarly, if the land of two different users drain into a shared waterway with no intervening treatment sites 310, the model 330 can advise the users on how and when to apply fertilizer to avoid excessive runoff into the shared waterway that would negatively affect other users who are downstream from the advised users, but upstream from any treatment sites 310.
[0040] In another example, if two operators of water treatment facilities at different locations in a given watershed are collectively tasked with reducing a microbe count in a waterway to or below a given point downstream to both facilities via individual control of the two facilities, the machine learning model 330 can identify how the two facilities can most effectively reach the goal, which may include identifying users within the watershed to communicate with to curtail certain activities at various times (e.g., to manage or reduce run off waters from agricultural lands to a manageable amount by the facilities). Additionally or alternatively, the machine learning model330 may be used to trade quality metrics throughout the watershed so that various actors may more efficiently reach the water quality goals.
[0041] In various embodiments, the machine learning model 330 generates the time series of estimates for control of the water system by identifying or calculating changepoints within the data. A changepoint may be identified by calculating a first standardized variable for a first segment of the data and a second standardized variable for a second segment of the data and determining that a difference between the first and second standardized variables exceeds an optimal threshold. Once this difference has been identified as exceeding the optimal threshold, the machine learning model 330 identifies a changepoint between the first segment and the second segment split the data into intervals at the changepoints and may then classify the intervals. These time series of estimates may identify one or more of predicted water quality, water volume, estimated carbon credits, and environmental benefits a water source based on the classified intervals using the data from a subset of the plurality of separate water sources.
[0042] For example, the machine learning model 330 may use these estimates to show compliance with or attainment of various carbon credit targets (e.g., to receive credit for these carbon credits) based on water quality or bioaccumulation in the watershed affected by water management policies. In another example, the machine learning model 330 may use the estimates to identify potential sources to receive some or all of a carbon credit, or be penalized (or identified as a target to work with) when land use policies by those entities affect water management policies in reaching (or nor reaching) a carbon credit target. Accordingly, the machine learning model 330 may attribute various effects in the bodies 110 of water to various actors, and help direct actions to improve land usage in the watershed with specific actors in need of positive or negative reinforcement.
[0043] In another example, the machine learning model 330 may identify the effects of a wildfire or other disaster (e.g., flood, hurricane, tornado) affecting the land of the watershed, and identify, using the time series of estimates, ways to reduce the effect on the water and downstream lands and actors of that disaster.
[0044] In various embodiments, the machine learning model 330 is configured to control various systems linked within the watershed to water quality based on the time series of estimates. The machine learning model 330 can determine which of the quality -linked systems or combinations thereof will have the largest, fastest, most cost-effective (or some combinationthereof) positive effect on water quality for the users or the watershed as a whole and direct the operation of those quality-linked systems at various times to meet various water quality goals. These quality-linked systems may include potable water treatment facilities, wastewater treatment facilities, irrigation equipment, farm equipment (such as fertilizer applicators or harvesters), dams (for water retainment or redirection), fences (e.g., to control the movement or location of livestock, wildlife, or humans), in-stream or in-lake algae treatment systems, water sources (e.g., pumps at wellheads), broadcast systems (e.g., to transmit advisories to persons or entities in the watershed or in neighboring watersheds), and the like.
[0045] For example, the machine learning model 330 may control an amount of water output by irrigation equipment located in the watershed, including at least one of a timing, a duration, and a location of irrigation. For example, the machine learning model 330 may control an amount of fertilizer or pesticide output by farm equipment located in the watershed, including at least one of a timing, a duration, an intensity, a chemical composition, and a location of application. For example, the machine learning model 330 may control movement of livestock within the watershed, including activating virtual or real electric fences. For example, the machine learning model 330 may transmit a boil-water advisory to persons and entities located in the watershed. For example, the machine learning model 330 may change an activation level (e.g., turn on, turn off, increase or decrease level of usage) of an algae treatment technology in the river or stream or a reservoir in the watershed. For example, the machine learning model 330 may release water from a dam fed by or feeding into the river or stream or control the dam to retain additional water from a current level. For example, the machine learning model 330 may direct a water utility to change a drinking water source used to supply users with. For example, the machine learning model 330 may direct a wastewater utility to change an activation level (e.g., turn on, turn off, increase or decrease level of usage) of treatment equipment or change a discharge level (e.g., turn on, turn off, increase or decrease level of usage) at one or more locations in the watershed.
[0046] Figure 4 illustrates a flowchart of an example method 400 for attributing contamination of a water source to a land-based contamination source, according to embodiments of the present disclosure. It will be appreciated that the method 400 is presented at a high level, and that actual embodiments of the method 400 may include additional steps not depicted herein. Additionally, actual embodiments of the method 400 may combine steps depicted herein or perform additional actions incorporated into but not explicitly discussed as elements of the steps illustrated.
[0047] At block 402, a contamination source detection system receives measurements of an optical property of water from a water fixture configured to monitor a water source. For example, a water fixture which includes a sensor 340 configured to measure a fluorescence of water from a well may transmit source data 220 from the sensor 340 to a contamination source detection system executing on a distributed computing system. The transmissions may be continuous or periodic, and may occur over radio frequency, optical communications, wired electrical communications, sonic communications, combinations thereof, or any other form of communication. The transmissions may be direct or by way of a network such as the internet, a cellular telephone network, a satellite telephony network, or combinations thereof.
[0048] At block 404, the contamination source detection system receives remote data including survey data about one or more land use metrics. For example, the contamination source detection system may request and receive satellite data about forest and biomass cover within a watershed area that feeds the well. The remote data may also include, but is not limited to, data about weather, land surface properties, and other qualitative and quantitative survey data.
[0049] At block 406, the contamination source detection system identifies, employing a process-based land-surface model ensemble and a machine learning-based model, a land-based source of predicted contamination of a water source based upon the remote data and the source data. For example, the contamination source detection system may include a machine learning system, such as a support vector machine, which may consult one or more mechanistic models of a watershed associated with the well in order to determine that, for example, future oil contamination is likely, and to identify a source of that contamination. The one or more mechanistic models may include but are not limited to a riparian filtration model, a hydrological flow model, a meteorological model, a model of a manmade water system such as a storm sewer system or a water treatment system, other models which may affect a quality of water from the well, or combinations thereof.
[0050] Figure 5 illustrates a block diagram of an example system 500 for attributing contamination of a water source to a land-based contamination source, according to embodiments of the present disclosure. In this example system, a first water fixture 510 and a second water fixture 520 monitor a first water source 514 and a second water source 524, respectively. The first water fixture 510 and the second water fixture 520 send first source data 512 and second source data 522, respectively, about one or more water quality metrics of the first water source 514 andthe second water source 524, respectively, to a contamination source detection system 540. The first water fixture 510 and the second water fixture 520 may each feature an optical sensor and a transmitting system configured to transmit the first source data 512 and the second source data 522, respectively to the contamination source detection system 540. One or both of the first water fixture 510 and the second water fixture 520 may contain additional sensors, including but not limited to thermometers, additional optical sensors, pH sensors, flow meters, salinity detectors, any other instrumentation which may be used to measure a water quality metric, or combinations thereof. In embodiments where the first water fixture 510 and the second water fixture 520 contain additional sensors, data from the additional sensors may be included in the first source data 512 and the second source data 522, respectively.
[0051] The contamination source detection system 540 may also receive remote data 532 from a remote data source 530. The remote data source 530 may be a singular source, such as a server configured to gather and send the remote data 532 to the contamination source detection system 540, or the remote data source 530 may be a collection of sources accessed by the contamination source detection system 540. For example, the contamination source detection system 540 may be configured to retrieve publicly available weather data, satellite imagery, and survey data as components of the remote data 532, and these may be retrieved from their various respective internet sources by the contamination source detection system. The remote data 532 may include but is not limited to biomass cover data, qualitative survey data, quantitative survey data, land use data, weather data, any other data that may be relevant to predicting water quality at the first water source 514 and the second water source 524, or combinations thereof.
[0052] The contamination source detection system 540 may combine the remote data 532, the first source data 512, and the second source data 522 to yield a prediction of future contamination of the first water source 514 or the second water source 524. The contamination source detection system 540 may then employ a machine learning-based model and a process-based land-surface model to attribute the predicted contamination to a particular source. For example, the contamination source detection system 540 may include a neural network configured to employ a mechanistic model of riparian filtration rates in a stream to estimate a concentration of fertilizer runoff caused by each of a plurality of farms in the stream’s watershed. By including riparian vegetation cover data along with fertilizer use statistics or land use data, the contamination source detection system 540 may produce a list of identified sources of contamination 542, which includeseach farm and specifies the degree to which each contributes to the contamination. This may be especially advantageous when differing contamination sources produce differing contaminants as the contamination source detection system 540 may facilitate more effective targeting of specific identified sources of contamination 542 which may not have otherwise been readily apparent, or which might otherwise have been masked by more obvious or more prolific sources of contamination.
[0053] The identified sources of contamination 542 may then be transmitted to a controller of a water or land management system configuration 550. For example, a system which recommends crop plantings may receive the identified sources of contamination 542 and direct farmers in the stream’s watershed to use a different fertilizer or to plant a different crop. Additionally or alternatively, a water intake system for a drinking water plant may be configured to forego water intake during periods of heavy predicted contamination or increase water intake and processing (e g., stocking a reservoir) during periods leading up to periods of heavy predicted contamination.
[0054] Figure 6 is a flowchart of an example method 600 for analyzing data in a contamination source detection system, according to embodiments of the present disclosure. It will be appreciated that the method 600 is presented at a high level, and that actual embodiments of the method 600 may include additional steps not depicted herein. Additionally, actual embodiments of the method 600 may combine steps depicted herein or perform additional actions incorporated into but not explicitly discussed as elements of the steps illustrated.
[0055] At block 602, an example contamination source detection system calculates a first standardized variable for a first segment of data and a second standardized variable for a second segment of data. For example, the contamination source detection system 540 may split the first source data 512, the second source data 522, and the remote data 532 into finite segments. The contamination source detection system 540 may perform statistical analysis on each segment to calculate standardized variables for each segment. The statistical analysis may include performing regression analysis to calculate the standardized variables.
[0056] At block 604, the example contamination source detection system 540 determines that a difference between the first standardized variable and the second standardized variable exceeds a predefined threshold. For example, the first standardized variable may be 1.5 and the second standardized variable may be 0.6. With a threshold value of 0.7, the contamination source detectionsystem 540 may determine that the difference between the first standardized variable and the second standardized variable exceeds the threshold.
[0057] At block 606, the example contamination source detection system identifies a changepoint between the first segment and the second segment. For example, responsive to the determination at block 604, the contamination source detection system 540 may mark a point at which the difference between the first standardized variable and the second standardized variable exceeds the predefined threshold as a changepoint. As used in this disclosure, a changepoint may be taken to mean a point at which a time series of data exhibits a change in a trend. Changepoints may be caused by one or more causal factors (both identifiable and unidentifiable) or may be random. For example, a start of a rain storm may significantly increase turbidity in a water source for a time, causing a changepoint to be placed in turbidity data at times roughly corresponding to a beginning and end of the rain storm.
[0058] At block 608, the example contamination source detection system splits the data into intervals at the changepoints. For example, once changepoints have been identified, the contamination source detection system 540 may sectionalize the data into intervals, where each interval contains a portion of the data which follows one or more common trends. This manner of data fragmentation may be advantageously leveraged to produce intervals which begin and end at points in a time series at which important events affecting the data occur. For example, the turbidity data may then be broken up into intervals, with a portion of the data corresponding to the rain storm being separated into a singular interval. An interval may comprise several segments.
[0059] At block 610, the example contamination source detection system classifies the intervals using a machine learning-based model. For example, the contamination source detection system 540 may feed the intervals created at block 608 into a random forest machine learning model, which may label the interval of the turbidity data corresponding to the rain storm as having an elevated degree of contamination, and may use weather data and survey included in remote data 532 to determine that a likely cause of the elevated turbidity was the rain storm coupled with a series of construction projects located within the watershed. The contamination source detection system may then take action to limit intake from a contaminated water source until contaminants return to normal levels or take action to mitigate the contamination, such as directing a regulatory agency to focus enforcement of runoff controls on specific sites.
[0060] Figure 7 illustrates an example computing device 700, as may be used MRV with sensor-remote sensing fusion, according to embodiments of the present disclosure. The computing device 700 may include at least one processor 710, a memory 720, and a communication interface 730.
[0061] The processor 710 may be any processing unit capable of performing the operations and procedures described in the present disclosure. In various embodiments, the processor 710 can represent a single processor, multiple processors, a processor with multiple cores, and combinations thereof.
[0062] The memory 720 is an apparatus that may be either volatile or non-volatile memory and may include RAM, flash, cache, disk drives, and other computer readable memory storage devices. Although shown as a single entity, the memory 720 may be divided into different memory storage elements such as RAM and one or more hard disk drives. As used herein, the memory 720 is an example of a device that includes computer-readable storage media, and is not to be interpreted as transmission media or signals per se.
[0063] As shown, the memory 720 includes various instructions that are executable by the processor 710 to provide an operating system 722 to manage various features of the computing device 700 and one or more programs 724 to provide various functionalities to users of the computing device 700, which include one or more of the features and functionalities described in the present disclosure. One of ordinary skill in the relevant art will recognize that different approaches can be taken in selecting or designing a program 724 to perform the operations described herein, including choice of programming language, the operating system 722 used by the computing device 700, and the architecture of the processor 710 and memory 720. Accordingly, the person of ordinary skill in the relevant art will be able to select or design an appropriate program 724 based on the details provided in the present disclosure. In various embodiments, the program 724 may include or make use of a machine learning model that is trained to make determinations as set forth in the present disclosure, and may be retrained or updated based on data collected as set forth in the present disclosure.
[0064] The communication interface 730 facilitates communications between the computing device 700 and other devices, which may also be computing devices as described in relation to Figure 7. In various embodiments, the communication interface 730 includes antennas for wireless communications and various wired communication ports. The computing device 700 may alsoinclude or be in communication, via the communication interface 730, one or more input devices (e g., a keyboard, mouse, pen, touch input device, etc.) and one or more output devices (e.g., a display, speakers, a printer, etc.).
[0065] Although not explicitly shown in Figure 7, it should be recognized that the computing device 700 may be connected to one or more public and / or private networks via appropriate network connections via the communication interface 730. It will also be recognized that software instructions may also be loaded into a non-transitory computer readable medium, such as the memory 720, from an appropriate storage medium or via wired or wireless means.
[0066] Accordingly, the computing device 700 is an example of a system that includes a processor 710 and a memory 720 that includes instructions that (when executed by the processor 710) perform various embodiments of the present disclosure. Similarly, the memory 720 is an apparatus that includes instructions that, when executed by a processor 710, perform various embodiments of the present disclosure.
[0067] Certain terms are used throughout the description and claims to refer to particular features or components. As one skilled in the art will appreciate, different persons may refer to the same feature or component by different names. This document does not intend to distinguish between components or features that differ in name but not function.
[0068] As used herein, the term “optimize” and variations thereof, is used in a sense understood by data scientists to refer to actions taken for continual improvement of a system relative to a goal. An optimized value will be understood to represent “near-best” value for a given reward framework, which may oscillate around a local maximum or a global maximum for a “best” value or set of values, which may change as the goal changes or as input conditions change. Accordingly, an optimal solution for a first goal at a given time may be suboptimal for a second goal at that time or suboptimal for the first goal at a later time.
[0069] As used herein, various chemical compounds are referred to by associated element abbreviations set by the International Union of Pure and Applied Chemistry (IUPAC), which one of ordinary skill in the relevant art will be familiar with. Similarly, various units of measure may be used herein, which are referred to by associated short forms as set by the International System of Units (SI), which one of ordinary skill in the relevant art will be familiar with.
[0070] As used herein, “about,” “approximately” and “substantially” are understood to refer to numbers in a range of the referenced number, for example the range of -10% to +10% of thereferenced number, preferably -5% to +5% of the referenced number, more preferably -1 % to +1% of the referenced number, most preferably -0.1% to +0.1% of the referenced number.
[0071] Furthermore, all numerical ranges herein should be understood to include all integers, whole numbers, or fractions, within the range. Moreover, these numerical ranges should be construed as providing support for a claim directed to any number or subset of numbers in that range. For example, a disclosure of a range from 1 to 10 should be construed as supporting ranges of any two numbers X and Y that fall into the initial range of from 1 to 10 where X > 1 and Y < 10.
[0072] As used in the present disclosure, a phrase referring to “at least one of’ a list of items refers to any set of those items, including sets with a single member, and every potential combination thereof. For example, when referencing “at least one of A, B, or C” or “at least one of A, B, and C”, the phrase is intended to cover the sets of: A, B, C, A-B, B-C, A-C, and A-B-C, where the sets may include one or multiple instances of a given member (e.g., A-A, A-A-A, A-A- B, A-A-B-B-C-C-C, etc.) and any ordering thereof. For avoidance of doubt, the phrase “at least one of A, B, and C” shall not be interpreted to mean “at least one of A, at least one of B, and at least one of C”.
[0073] As used in the present disclosure, the term “determining” encompasses a variety of actions that may include calculating, computing, processing, deriving, investigating, looking up (e g., via a table, database, or other data structure), ascertaining, receiving (e.g., receiving information), accessing (e.g., accessing data in a memory), retrieving, resolving, selecting, choosing, establishing, and the like.
[0074] Without further elaboration, it is believed that one skilled in the art can use the preceding description to use the claimed inventions to their fullest extent. The examples and aspects disclosed herein are to be construed as merely illustrative and not a limitation of the scope of the present disclosure in any way. It will be apparent to those having skill in the art that changes may be made to the details of the above-described examples without departing from the underlying principles discussed. In other words, various modifications and improvements of the examples specifically disclosed in the description above are within the scope of the appended claims. For instance, any suitable combination of features of the various examples described is contemplated.
[0075] Within the claims, reference to an element in the singular is not intended to mean “one and only one” unless specifically stated as such, but rather as “one or more” or “at least one”.Unless specifically stated otherwise, the term “some” refers to one or more. No claim element is to be construed under the provision of 35 U.S.C. § 112(f) unless the element is expressly recited using the phrase “means for” or “step for”. All structural and functional equivalents to the elements of the various embodiments described in the present disclosure that are known or come later to be known to those of ordinary skill in the relevant art are expressly incorporated herein by reference and are intended to be encompassed by the claims. Moreover, nothing disclosed in the present disclosure is intended to be dedicated to the public regardless of whether such disclosure is explicitly recited in the claims.
Claims
CLAIMSThe invention is claimed as follows:
1. A system, comprising: a plurality of separate water fixtures, wherein each of the plurality of separate fixtures includes an optical sensor configured to measure a water quality metric of a water source and a data transmission system configured to transmit source data from each of the plurality of water fixtures, respectively; a remote data source configured to transmit remote data which includes survey data about one or more land use metrics; a contamination source detection system, configured to: receive the source data from the plurality of water fixtures and the remote data from the remote data source; and employ a process-based land-surface model ensemble and a machine learningbased model to identify a land-based source of predicted contamination of a water source based upon the remote data and the source data.
2. The system of claim 1, wherein identifying the land-based source of predicted contamination further comprises: calculating changepoints in the source data and the remote data by: calculating a first standardized variable for a first segment of the data and a second standardized variable for a second segment of the data; determining that a difference between the first standardized variable and the second standardized variable exceeds a predefined threshold;identifying a changepoint between the first segment and the second segment; and splitting the data into intervals at the changepoints; and classifying the intervals using the machine learning-based model.
3. The system of claim 1, wherein a given optical sensor is configured to measure at least one of a transmittance, a reflectance, and a fluorescence.
4. The system of claim 1, wherein responsive to identifying the land-based source, the contamination source detection system modifies a configuration of a water system.
5. The system of claim 1, wherein responsive to identifying the land-based source, the contamination source detection system modifies a configuration of a land management system.
6. The system of claim 1, wherein the process-based land-surface model ensemble includes at least one mechanistic model of a watershed.
7. The system of claim 1, wherein the contamination source detection system includes a microcontroller having a processor and a memory storing instructions.
8. The system of claim 1, wherein the contamination source detection system executes on a server or a distributed computing system.
9. A method, comprising:receiving source data about an optical property of water from a water fixture configured to monitor a water source; receiving remote data including survey data about one or more land use metrics; and identifying, employing a process-based land-surface model ensemble and a machine learning-based model, a land-based source of predicted contamination of a water source based upon the remote data and the source data.
10. The method of claim 9, further comprising: calculating changepoints in the source data and the remote data by: calculating a first standardized variable for a first segment of the data and a second standardized variable for a second segment of the data; determining that a difference between the first standardized variable and the second standardized variables exceeds a predefined threshold; identifying a changepoint between the first segment and the second segment; and splitting the data into intervals at the changepoints; and classifying the intervals using the machine learning-based model.
11. The method of claim 9, wherein the optical property is at least one of a transmittance, a reflectance, and a fluorescence.
12. The method of claim 9, further comprising modifying a configuration of a water system, responsive to identifying the land-based source.
13. The method of claim 9, further comprising modifying a configuration of a land management system, responsive to identifying the land-based source.
14. The method of claim 9, wherein the process-based land-surface model ensemble includes at least one mechanistic model of a watershed.
15. A non-transitory computer-readable medium storing instructions which, when executed by a processing device, cause the processing device to: receive source data about an optical property of water from a water fixture configured to monitor a water source; receive remote data including survey data about one or more land use metrics; and identify, employing a process-based land-surface model ensemble and a machine learning-based model, a land-based source of predicted contamination of a water source based upon the remote data and the source data.
16. The non-transitory computer-readable medium of claim 15 storing further instructions which, when executed by the processing device, cause the processing device to: calculate changepoints in the source data and the remote data by: calculating a first standardized variable for a first segment of the data and a second standardized variable for a second segment of the data; determining that a difference between the first standardized variable and the second standardized variables exceeds a predefined threshold; identifying a changepoint between the first segment and the second segment; andsplitting the data into intervals at the changepoints; and classify the intervals using the machine learning-based model.
Citation Information
Patent Citations
Systems and methods for forecasting bacterial water quality
US20150323514A1
Simulation of soil condition response to expected weather conditions for forecasting temporal opportunity windows for suitability of agricultural and field operations
US20160247076A1
Accessing agriculture productivity and sustainability
US20220061236A1