Water quality prediction and evaluation method based on multi-source data
By analyzing multi-source data, constructing an uncertainty index, and dynamically adjusting sensor strategies, the problems of data integrity and environmental adaptability in traditional water quality prediction and assessment are solved, achieving efficient and accurate water quality monitoring.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-18
- Publication Date
- 2026-04-03
AI Technical Summary
Traditional water quality prediction and assessment methods rely on single water quality parameter data, neglect the integrity of data in time and space, and cannot adapt to environmental changes, resulting in insufficient assessment reliability and waste of resources.
A water quality prediction and assessment method based on multi-source data is adopted. By acquiring data quality and environmental quality data, the spatiotemporal deviation coefficient and environmental difference coefficient are calculated to construct an uncertainty index and dynamically adjust the sensor sampling strategy and power management.
It enables comprehensive assessment of water quality prediction, dynamically adapts to the needs of different water areas, improves monitoring accuracy and resource utilization efficiency, and avoids the resource waste and insufficient accuracy of traditional methods.
Smart Images

Figure CN121786355A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of water quality monitoring technology, and in particular to a water quality prediction and assessment method based on multi-source data. Background Technology
[0002] Water quality monitoring is a core support for water resource protection and water environment management. Traditional water quality prediction and assessment methods often rely on the analysis of single water quality parameter data, which has significant limitations in assessment dimensions. On the one hand, data quality assessment only uses the absolute number of sensors or simple data integrity as the judgment standard, without considering the sampling sparsity difference in the time dimension and the local coverage gap in the spatial dimension, which leads to misjudgment of the spatiotemporal integrity of data, and thus affects the basic reliability of prediction and assessment. On the other hand, environmental impact analysis only focuses on the surface fluctuations of real-time environmental parameters and does not take into account the characteristics and patterns of historical water quality events. It is difficult to identify hidden risks that are stable in real-time data but close to the risk threshold, resulting in a lag in early warning of sudden changes in water quality.
[0003] Meanwhile, existing sensor monitoring strategies are mostly fixed modes and cannot be adaptively adjusted according to data quality defects and dynamic environmental changes: high-frequency sampling is maintained even when data quality is sufficient and the environment is stable, resulting in energy consumption and resource waste; when data coverage is insufficient or the environment fluctuates drastically, the monitoring accuracy is insufficient due to the low sampling frequency and insufficient sensor activation.
[0004] Furthermore, the lack of a unified multi-dimensional fusion standard for uncertainty quantification and the fragmented assessment of various factors related to data quality and environmental impact lead to a lack of scientific and unified decision-making basis for strategy adjustments, further exacerbating the imbalance between monitoring accuracy and energy consumption, and making it difficult to adapt to the complex monitoring needs of different water bodies such as rivers, lakes, and reservoirs.
[0005] Therefore, a water quality prediction and assessment method based on multi-source data is proposed to address the aforementioned problems. Summary of the Invention
[0006] The purpose of this invention is to propose a water quality prediction and assessment method based on multi-source data in order to solve the above-mentioned problems.
[0007] To achieve the above objectives, the present invention adopts the following technical solution: Water quality prediction and assessment methods based on multi-source data include: Acquire data quality data and environmental quality data; The spatiotemporal deviation coefficient is obtained after analyzing the data quality data; The environmental difference coefficient is obtained after analyzing the environmental quality data; The uncertainty index is obtained by combining the two coefficients. The sampling strategy and power management status of the sensor are controlled based on the uncertainty index.
[0008] Preferably, the acquisition of data quality data and environmental quality data specifically includes: Data quality data includes: geographical boundary information of the water quality monitoring area; current sensor operating status data; Environmental quality data includes key environmental variable parameters filtered according to the characteristics of the water area.
[0009] Preferably, the step of analyzing the data quality data to obtain the spatiotemporal deviation coefficient specifically includes: Acquire and count the number of sensors that are in normal working condition within the water quality monitoring area within the current preset time period, and obtain the total number of sensors; The current sensor density is obtained by dividing the total number of sensors by the area of the corresponding water quality monitoring area. The average sensor density is obtained by dividing the number of sensors that are in normal working condition within the current preset time period by the corresponding water quality monitoring area. Subtract the average sensor density from the current sensor density. If the resulting value is less than 0, take its absolute value and record it as the sensor difference.
[0010] Preferably, the method further includes: Obtain the sensor's base sampling interval and maximum allowable sampling interval; Based on the timestamp of the actual data transmitted by the sensor, obtain the sensor sampling interval within the current preset time period to get the current sampling interval; The result of (current sampling interval - basic sampling interval) / (maximum allowed sampling interval - basic sampling interval) is normalized to [0,1] and recorded as sparsity; Obtain the sparsity of each sensor, sort them in descending order according to the value of the sparsity, and extract the k largest sparsity values, where k≥3. Extract the three largest sparsities from the k sparsities and locate the sensor positions of the three sparsities; Using the locations of the three sensors as endpoints, connect the endpoints with straight lines to form a closed triangle; calculate the area of the triangle and divide the area of the triangle by the area of the water quality monitoring area to obtain the coverage anomaly. After normalizing the sensor difference and coverage anomaly, a weighted summation is performed to obtain the spatiotemporal deviation coefficient.
[0011] Preferably, the process of analyzing environmental quality data to obtain the environmental difference coefficient specifically includes: Based on the characteristics of the monitored water area, core variable parameters and environmental parameter sensors are determined; Extract environmental parameter values from three consecutive samples, with the sampling interval matching that of the water quality parameter monitoring sensor; Calculate the rate of change of environmental parameters for two consecutive times to obtain two rates of change, and take the maximum value of the two rates of change as the maximum rate of change; Statistically calculate the rate of change of all environmental parameters of environmental variables within a preset time period, and take the 90th percentile value as the threshold for environmental parameter changes. The uncertainty of the rate of change of environmental parameters is calculated, and the results are normalized to [0,1] to obtain the unstable values.
[0012] Preferably, the method further includes: Construct a historical event feature library: acquire typical water quality events within a preset time period in the past, extract the combination of environmental features within a preset time period before each event, and form a feature vector before the event; wherein, the feature dimension is uniform, and each feature is the average value within 1 hour; Collect the average environmental parameters within the current preset time period according to the dimensions consistent with the historical event feature database, and form the current feature vector; Calculate the dot product of the feature vector before the event and the current feature vector, and the product of the magnitudes of the feature vector before the event and the current feature vector, respectively. The similarity is obtained by dividing the dot product of two vectors by the product of their magnitudes. After normalizing the unstable values and similarity, a weighted summation is performed to obtain the environmental difference coefficient.
[0013] Preferably, the weighting factors for the preset spatiotemporal deviation coefficient and environmental difference coefficient are calculated by multiplying the spatiotemporal deviation coefficient and environmental difference coefficient with their corresponding weighting factors, and then summing them to obtain the uncertainty index.
[0014] Preferably, the sampling strategy and power management state based on the uncertainty exponent control sensor specifically include: Three threshold ranges are preset, and each threshold range corresponds to an uncertainty level, namely, level 1 uncertainty, level 2 uncertainty, and level 3 uncertainty. By matching the uncertainty index with the range of values of the three sets of thresholds, the uncertainty level corresponding to the uncertainty index is obtained.
[0015] Preferably, the method further includes sampling strategies and power management state designs corresponding to each uncertainty level: Level 1 Uncertainty: The sampling frequency maintains the basic sampling interval; only critical sensors remain active, while non-critical sensors operate in intermittent activation mode. Level 2 uncertainty: Sampling frequency strictly maintains the basic sampling interval; all critical sensors are activated; Level 3 uncertainty: The sampling frequency is switched to a high-frequency sampling interval; all critical sensors, non-critical sensors, and backup sensors are activated.
[0016] In summary, due to the adoption of the above technical solution, the beneficial effects of the present invention are: 1. This invention solves the core problem of the one-sided assessment of traditional methods by constructing a two-dimensional assessment system for data quality spatiotemporal integrity and environmental impact. In terms of data quality, it integrates sensor density comparison, temporal sparsity quantification, and local coverage anomaly analysis to accurately characterize data spatiotemporal defects and avoid misjudgment by a single indicator. In terms of environmental impact, it combines the real-time change rate of environmental parameters with the similarity of historical water quality events to capture the impact of explicit environmental fluctuations on water quality and identify implicit risks that are stable in real time but close to the risk threshold.
[0017] 2. This invention achieves dynamic adaptation of monitoring resources by designing differentiated sampling strategies and power management modes based on three levels of uncertainty. Under low uncertainty, an energy-saving mode is adopted, which reduces the sampling frequency and puts non-critical sensors into hibernation to minimize energy consumption. Under medium uncertainty, basic sampling is maintained and local coverage is supplemented to ensure data quality. Under high uncertainty, high-frequency sampling is switched, all sensors are activated, and temporary mobile sensors are activated to enhance monitoring accuracy. This not only avoids the resource waste or insufficient accuracy of traditional fixed strategies, but also adapts to the monitoring needs of different water areas such as rivers, lakes, and reservoirs by adjusting parameters such as weighting factors and thresholds. Attached Figure Description
[0018] Further details, features, and advantages of this application are disclosed in the following description of exemplary embodiments in conjunction with the accompanying drawings, in which: Figure 1 This is a flowchart of the method of the present invention. Detailed Implementation
[0019] Several embodiments of this application will now be described in more detail with reference to the accompanying drawings to enable those skilled in the art to implement this application. This application may be embodied in many different forms and for various purposes and should not be limited to the embodiments set forth herein. These embodiments are provided to make this application thorough and complete, and to fully convey the scope of this application to those skilled in the art. The embodiments described do not limit this application.
[0020] Unless otherwise defined, all terms used herein (including technical and scientific terms) shall have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains. It will be further understood that terms such as those defined in commonly used dictionaries shall be interpreted as having a meaning consistent with their meaning in the relevant field and / or the context of this specification, and shall not be interpreted in an idealized or overly formal sense unless expressly defined herein.
[0021] Example 1
[0022] Its specific implementation method is combined with the appendix Figure 1 Please provide a detailed explanation.
[0023] Appendix Figure 1 The flowchart of the water quality prediction and assessment method based on multi-source data provided in this embodiment of the invention illustrates the complete steps from acquiring data quality data and environmental quality data to controlling the sampling strategy and power management status of the sensor based on the uncertainty index.
[0024] In this embodiment, it includes: Acquire data quality data and environmental quality data; Specifically, it includes: Data quality data includes: basic configuration data, i.e., the geographical boundary information of the water quality monitoring area (such as the area of the reservoir / river section, unit: square kilometers), which needs to be obtained through geographic mapping or preliminary deployment planning documents, and the boundary needs to be fixed (to avoid fluctuations in the subsequent calculation range); real-time operation data, i.e. the current working status data of the sensor; and historical benchmark data, i.e. the sensor operation data of the same period within a preset time period (such as the same season or the same time period). The sensors involved in the real-time operational data and historical benchmark data of the data quality data are water quality parameter monitoring sensors. These sensors are used to monitor core water quality indicators (such as dissolved oxygen, COD, pH, ammonia nitrogen, etc.), and their operational status data (such as whether they are uploaded normally, whether there are outliers, etc.) constitute the key content reflecting the sensor's operational status in the data quality data and are the direct carrier for assessing data quality. Environmental quality data includes key environmental variable parameters selected based on water characteristics (such as flow velocity and rainfall for rivers; and light intensity and wind speed for lakes). These parameters need to be determined in advance through correlation analysis to avoid redundancy.
[0025] The acquisition of environmental quality data relies on environmental parameter sensors. These sensors are deployed specifically according to the characteristics of the water area. For example, sensors in rivers are used to monitor flow velocity and rainfall, and sensors in lakes are used to monitor light intensity and wind speed. They are designed to collect external environmental variables that are strongly correlated with changes in water quality, providing raw data for analyzing the impact of the environment on water quality.
[0026] Clearly defining the specific composition and acquisition methods of data quality data and environmental quality data lays a precise data foundation for water quality prediction and assessment. Among them, the data quality data includes fixed geographical boundary information, real-time sensor operating status and historical benchmark data. This ensures the stability of the monitoring range to avoid calculation deviations, and can accurately reflect the current real status of sensor operation by comparing real-time and historical data. This reduces analysis errors caused by ambiguous data definitions or chaotic sources from the source, and provides a premise for the reliable calculation of subsequent spatiotemporal deviation coefficients.
[0027] Meanwhile, environmental quality data focuses on key environmental variables selected according to the characteristics of water bodies. Redundant parameters are eliminated through correlation analysis, which not only reduces the complexity of data processing and improves efficiency, but also accurately identifies the core factors that are strongly correlated with water quality (such as river flow velocity, rainfall, and lake sunshine and wind speed). This targeted data collection method ensures that environmental quality data can truly reflect the external dynamics affecting water quality, providing high-quality input for the effective calculation of environmental difference coefficients, avoiding irrelevant data from interfering with the assessment results, and making the analysis of environmental fluctuations more targeted and accurate.
[0028] The spatiotemporal deviation coefficient is obtained after analyzing the data quality data; Specifically, it includes: Acquire and count the number of sensors in normal working condition within the water quality monitoring area within the current preset time period, obtain the total number of sensors, exclude nodes that are offline due to faults, have low battery and have not uploaded data, or have continuous abnormal data, and only count the number of sensors that have uploaded valid data within the preset time period. The current sensor density is obtained by dividing the total number of sensors by the area of the corresponding water quality monitoring area. The average sensor density is obtained by dividing the number of sensors that are in normal working condition within the current preset time period by the corresponding water quality monitoring area. Subtract the average sensor density from the current sensor density. If the result is less than 0, take its absolute value and record it as the sensor difference. If the result is greater than 0, discard it, and the sensor difference is set to 0.
[0029] The comparative analysis of sensor density provides a precise method for quantifying the spatial coverage quality of data. Its core lies in comparing the current working sensor density (strictly excluding invalid nodes such as faults and low battery) with the historical average density for the same period. Sensor differential is extracted only when the current density is insufficient. This avoids the one-sidedness of judging coverage by absolute values and eliminates the impact of environmental differences such as seasons and time periods on the normal operation of sensors by using historical benchmarks. This makes the assessment of insufficient spatial coverage more in line with the actual monitoring scenario and provides a reliable spatial dimension quantitative basis for the subsequent calculation of the spatiotemporal deviation coefficient, ensuring that the judgment of coverage adequacy in data quality analysis has a clear standard.
[0030] Meanwhile, this analysis method ensures the accuracy of density calculation and avoids invalid data interfering with the evaluation results by strictly screening valid sensors (only counting nodes that upload valid data within a preset time period). The logic of removing data with current density higher than historical levels focuses on core scenarios with insufficient data quality, reducing the interference of redundant information on subsequent analysis. This makes the calculation of the spatiotemporal deviation coefficient more targeted, laying a rigorous foundation for adjusting monitoring strategies based on data quality and improving the overall evaluation method's ability to capture actual data defects.
[0031] Obtain the sensor's base sampling interval and maximum allowable sampling interval; Based on the timestamp of the data actually transmitted by the sensor, the sensor sampling interval within the current preset time period is obtained, and the current sampling interval is obtained. If the sensor is in a sleep state, the interval is calculated according to the sleep strategy. The result of (current sampling interval - basic sampling interval) / (maximum allowed sampling interval - basic sampling interval) is normalized to [0,1] and recorded as sparsity; If the current sampling interval is less than or equal to the basic sampling interval, the sparsity is set to 0. Obtain the sparsity of each sensor, sort them in descending order according to the value of the sparsity, and extract the k largest sparsity values, where k≥3. Extract the three largest sparsities from the k sparsities and locate the sensor positions of the three sparsities; A threshold for the straight-line distance between two sensor locations is preset. If the straight-line distance between a sensor and two adjacent sensors is greater than the threshold, the sensor is removed, and a new sparsity and its corresponding sensor are extracted from the k sparsity values in order. Using the locations of the three sensors as endpoints, connect the endpoints with straight lines to form a closed triangle; calculate the area of the triangle and divide the area of the triangle by the area of the water quality monitoring area to obtain the coverage anomaly. After normalizing the sensor difference and coverage anomaly, a weighted summation is performed to obtain the spatiotemporal deviation coefficient. The weighted summation process is as follows: preset the weighting factors for sensor difference and coverage anomaly, multiply the sensor difference and coverage anomaly by their corresponding weighting factors respectively, and then sum them to obtain the spatiotemporal deviation coefficient.
[0032] By introducing temporal sparsity calculation and spatial coverage anomaly analysis, the limitations of relying solely on sensor density to assess data quality are overcome, enabling a comprehensive characterization of the spatiotemporal integrity of data.
[0033] Its core lies in: normalizing the sampling density in the time dimension by using the ratio of (current sampling interval - basic sampling interval) to (maximum allowable interval - basic interval), thus accurately quantifying the problem of temporal sparsity; at the same time, by selecting high-sparseness sensors to construct triangular coverage areas and calculating coverage anomalies, it can effectively identify monitoring loopholes in local spaces (such as coverage blind spots caused by overly scattered sensor distribution), and prevent situations where the overall density meets the standard but local data is missing from being masked.
[0034] Furthermore, by using weighted summation to integrate sensor difference (overall insufficient spatial coverage) and coverage anomalies (local spatial gaps) into a spatiotemporal deviation coefficient, we can both preserve the weight differences of data defects in different dimensions (the weights can be adjusted according to the actual scenario) and achieve the unification of quantitative indicators, thus providing a scientific input for the subsequent calculation of uncertainty indices.
[0035] This design allows data quality analysis to no longer view spatial or temporal issues in isolation, but to integrate the impact of both on prediction, making the adjustment of sampling strategies based on this coefficient more targeted; For example, supplementing sensors for areas with high coverage outliers or shortening the sampling interval for highly sparse sensors can improve data quality while avoiding resource waste.
[0036] The environmental difference coefficient is obtained after analyzing the environmental quality data; Specifically, it includes: Water quality is strongly correlated with environmental parameters (flow velocity, rainfall, sunlight, wind speed, etc.). For example, heavy rain washes surface pollutants into rivers, a sudden increase in flow velocity accelerates dissolved oxygen diffusion, and abrupt changes in sunlight affect algal photosynthesis. Rapid changes in these environmental parameters can trigger abrupt changes in water quality trends. Based on the characteristics of the monitored water area, core variable parameters are determined (flow velocity and rainfall are prioritized for rivers, and light intensity and wind speed are prioritized for lakes) to avoid redundant variables that increase the computational load and to determine the environmental parameter sensors. Extract environmental parameter values from three consecutive samples, with the sampling interval matching that of the water quality parameter monitoring sensor (to ensure time synchronization). Calculate the rate of change of environmental parameters for two consecutive times to obtain two rates of change, and take the maximum value of the two rates of change as the maximum rate of change.
[0037] Statistically calculate the rate of change of all environmental parameters of environmental variables within a preset time period, and take the 90th percentile value as the threshold for environmental parameter changes. The uncertainty of the rate of change of environmental parameters is calculated, and the results are normalized to [0,1] to obtain the unstable values; Unstable value = min(maximum rate of change / environmental parameter change threshold, 1.0) (i.e., take 1.0 when the maximum rate of change exceeds the threshold, otherwise calculate according to the ratio), to ensure that the quantification process is reproducible.
[0038] Focusing on core environmental variables strongly correlated with water quality, a scientific mechanism for quantifying environmental dynamic fluctuations was established, providing a precise basis for assessing the impact of the environment on water quality prediction.
[0039] Its core advantages lie in: selecting only key parameters directly related to water quality changes (such as river flow velocity and rainfall), eliminating redundant variables, which reduces the complexity and energy consumption of data processing, and avoids irrelevant information interfering with the assessment results; at the same time, by calculating the rate of change through three consecutive synchronous samplings and taking the 90th percentile value as the threshold, it provides an objective standard for the severity of environmental fluctuations, rather than subjective judgment, ensuring that unstable values can truly reflect the potential impact of sudden changes in environmental parameters on water quality, making the calculation of the environmental difference coefficient more targeted and reliable.
[0040] Normalizing the uncertainty of the rate of change of environmental parameters to the [0,1] interval forms a standardized unstable value, which not only solves the problem of not being able to directly compare different environmental variables (such as flow rate and light intensity) due to differences in magnitude, but also lays the foundation for subsequent weighted fusion with similarity.
[0041] This not only enhances the scientific rigor of the entire assessment method but also accurately captures the risks of water quality trend changes caused by sudden environmental changes. It provides key environmental dimension support for subsequent adjustments to monitoring strategies based on uncertainty indices, ensuring that monitoring strategies can respond to environmental dynamics in a timely manner.
[0042] Construct a historical event feature database: acquire typical water quality events within a preset time period in the past (such as 5 dissolved oxygen hypoxia events and 3 cyanobacterial bloom events that occurred in the past year), extract the combination of environmental features within a preset time period before each event, and form a feature vector before the event; among them, the feature dimensions are unified (such as water temperature, flow rate, dissolved oxygen, ammonia nitrogen, and rainfall, a total of 5 dimensions), and each feature is taken as the average value within 1 hour (to avoid interference from instantaneous fluctuations). The preset duration is the average interval between the appearance of a feature and its occurrence in historical events, thus avoiding distortion in feature extraction caused by uniform duration. Collect the average environmental parameters within the current preset time period according to the dimensions consistent with the historical event feature database, and form the current feature vector; Historical and current features are unified and normalized to avoid the impact of differences in feature magnitude on similarity calculation; Calculate the dot product of the feature vector before the event and the current feature vector, and the product of the magnitudes of the feature vector before the event and the current feature vector, respectively. The similarity is obtained by dividing the dot product of two vectors by the product of their magnitudes. The similarity range is [0,1], and the closer it is to 1, the more similar the scenes are. After normalizing the unstable values and similarity, a weighted summation is performed to obtain the environmental difference coefficient. The weighted summation calculation process is as follows: pre-set the weight factors for unstable values and similarity, multiply the unstable values and similarity with their corresponding weight factors respectively, and then sum them to obtain the environmental difference coefficient.
[0043] By constructing a historical event feature database and calculating the similarity between the current environment and historical risk scenarios, historical experience is transformed into quantifiable assessment criteria, thereby enhancing the foresight and risk warning capabilities of the environmental difference coefficient.
[0044] By extracting combinations of environmental features before typical water quality events (such as water temperature and dissolved oxygen before hypoxia), standardized feature vectors are formed, allowing past event patterns to be reused. By comparing the matching degree between the current environment and historical scenarios using cosine similarity, we can identify hidden risks that appear stable in real-time data but are actually close to the risk threshold in advance, thus avoiding the omission of potential water quality events due to the model relying solely on real-time fluctuations.
[0045] The unstable value (reflecting real-time environmental fluctuations) and similarity (reflecting historical risk associations) are weighted and fused to form an environmental difference coefficient, which preserves the independent impact of the two on water quality prediction (the weights can be adjusted to suit different water characteristics).
[0046] This design addresses the limitations that may arise from relying solely on real-time data. For example, when environmental fluctuations are small, if the environment is highly similar to historical high-risk scenarios, the environmental difference coefficient can still be improved by using similarity weights to ensure a more comprehensive assessment result. Normalization processing eliminates the interference of differences in feature magnitudes, making the fusion calculation more scientific and providing key support for the accuracy of subsequent uncertainty indices.
[0047] The uncertainty index is obtained by combining the two coefficients. The uncertainty index is obtained by multiplying the spatiotemporal deviation coefficient and the environmental difference coefficient by their respective weighting factors and summing the products.
[0048] The preset weighting factor integrates the spatiotemporal deviation coefficient and the environmental difference coefficient to form a unified uncertainty index, which solves the problem of the one-sidedness of single-dimensional assessment and realizes a comprehensive characterization of the uncertainty of water quality prediction.
[0049] The impact of data quality (spatial-temporal bias) and environmental conditions (environmental differences) on prediction reliability is not equivalent. By using weighting factors, different monitoring scenarios can be flexibly adapted (e.g., rivers focus more on environmental fluctuations, so the weight is tilted towards the environmental difference coefficient; lakes focus more on data coverage, so the weight is tilted towards the spatio-temporal bias coefficient). This allows the uncertainty index to reflect both the defects of the data itself and the interference of the external environment, avoiding the distortion of the assessment caused by ignoring a certain dimension, and making the quantification of uncertainty more in line with actual monitoring needs.
[0050] At the same time, this weighted summation approach integrates the quantitative indicators of the two dimensions into a single index, providing a concise and clear decision-making basis for subsequent sensor strategy adjustments.
[0051] Compared to controlling based on two separate coefficients, a unified uncertainty index avoids confusion in decision-making logic (such as conflicts when spatiotemporal deviations are high but environmental differences are low), enabling the adjustment of sampling strategies and power management to be precisely executed based on the overall uncertainty level. This improves the operability and practicality of the method, ensuring that resource investment (such as high-frequency sampling and power supply assurance) is prioritized for scenarios with the highest uncertainty, thus achieving a balance between monitoring efficiency and cost.
[0052] The sampling strategy and power management status of the sensor are controlled based on the uncertainty index. Specifically, it includes: Three threshold ranges are preset, and each threshold range corresponds to an uncertainty level, namely, level 1 uncertainty, level 2 uncertainty, and level 3 uncertainty. By matching the uncertainty index with the range of values of the three sets of thresholds, the uncertainty level corresponding to the uncertainty index is obtained.
[0053] By pre-setting three threshold ranges and corresponding three levels of uncertainty, the abstract uncertainty index is transformed into a concrete and distinguishable evaluation level, providing a clear boundary standard for the quantification of uncertainty. This design avoids the ambiguity in judgment caused by the continuous values of the uncertainty index; For example, instead of getting bogged down in whether the difference between 0.36 and 0.28 is significant, the boundary between medium and low uncertainty is clearly defined through the classification of levels. This provides a unified benchmark for uncertainty assessment in different scenarios and at different times, improving the standardization and interpretability of the method and enabling operators to quickly understand the credibility of the current water quality prediction.
[0054] This classification bridges the uncertainty index with subsequent control strategies, providing a direct basis for the precise execution of sampling strategies and power management.
[0055] Since different levels correspond to specific threshold ranges, the system can automatically match preset control schemes based on the level (such as energy-saving strategies for level 1 and high-frequency monitoring for level 3), avoiding the subjectivity and delay of manual judgment.
[0056] This ensures that uncertainty assessment can be efficiently translated into practical operation, allowing the investment of sensor resources (sampling frequency, power distribution) to be precisely matched with the degree of uncertainty, thus ensuring monitoring quality while avoiding resource waste.
[0057] It also includes the sampling strategy and power management state design corresponding to each level of uncertainty: Level 1 Uncertainty: The sampling frequency is maintained at the basic sampling interval or reduced to the energy-saving sampling interval; only critical sensors (such as dissolved oxygen, COD, and pH) are kept active in the sensor activation state, while non-critical sensors (such as auxiliary monitoring sensors for water temperature and wind speed) are operated in an intermittent activation mode (activated once every 2 basic sampling cycles); all backup sensors are kept in a dormant state (no additional activation is required to avoid redundant coverage). Level 2 Uncertainty: Sampling frequency must be strictly maintained at the basic sampling interval and must not be reduced; all critical sensors and 80% of non-critical sensors must be activated (prioritizing the activation of non-critical sensors in historically sparse areas); after troubleshooting faulty or low-power nodes, if the sensor density is still lower than 70% of the historical average, wake up 1-2 backup sensors (to supplement spatial coverage and reduce sensor differential); increase real-time data verification (e.g., use data directly when the probability of outliers is <0.5, and trigger one supplementary sampling when it is ≥0.5) to avoid amplifying uncertainty caused by moderate data quality; Level 3 uncertainty: Switch the sampling frequency to a high-frequency sampling interval; wake up all critical sensors, non-critical sensors and backup sensors (including dormant nodes) to ensure that the sensor density is ≥90% of the historical average and reduce spatial sparsity; for areas with high coverage anomalies (triangular coverage area exceeds the preset threshold), activate 2 to 3 additional temporarily deployed mobile sensors (such as buoy sensors) to supplement local coverage.
[0058] The above formulas are all dimensionless calculations. The formulas are derived from software simulations based on a large amount of collected data to obtain the most recent real-world results. The preset parameters in the formulas are set by those skilled in the art according to the actual situation.
[0059] The foregoing has only described certain exemplary embodiments of the present invention by way of illustration. Undoubtedly, those skilled in the art can modify the described embodiments in various ways without departing from the spirit and scope of the present invention. Therefore, the foregoing drawings and descriptions are illustrative in nature and should not be construed as limiting the scope of protection of the claims of the present invention.
[0060] It should be noted that, in this document, the use of relational terms such as "first" and "second" is merely for distinguishing one entity or operation from another, and does not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "include," "contain," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that includes a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the statement "includes a…" does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes the element.
[0061] It should be understood that in the various embodiments of this application, the order of the above-mentioned processes does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0062] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0063] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0064] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0065] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0066] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
[0067] The foregoing has only described certain exemplary embodiments of the present invention by way of illustration. Undoubtedly, those skilled in the art can modify the described embodiments in various ways without departing from the spirit and scope of the present invention. Therefore, the foregoing drawings and descriptions are illustrative in nature and should not be construed as limiting the scope of protection of the claims of the present invention.
Claims
1. A water quality prediction and assessment method based on multi-source data, characterized in that, include: Acquire data quality data and environmental quality data; The spatiotemporal deviation coefficient is obtained after analyzing the data quality data; The environmental difference coefficient is obtained after analyzing the environmental quality data; The uncertainty index is obtained by combining the two coefficients. The sampling strategy and power management status of the sensor are controlled based on the uncertainty index.
2. The water quality prediction and assessment method based on multi-source data according to claim 1, characterized in that, Acquiring data quality data and environmental quality data, specifically including: Data quality data includes: geographical boundary information of the water quality monitoring area; current sensor operating status data; Environmental quality data includes key environmental variable parameters filtered according to the characteristics of the water area.
3. The water quality prediction and assessment method based on multi-source data according to claim 1, characterized in that, The spatiotemporal bias coefficient is obtained after analyzing the data quality data, specifically including: Acquire and count the number of sensors that are in normal working condition within the water quality monitoring area within the current preset time period, and obtain the total number of sensors; The current sensor density is obtained by dividing the total number of sensors by the area of the corresponding water quality monitoring area. The average sensor density is obtained by dividing the number of sensors that are in normal working condition within the current preset time period by the corresponding water quality monitoring area. Subtract the average sensor density from the current sensor density. If the resulting value is less than 0, take its absolute value and record it as the sensor difference.
4. The water quality prediction and assessment method based on multi-source data according to claim 3, characterized in that, Also includes: Obtain the sensor's base sampling interval and maximum allowable sampling interval; Based on the timestamp of the actual data transmitted by the sensor, obtain the sensor sampling interval within the current preset time period to get the current sampling interval; The result of (current sampling interval - basic sampling interval) / (maximum allowed sampling interval - basic sampling interval) is normalized to [0,1] and recorded as sparsity; Obtain the sparsity of each sensor, sort them in descending order according to the value of the sparsity, and extract the k largest sparsity values, where k≥3. Extract the three largest sparsities from the k sparsities and locate the sensor positions of the three sparsities; Using the positions of the three sensors as endpoints, and connecting the endpoints with straight lines to form a closed triangle; Calculate the area of the triangle and divide it by the area of the water quality monitoring area to obtain the coverage anomaly. After normalizing the sensor difference and coverage anomaly, a weighted summation is performed to obtain the spatiotemporal deviation coefficient.
5. The water quality prediction and assessment method based on multi-source data according to claim 1, characterized in that, The environmental difference coefficient is obtained after analyzing the environmental quality data, specifically including: Based on the characteristics of the monitored water area, core variable parameters and environmental parameter sensors are determined; Extract environmental parameter values from three consecutive samples, with the sampling interval matching that of the water quality parameter monitoring sensor; Calculate the rate of change of environmental parameters for two consecutive times to obtain two rates of change, and take the maximum value of the two rates of change as the maximum rate of change; Statistically calculate the rate of change of all environmental parameters of environmental variables within a preset time period, and take the 90th percentile value as the threshold for environmental parameter changes. The uncertainty of the rate of change of environmental parameters is calculated, and the results are normalized to [0,1] to obtain the unstable values.
6. The water quality prediction and assessment method based on multi-source data according to claim 5, characterized in that, Also includes: Construct a historical event feature library: acquire typical water quality events within a preset time period in the past, extract the combination of environmental features within a preset time period before each event, and form a feature vector before the event; wherein, the feature dimension is uniform, and each feature is the average value within 1 hour; Collect the average environmental parameters within the current preset time period according to the dimensions consistent with the historical event feature database, and form the current feature vector; Calculate the dot product of the feature vector before the event and the current feature vector, and the product of the magnitudes of the feature vector before the event and the current feature vector, respectively. The similarity is obtained by dividing the dot product of two vectors by the product of their magnitudes. After normalizing the unstable values and similarity, a weighted summation is performed to obtain the environmental difference coefficient.
7. The water quality prediction and assessment method based on multi-source data according to claim 1, characterized in that, The uncertainty index is obtained by multiplying the spatiotemporal deviation coefficient and the environmental difference coefficient by their respective weighting factors and summing the products.
8. The water quality prediction and assessment method based on multi-source data according to claim 1, characterized in that, The sampling strategy and power management status of the sensor are based on uncertainty exponent control, specifically including: Three threshold ranges are preset, and each threshold range corresponds to an uncertainty level, namely, level 1 uncertainty, level 2 uncertainty, and level 3 uncertainty. By matching the uncertainty index with the range of values of the three sets of thresholds, the uncertainty level corresponding to the uncertainty index is obtained.
9. The water quality prediction and assessment method based on multi-source data according to claim 1, characterized in that, It also includes the sampling strategy and power management state design corresponding to each level of uncertainty: Level 1 Uncertainty: The sampling frequency maintains the basic sampling interval; only critical sensors remain active, while non-critical sensors operate in intermittent activation mode. Level 2 uncertainty: Sampling frequency strictly maintains the basic sampling interval; all critical sensors are activated; Level 3 uncertainty: The sampling frequency is switched to a high-frequency sampling interval; all critical sensors, non-critical sensors, and backup sensors are activated.
Citation Information
Patent Citations
Unmanned aerial vehicle-based river hydrological sampling inspection method and system
CN119151387A
Marine ranch water quality parameter real-time correction and compensation method and system of multi-source sensor
CN120448769A
Intelligent fishery comprehensive evaluation system based on multi-dimensional indexes
CN120450530A
River water quality early warning and sampling optimization method, system, equipment and medium
CN120725485A
Water quality parameter prediction method and system based on multi-sensor data fusion
CN120805047A