Lake water quality multi-parameter real-time monitoring method based on Internet of Things
By combining the generation of three-dimensional binding relationship data, distributed computing, and multi-level network structure, the problem of binding multiple parameters with geographical location in lake water quality monitoring is solved, achieving efficient data aggregation and analysis, improving the accuracy and efficiency of lake water quality monitoring, and supporting the scientific nature and reliability of environmental governance.
Patent Information
- Application Number
- CN202511873919.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-12
- Publication Date
- 2026-03-06
AI Technical Summary
Existing lake water quality monitoring methods struggle to accurately link multiple parameters, devices, and geographical locations, leading to data confusion and low analysis efficiency. Furthermore, the lack of regional-level data processing and transmission mechanisms hinders rapid response to water quality anomalies.
By generating three-dimensional binding relationship data, cleaning the data using distributed computing, constructing a multi-level network structure for data aggregation and analysis, grouping data based on geographical and hydrological characteristics, and optimizing data processing using fusion learning and dimensionality reduction methods, the distribution of potential risk sources can be predicted.
It has achieved precise binding and efficient data aggregation of multi-parameter water quality monitoring, improved the accuracy of environmental monitoring and the efficiency of equipment layout, and ensured the scientific nature and reliability of environmental governance.
Smart Images

Figure CN121614901A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of water quality monitoring technology, and in particular relates to a method for real-time monitoring of multiple parameters of lake water quality based on the Internet of Things. Background Technology
[0002] Lakes, as vital freshwater resources, directly impact ecological balance, drinking water safety, and regional sustainable development. With rapid industrialization and urbanization, lake pollution is becoming increasingly severe, making real-time monitoring of water quality crucial for protecting the ecological environment and public health. However, lake water quality is influenced by numerous factors, such as chemical oxygen demand (COD), chlorophyll concentration, and electrical conductivity, as well as the interaction of environmental conditions like wind speed and rainfall. This necessitates multi-dimensional and multi-level data collection and analysis for monitoring. Scientific and effective monitoring methods are not only essential for lake ecological protection but also have profound implications for water resource management.
[0003] Current water quality monitoring methods have significant limitations in practical applications. Many systems rely on sensors at single sites or fixed depths, making it difficult to comprehensively reflect the complex changes in lake water across different areas and depths. For example, some monitoring systems only collect data from the lake surface, ignoring the differences in water quality between the middle and bottom layers, leading to difficulties in locating pollution sources. Furthermore, the correlation between monitoring equipment and water quality parameters is insufficient, and the traceability of data sources is weak, making it difficult to accurately determine which device or sensor location a particular abnormal value originates from. This information silos result in inefficient data analysis and hinder rapid response to water quality anomalies.
[0004] Achieving precise binding of multiple parameters, devices, and geographical locations, and efficiently integrating data among dispersed monitoring points, is one of the core technical challenges in lake water quality monitoring. Firstly, lake water quality monitoring requires establishing accurate mapping relationships between different types of sensors (such as probes measuring chemical oxygen demand or conductivity) and their device identifiers, geographical locations, and installation depths. Due to the complexity of the lake environment, the wide distribution of monitoring points, and the large number of devices, the lack of a unified parameter-device-location association mechanism easily leads to data confusion. For example, an abnormal chemical oxygen demand may be detected in a lake area, but because the data cannot be traced back to the specific sensor and location, it is difficult to determine whether the pollution originates from point source emissions or natural changes.
[0005] Another key challenge lies in the efficient aggregation and regional analysis of data. Lakes are typically divided into multiple sub-lake areas, each generating a large amount of data from monitoring points, involving different water depths and environmental factors. Existing methods often lack regional-level preprocessing and compression mechanisms for data transmission and processing, leading to excessive load on the central platform and inefficient data analysis. For example, when multiple monitoring points in a large lake simultaneously report raw data, the central platform experiences processing delays due to the massive data volume, making it difficult to promptly detect cross-regional water pollution trends.
[0006] Therefore, how to accurately bind parameters and equipment in a multi-site, multi-depth monitoring network, and how to form a global water quality situation perception through regional data aggregation, has become a key issue in real-time monitoring of lake water quality. Summary of the Invention
[0007] To address the aforementioned technical problems, this invention proposes a real-time monitoring method for multiple parameters of lake water quality based on the Internet of Things (IoT), thereby resolving the issues present in the existing technologies.
[0008] To achieve the above objectives, the present invention provides a method for real-time monitoring of multiple parameters of lake water quality based on the Internet of Things, comprising: Based on the initial environmental indicators and auxiliary parameter data obtained from the data acquisition device, combined with the device's location coordinates and installation depth, three-dimensional binding relationship data is generated; Distributed computing is used to clean the three-dimensional binding relationship data of the neighboring area. If the cleaned data exceeds a preset threshold, it is marked as abnormal, and the abnormality detection result is obtained. Based on the anomaly detection results, a normal data subset is extracted from the multi-level network structure, and the normal data subset is reported to the center through the transmission link to obtain the aggregated regional feature parameters. Based on the geographical and hydrological characteristics represented by the regional feature parameters, a grouping analysis method is used to divide the monitoring zones; the changing trends of regional feature parameters within each zone are analyzed to obtain a local dynamic sensing model. If the local dynamic perception model detects abnormal parameter linkage, the influence weight of auxiliary parameters on environmental indicators is analyzed through fusion learning method to obtain comprehensive monitoring indicators. Data from different levels is extracted from the comprehensive monitoring indicators, and the data is reduced in dimensionality using a data simplification method to obtain compressed global perception link data. The compressed global sensing link data is analyzed using a classification prediction method. Based on historical patterns, the distribution of potential risk sources is predicted, resulting in an optimized deployment density configuration scheme.
[0009] Optionally, the process of generating three-dimensional binding relationship data based on the initial environmental indicators and auxiliary parameter data obtained from the data acquisition device, combined with the device's location coordinates and installation depth, includes: Initial environmental indicators and auxiliary parameter data are acquired from the data acquisition device to obtain a data acquisition dataset. The data acquisition dataset is then bound to the device's location coordinates and installation depth using preset association rules to generate initial three-dimensional binding relationship data. A clustering algorithm is used to group the initial three-dimensional binding relationship data to determine spatial distribution groups. The pollutant concentration change trend is analyzed based on the spatial distribution groups; if the change trend exceeds a preset threshold, abnormal areas are marked to obtain an abnormal marker set. The location coordinates and installation depth in the abnormal marker set are acquired, and a depth adjustment command is generated. The installation depth in the initial three-dimensional binding relationship data is updated using the depth adjustment command to obtain optimized three-dimensional binding relationship data.
[0010] Optionally, the process of obtaining anomaly detection results includes: The three-dimensional binding relationship data is grouped into regions according to the neighboring region division rules; distributed computing nodes are used to perform parallel cleaning operations on the data of each region. If there are continuous zero values or over-range readings caused by sensor failure in the data, they are marked as invalid values, and a preliminary cleaned dataset is obtained; the numerical range of each monitoring point in the preliminary cleaned dataset is obtained, and the dissolved oxygen content and pollutant concentration are numerically verified according to the preset threshold boundary. If the monitoring value exceeds the normal environmental parameter range, an anomaly marking mechanism is triggered to determine the abnormal detection result.
[0011] Optionally, the process of extracting a normal data subset from the multi-level network structure based on the anomaly detection results, and reporting the normal data subset to the center via a transmission link to obtain the aggregated regional feature parameters includes: Based on the anomaly detection results, a normal data subset is extracted from the multi-level network structure, and the integrity of the normal data subset is determined. According to the integrity of the normal data subset, it is reported to the center via a transmission link, and a reporting confirmation signal is obtained. If the reporting confirmation signal indicates success, a convergence index is extracted from the reported normal data subset, and the stability of the convergence index is determined. Based on the stable convergence index, regional feature parameters are extracted, and the consistency of the regional feature parameters is determined. Based on the consistency of the regional feature parameters, the k-means clustering algorithm is used to cluster the regional feature parameters to obtain cluster groups. If the number of cluster groups exceeds a preset threshold, dominant features are selected from each group to obtain an optimized parameter set. Based on the optimized parameter set, the multi-level network structure information is integrated to obtain the final regional feature parameters.
[0012] Optionally, the process of analyzing the changing trends of regional feature parameters within each partition to obtain a local dynamic sensing model includes: The changes in regional characteristic parameters of each zone are analyzed to generate a trend sequence; key nodes indicating fluctuations are extracted from the trend sequence, and the spatial distribution of the key nodes is obtained; based on the spatial distribution of the key nodes, hydrological characteristic data are integrated to determine the consistency of the distribution; based on the consistent distribution, geographical characteristic data are fused to obtain a feature fusion set; the feature fusion set is classified using a random forest algorithm to determine the evolution path of regional environmental indicators; based on the evolution path, a local dynamic perception model is constructed.
[0013] Optionally, the process of obtaining comprehensive monitoring indicators includes: The real-time collected data is input into the local dynamic perception model to determine whether parameter linkage anomalies have occurred. If a parameter linkage anomaly is determined to have occurred, auxiliary parameters are extracted from the real-time data that triggered the anomaly to form an auxiliary parameter set. Based on the auxiliary parameter set, a fusion learning method is used to calculate the influence weight of each auxiliary parameter on the environmental indicator to obtain a weight distribution. Based on the weight distribution, the auxiliary parameters and the environmental indicator are weighted and fused to generate a comprehensive monitoring indicator.
[0014] Optionally, the process of extracting data from different levels of the comprehensive monitoring indicators and performing dimensionality reduction on the data using data simplification methods to obtain compressed global perception link data includes: Data from different levels covering the full depth of the monitored objects are extracted from the comprehensive monitoring indicators to determine the dataset to be reduced in dimension. Principal component analysis is used to reduce the dimension of the dataset to obtain simplified data. The covariance matrix is calculated based on the simplified data, and eigenvectors are extracted. The top few eigenvectors with the highest variance contribution rate are retained to form a principal component set. The data is reconstructed through the principal component set to obtain compressed data. Cluster analysis is performed on the compressed data to obtain link groups. The link groups are fused with historical monitoring data and feature parameters of other monitoring zones to generate compressed global sensing link data.
[0015] Optionally, the process of analyzing the compressed global sensing link data using a classification prediction method, predicting the distribution of potential risk sources based on historical patterns, and obtaining an optimized deployment density configuration scheme includes: The compressed global sensing link data is processed using a classification prediction method to obtain the distribution of potential risk sources; high-density risk areas are determined based on the distribution of potential risk sources, and these high-density risk areas are prioritized; deployment density parameters are calculated based on the priority ranking to generate a preliminary configuration scheme; if the density parameters in the preliminary configuration scheme exceed a preset threshold, historical pattern data is fused to adjust the parameters to obtain a corrected configuration scheme; the corrected configuration scheme is trained using a random forest algorithm to obtain an optimized deployment density configuration scheme.
[0016] Compared with the prior art, the present invention has the following advantages and technical effects: This invention collects environmental indicators such as dissolved oxygen, temperature, and pollutant concentration, along with auxiliary parameters, and combines these with location coordinates and installation depth to form a three-dimensional bound data structure, characterizing the spatial distribution features of the monitored objects. Distributed computing is used to clean neighboring area data, removing noise and invalid values, marking abnormal data to filter valid subsets, and reporting to the center through a multi-level network structure to generate regional characteristic parameters. A grouping analysis method based on geographical and hydrological characteristics is used to divide monitoring zones, constructing a local dynamic sensing model to capture parameter change trends. If a linkage anomaly is detected, the influence weight of auxiliary parameters on environmental indicators is analyzed through fusion learning to generate comprehensive monitoring indicators. Subsequently, a data simplification method is used to reduce the dimensionality of the full-depth data, obtaining compressed global sensing link data. A classification prediction method is then used to predict the distribution of potential risk sources based on historical patterns, optimizing the layout density configuration of data acquisition devices. This invention, through multi-level data processing and intelligent analysis, significantly improves the accuracy of water environment monitoring and the efficiency of equipment layout, ensuring the scientific and reliable nature of environmental governance. Attached Figure Description
[0017] The accompanying drawings, which form part of this application, are used to provide a further understanding of this application. The illustrative embodiments and descriptions of this application are used to explain this application and do not constitute an undue limitation of this application. In the drawings: Figure 1 This is a flowchart of a method according to an embodiment of the present invention. Detailed Implementation
[0018] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. This application will now be described in detail with reference to the accompanying drawings and embodiments.
[0019] It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases the steps shown or described may be executed in a different order than that shown here.
[0020] Example 1 like Figure 1 As shown, this embodiment provides a method for real-time monitoring of multiple parameters of lake water quality based on the Internet of Things, including: Step S101: Obtain initial environmental indicators and auxiliary parameter data from the data acquisition device. The data includes dissolved oxygen content, temperature, and pollutant concentration. Bind the data with the device location coordinates and installation depth according to preset association rules to obtain three-dimensional binding relationship data. The relationship data characterizes the spatial distribution characteristics of the monitored object.
[0021] Initial environmental indicators and auxiliary parameter data, including dissolved oxygen content, temperature, and pollutant concentration, are acquired from a data acquisition device to obtain a collected dataset. The collected dataset is then bound to the device's location coordinates and installation depth using preset association rules, generating three-dimensional binding relationship data. A clustering algorithm is used to group the three-dimensional binding relationship data, determining spatial distribution groups. Based on these spatial distribution groups, pollutant concentration trends are analyzed; if a trend exceeds a preset threshold, anomaly regions are marked, resulting in an anomaly marker set. The location coordinates and installation depth of the anomaly marker set are obtained, generating a depth adjustment command. The installation depth in the three-dimensional binding relationship data is updated using this depth adjustment command, resulting in an optimized binding dataset. The spatial distribution uniformity of the optimized binding dataset is calculated; if the uniformity is lower than a preset threshold, the clustering algorithm is repeated to obtain the final distribution model.
[0022] In one possible implementation, initial environmental indicators and auxiliary parameter data are acquired from a data acquisition device. For example, in a river monitoring scenario, the device collects data in real time, such as dissolved oxygen content (e.g., 5.2 mg / L), temperature (e.g., 22.3℃), and pollutant concentration (e.g., ammonia nitrogen, 0.8 mg / L), forming a dataset. This acquisition can promptly capture dynamic changes in water bodies, helping to identify pollution sources early and thus improving the response speed of environmental management.
[0023] Specifically, by using preset association rules, the collected dataset is bound to the device's location coordinates, such as latitude and longitude (120.5, 30.2), and installation depth, such as 1.5m, to generate three-dimensional binding relationship data. This binding establishes multi-dimensional associations, facilitating subsequent spatial analysis and resulting in a more accurate mapping of pollutant distribution.
[0024] For example, in rivers, bound data can reveal how upstream pollutants spread downstream, avoiding the limitations of planar analysis.
[0025] In one possible implementation, clustering algorithms are used to group the three-dimensional binding relationship data to determine spatial distribution groups. For example, the K-means algorithm can be used to divide the data into high-concentration groups and low-concentration groups. This grouping can identify pollutant accumulation areas and improve the targeting of monitoring.
[0026] Specifically, based on the analysis of spatial distribution groups, the trend of pollutant concentration changes is determined. If the trend exceeds a preset threshold, such as a 20% increase in concentration, an abnormal area is marked, resulting in an abnormal marker set. This analysis reveals trends such as a concentration increase from 0.5 mg / L to 1.2 mg / L. After marking, pollution hotspots can be quickly located, leading to beneficial preventative interventions.
[0027] For example, in the lower reaches of a river, if the trend is abnormal, the marker set includes coordinates (120.6, 30.1) and a depth of 2.0m, which supports decision-makers in deploying targeted measures.
[0028] In one possible implementation, the location coordinates and installation depth of the anomaly marker set are obtained, and a depth adjustment command is generated, for example, adjusting the depth from 1.5m to 3.0m to cover deeper contamination. This command optimizes monitoring coverage and enhances data comprehensiveness.
[0029] Specifically, by updating the installation depth in the 3D binding relationship data through depth adjustment commands, an optimized binding dataset is obtained. This update dynamically adapts to environmental changes, enhancing the system's self-adaptability.
[0030] For example, the updated dataset shows that adjusting the depth results in more uniform pollutant concentration monitoring and avoids blind spots.
[0031] In one possible implementation, the spatial distribution uniformity of the optimized bound dataset is calculated. If the uniformity falls below a preset threshold, such as 0.7, the clustering algorithm is repeated to group the data, resulting in the final distribution model. This calculation stops once the uniformity increases from 0.6 to 0.85, and the repeated process ensures model optimization, leading to more reliable long-term monitoring results.
[0032] For example, along the entire length of the river, the final model supports the prediction of pollutant diffusion paths, thus supporting sustainable environmental governance.
[0033] Step S102: Distributed computing is used to clean the three-dimensional binding relationship data of the neighboring area. The cleaning includes removing noise and invalid values. If the cleaned data exceeds a preset threshold, it is marked as abnormal, and the abnormal detection result is obtained. The result is used to filter valid data.
[0034] The three-dimensional binding relationship data is grouped into regions according to the neighboring region division rules. Distributed computing nodes are used to perform parallel cleaning operations on the data of each region. If there are continuous zero values or over-range readings caused by sensor malfunctions in the data, they are marked as invalid values, resulting in a preliminary cleaned dataset. The numerical range of each monitoring point in the preliminary cleaned dataset is obtained, and the dissolved oxygen content and pollutant concentration are numerically verified according to preset threshold boundaries. If the monitored values exceed the normal environmental parameter range, an anomaly marking mechanism is triggered to determine the anomaly detection result. The preliminary cleaned dataset is classified and processed using the marking information in the anomaly detection result. If a data point is marked as an anomaly, it is moved to the verification queue. The anomaly data is corrected and verified using the neighboring region data interpolation method to obtain a quality control report. Valid data that meets the monitoring requirements is screened according to the data rating criteria in the quality control report. If the data quality level reaches the preset standard, it is included in the valid data set to determine the final screening result.
[0035] Specifically, the principle of grouping three-dimensional binding relationship data into regions based on neighboring area division rules lies in using location coordinates and installation depth to define spatial neighborhoods, thereby dividing the monitoring data into logical units. This grouping facilitates targeted treatment based on the distribution of pollutants in different areas of the water body.
[0036] For example, in a lake monitoring scenario, dividing the shallow waters from 0 to 10 meters in depth into one group and the deeper waters from 10 to 20 meters into another can reveal the vertical gradient of pollutants, which helps improve the accuracy of data analysis and reduce computational redundancy.
[0037] In one embodiment, when using distributed computing nodes to perform parallel cleaning operations on data from different regions, if a sensor malfunction causes consecutive zero values, such as dissolved oxygen readings of 0.0 mg / L for more than 5 consecutive sampling points, the data is marked as invalid. This parallel processing can accelerate the cleaning process. For example, when processing a lake dataset containing 1000 monitoring points, by operating multiple nodes simultaneously, the cleaning time can be reduced from several hours to minutes, thereby improving overall monitoring efficiency and ensuring data reliability, and preventing faulty data from interfering with subsequent pollutant concentration trend analysis.
[0038] Specifically, the mechanism of verifying the dissolved oxygen content and pollutant concentration based on preset threshold boundaries after obtaining the numerical range of each monitoring point in the preliminary cleaning dataset is to verify whether the data complies with environmental regulations.
[0039] For example, if the normal range for dissolved oxygen is set to 4.0 to 8.0 mg / L, and a reading of 12.0 mg / L exceeds the threshold, an anomaly marker is triggered. This verification can identify sensor bias or sudden pollution events early, which is beneficial for maintaining data integrity and allows for further optimization of the marker set for anomaly areas based on historical clustering.
[0040] For example, in the classification stage, the initial cleaned dataset is classified based on the labeling information in the anomaly detection results. If a data point is marked as an anomaly, such as a sudden increase in pollutant concentration to 50 ppm, it is moved to the verification queue. Subsequently, a neighboring region data interpolation method is used for correction, for example, using the average of adjacent depths of 5 meters and 7 meters to fill the outlier at the 6-meter depth. This method can restore data continuity, generate quality control reports, and thus support the generation of subsequent depth adjustment instructions, improving the accuracy of the three-dimensional binding relationship.
[0041] Specifically, the process of screening valid data based on the data rating standards in the quality control report involves ratings such as Grade A, which indicates an error of less than 5%. If the preset standard is met, the data is included in the valid set.
[0042] For example, in lake monitoring, the filtered valid data can be used to calculate the spatial distribution evenness; if the evenness is higher than 0.8, it indicates successful optimization. This filtering ensures the reliability of the final results, facilitates the formation of the final distribution model, avoids misjudgments caused by low-quality data, and enhances the decision-making value of environmental monitoring through multi-faceted verification such as numerical range and interpolation correction.
[0043] Step S103: Extract a normal data subset from the multi-level network structure based on the anomaly detection results, and report the normal data subset to the center through the transmission link to obtain the aggregated regional feature parameters.
[0044] Anomaly detection results are used to extract a subset of normal data from the multi-level network structure, and the integrity of this subset is determined. Based on the integrity of the normal data subset, a transmission link is used to report the data, obtaining a reporting confirmation signal. If the reporting confirmation signal indicates success, aggregation indicators are extracted from the reported data to assess aggregation stability. Regional feature parameters are obtained based on aggregation stability to determine parameter consistency. Based on parameter consistency, the aggregation data is processed using a k-means clustering algorithm to obtain cluster groups. If the number of cluster groups exceeds a preset threshold, dominant features are selected from the groups to obtain an optimized parameter set. The optimized parameter set is then integrated with the multi-level network structure information to obtain the final regional feature parameters.
[0045] For example, in the field of environmental monitoring, the process of extracting a subset of normal data from a multi-level network structure using anomaly detection results first involves filtering historically cleaned, three-dimensional binding relationship data. Assuming a river monitoring network with a multi-level structure including a sensor layer, a regional convergence layer, and a central control layer, anomaly detection results mark invalid values such as dissolved oxygen levels below 2 mg / L or pollutant concentrations exceeding 50 ppm. Unmarked normal data is extracted from these results to form a subset, and then the integrity of the subset is checked, for example, verifying whether the data coverage reaches 90%, ensuring that the subset contains data from at least 80% of the monitoring points. This integrity determination helps avoid analytical biases caused by missing data, thereby improving the reliability of subsequent processing.
[0046] In one possible implementation, reporting is performed using a transmission link based on the integrity of the normal data subset.
[0047] Specifically, if the integrity of the subset exceeds a preset threshold, such as 85%, the data is reported to the central server via a wireless transmission link, and a reporting confirmation signal is received. For example, at an upstream monitoring station, the subset data includes dissolved oxygen readings at multiple time points. After reporting, the server returns a success signal, which not only confirms that the data transmission is error-free but also provides timely feedback on potential network problems, thus helping to maintain the continuity of the monitoring system.
[0048] Specifically, if the reported confirmation signal indicates success, then convergence indicators, such as average dissolved oxygen level and average pollutant concentration, are extracted from the reported data to determine the convergence stability.
[0049] For example, the variance of the indicator is calculated. If the variance is less than 0.5, it is considered stable. This step can identify abnormal data fluctuations, ensure the credibility of the aggregated data, and thus provide a solid foundation for regional analysis.
[0050] For example, regional characteristic parameters can be obtained through convergence stability to determine parameter consistency. In downstream river regions, characteristic parameters such as pH range of 7-8 and temperature of 20-25℃ are extracted after stable convergence. Consistency is judged by comparing the correlation between parameters; if the correlation coefficient is greater than 0.8, they are considered consistent. This helps to reveal the intrinsic relationship between environmental parameters and improve the accuracy of monitoring.
[0051] In one possible implementation, the k-means clustering algorithm is used to process the aggregated data based on parameter consistency to obtain cluster groups.
[0052] Specifically, by inputting consistent parameters into the algorithm and setting k=3, clustering into high-pollution, medium-pollution, and low-pollution groups, this grouping simplifies complex data, facilitates the identification of pollution patterns, and is beneficial for targeted intervention.
[0053] For example, if the number of cluster groups exceeds a preset threshold, such as 4, then the dominant feature is selected from the groups, such as prioritizing pollutant concentration as the dominant feature, to obtain an optimized parameter set. This reduces computational load and improves efficiency by removing redundant features, such as secondary temperature data, and refining the set to core parameters.
[0054] In one possible implementation, the final regional feature parameters are obtained by integrating multi-level network structure information through the optimization of the parameter set.
[0055] For example, by integrating the optimized set with data from each layer of the network, a comprehensive parameter such as an overall water quality index of 80 can be generated. This integration can provide a comprehensive view, support decision-making, and enhance the robustness of the system.
[0056] Step S104: Based on the regional characteristic parameters, a grouping analysis method is used to divide the monitoring zones according to geographical and hydrological characteristics, analyze the parameter change trends within the zones, and obtain a local dynamic sensing model. The model characterizes the evolution of regional environmental indicators.
[0057] Based on regional characteristic parameters, a grouping analysis method is used to divide monitoring zones according to geographic and hydrological features, resulting in a set of zones. The parameter change trends are analyzed through these zones to determine trend sequences. If the trend sequence indicates fluctuations, key nodes are extracted from the sequence to obtain node distribution. Hydrological features are integrated based on node distribution to assess distribution consistency. Geographic features are processed through distribution consistency analysis to obtain a feature fusion set. A random forest algorithm is used to classify the evolution of indicators from the fusion set, determining the evolution path. A local dynamic sensing model is constructed based on the evolution path to obtain model parameters.
[0058] In one possible implementation, a grouping analysis method based on regional characteristic parameters is used to divide monitoring zones according to geographical and hydrological characteristics, which can effectively subdivide the data in the multi-level network structure into more manageable units.
[0059] For example, in the field of hydrological monitoring, assuming that regional characteristic parameters include river flow and rainfall data, monitoring points can be divided into upstream, midstream, and downstream zones based on elevation and soil type using grouping analysis methods such as hierarchical clustering, resulting in a set of zones. This approach helps improve the targeting of data processing, avoids the complexity of overall analysis, and thus improves the accuracy of anomaly detection.
[0060] Specifically, the process of determining the trend sequence by analyzing parameter change trends through partitioned set analysis can start with time series data to identify long-term patterns of parameters.
[0061] In one possible implementation, for the upstream region, flow data from the past 12 months is collected, observing an upward trend from January to June and a downward trend from July to December, forming a complete trend sequence. This analysis helps predict potential risks, such as flood warnings, enhancing the practical value of converged regional characteristic parameters.
[0062] For example, if a trend sequence indicates fluctuations, the step of extracting key nodes from the sequence and obtaining node distribution can focus on peak and trough points. In hydrological monitoring, for a sequence in the middle reaches of the river, if the fluctuation exceeds 20%, peak flow nodes, such as the record on the 15th of each month, can be extracted, and then the geographical distribution of these nodes can be mapped. This extraction can reveal uneven distribution issues, support subsequent optimization of normal data subsets, and improve the efficiency of the reporting chain.
[0063] In one possible implementation, the method for determining distribution consistency involves integrating hydrological features based on node distribution and comparing the similarity between nodes.
[0064] For example, by integrating node distribution with historical hydrological characteristics such as river width, a consistency score of 0.8 or higher is considered stable. This assessment helps filter noisy data, ensuring that the normal data subset obtained from multi-level networks is more reliable, resulting in better convergence stability.
[0065] Specifically, the process of processing geographical features through distribution consistency to obtain a feature fusion set can integrate geographical data such as altitude and vegetation cover.
[0066] In one possible implementation, for downstream partitions, if consistency is high, these features are weighted and averaged to form a fusion set, such as a comprehensive risk index of 0.65. This fusion improves the accuracy of parameter consistency assessment, which is beneficial for the subsequent application of the k-means clustering algorithm.
[0067] For example, the Random Forest algorithm can be used to evolve classification indicators from the fusion set and determine the implementation of the evolution path. Indicator changes can be handled through decision tree ensembles. In hydrological scenarios, the flow indicators in the fusion set are classified, evolving into paths from low risk to high risk, such as paths with a length of 5 stages. This determination helps to dynamically adjust monitoring strategies and optimize the selection of dominant features.
[0068] In one possible implementation, the steps of constructing a local dynamic sensing model based on the evolution path and obtaining model parameters include estimating parameters such as sensitivity coefficients.
[0069] For example, a path-based model with a parameter value of 0.75 represents the response speed to fluctuations. This approach enables real-time perception, improves the robustness of the overall network structure, and supports continuous optimization of regional feature parameters.
[0070] Step S105: If the local dynamic perception model detects abnormal parameter linkage, the influence weight of auxiliary parameters on environmental indicators is analyzed by the fusion learning method to obtain comprehensive monitoring indicators.
[0071] Real-time data is collected through a local dynamic sensing model to identify abnormal parameter linkages. If an abnormality is detected, auxiliary parameters are extracted from the abnormal data to obtain an auxiliary parameter set. Based on the auxiliary parameter set, a fusion learning method is used to calculate the influence weight of each auxiliary parameter on environmental indicators, determining the weight distribution. The fluctuation range of environmental indicators is obtained through the weight distribution, yielding the fluctuation range value. If the fluctuation range value exceeds a preset threshold, the weight distribution is adjusted to obtain the adjusted weights. Based on the adjusted weights, the auxiliary parameters and environmental indicators are fused to determine the comprehensive monitoring indicator.
[0072] Real-time monitoring data acquired through data acquisition devices is input into a pre-established local dynamic sensing model for each monitoring zone. This model analyzes and judges whether the coordinated changes between multiple environmental parameters deviate from the inherent pattern that the zone should follow under normal conditions. If a deviation occurs, it is determined that an abnormal parameter linkage has occurred.
[0073] In one possible implementation, real-time data is collected through a local dynamic sensing model. This initially involves the dynamic monitoring of geographical and hydrological characteristics within the region. For example, in a river basin environment, the model collects parameters such as water level, rainfall, and soil moisture in real time. These parameters are then linked and analyzed based on trend sequences obtained from historical zoning analysis. The principle of identifying anomalies in parameter linkages lies in detecting unusual correlations between parameters, such as a sudden rise in water level while rainfall remains unchanged, which may indicate a potential flood risk. Through this assessment, anomalies in environmental evolution can be identified early, thereby improving the timeliness of monitoring and contributing to disaster prevention.
[0074] Specifically, if the correlation coefficient between parameters exceeds 0.8 but the actual data deviation reaches 15%, it is judged as an anomaly, which helps to build a more robust local dynamic perception model.
[0075] For example, when extracting auxiliary parameters from anomalous data, in a river basin scenario, anomalous data may include sudden changes in pollutant concentration or temperature. Parameters such as pH and dissolved oxygen can be extracted as auxiliary parameters to form a set. This extraction, based on the consistency judgment of node distribution, can integrate geographical features, ensure strong parameter correlation, and thus obtain a more comprehensive set of auxiliary parameters, which is beneficial to the accuracy of subsequent analysis.
[0076] In one possible implementation, a fusion learning method is used to calculate the weights. For example, an ensemble learning algorithm can be used to fuse the outputs of multiple models, calculating the influence weight of each auxiliary parameter, such as pH value with a weight of 0.4, dissolved oxygen with 0.3, and temperature with 0.3, for a total of 1.0. This method reveals the degree of influence of parameters on environmental indicators such as water quality indices through weight distribution, which is beneficial for quantifying environmental evolution paths and improving the prediction accuracy of the model.
[0077] Specifically, the fluctuation range is obtained through weighted distribution. For example, if the weighted distribution shows that pH fluctuations cause the water quality index to range from 6.5 to 8.5, then the fluctuation range value is 2.0. This acquisition process integrates feature fusion sets, which can support the stability analysis of environmental indicators from multiple perspectives. For example, combining historical trend sequences to determine whether fluctuations exceed normal levels is beneficial for dynamically adjusting monitoring strategies.
[0078] In one possible implementation, if the fluctuation range exceeds a preset threshold such as 1.5, the weight distribution is adjusted, for example, reducing the weight of high-fluctuation parameters from 0.4 to 0.25 to balance the impact. This adjustment, based on distribution consistency processing, optimizes the feature fusion set, forming a more stable evolution path, which helps reduce false alarms and improve the reliability of model parameters.
[0079] For example, by integrating adjusted weighted auxiliary parameters and environmental indicators in a river basin, the weighted pH value, dissolved oxygen, and water quality index are combined to determine a comprehensive monitoring indicator, such as an overall environmental health score of 85. This integration, through the classification of indicator evolution using a random forest algorithm, can be extended from the core solution to optional implementations, such as incorporating node distribution judgments based on geographical features, ensuring the comprehensiveness of the indicators and benefiting the long-term application of local dynamic perception models. From multiple perspectives, this logical progression supports the integrity of environmental monitoring. For instance, the mutual support between early anomaly detection and subsequent weight adjustment enhances the perception capability of regional environmental indicator evolution, ultimately achieving more effective disaster early warning and resource management.
[0080] Step S106: Extract data from different levels from the comprehensive monitoring indicators. The data covers the full depth of the monitoring object. Use a data simplification method to perform dimensionality reduction on the data to obtain compressed global perception link data.
[0081] Hierarchical data is obtained through comprehensive monitoring indicators, covering the full depth of the monitored objects, to determine the data to be extracted. Principal component analysis is used to reduce the dimensionality of the extracted data, resulting in simplified data. A covariance matrix is calculated based on the simplified data, representing the correlation between the simplified data, yielding eigenvectors. If the number of eigenvectors exceeds a preset threshold, the top few eigenvectors are retained, resulting in a principal component set. The simplified data is reconstructed using the principal component set, resulting in compressed data. Cluster analysis is applied to the compressed data to group the compressed data, resulting in link groups. Global data is then fused based on these link groups to obtain globally perceived link data.
[0082] In one possible implementation, hierarchical data can be obtained through comprehensive monitoring indicators, which can cover the full depth of the monitored object. For example, in the field of environmental monitoring, comprehensive monitoring indicators are derived from previously fused auxiliary parameters and environmental indicators, which can be used to extract hierarchical data of air pollutants, such as the distribution of pollutant concentrations from the ground to the upper atmosphere, thereby determining the data to be extracted. This helps to reveal the vertical propagation path of pollution sources and improve the comprehensiveness of monitoring.
[0083] Specifically, principal component analysis (PCA) can be used to reduce the dimensionality of extracted data, resulting in simplified data. For example, when analyzing multi-dimensional indicators such as temperature, humidity, and wind speed for air pollutants, PCA can remove redundancy and simplify the data into a few principal components. This not only reduces computational complexity but also improves data processing efficiency and avoids information loss.
[0084] For example, a covariance matrix can be calculated based on simplified data, which represents the correlation between data points, and then an eigenvector can be obtained. In environmental monitoring, if the simplified data includes pollutant concentrations and meteorological factors, the covariance matrix can quantify the correlation strength between concentration and wind speed, while the eigenvector captures the main direction of variation. This helps to identify key influencing factors and improve the accuracy of anomaly detection.
[0085] In one possible implementation, if the number of feature vectors is greater than a preset threshold, such as a threshold of 5, then the first 3 feature vectors are retained to form a principal component set. For example, when processing 10 feature vectors, the first few with a cumulative variance contribution rate of 85% are retained. This ensures that core information is preserved while compressing the data size and optimizing the resource consumption of subsequent analysis.
[0086] Specifically, data is simplified by reconstructing principal component sets to obtain compressed data. For example, by using selected principal components, the original pollutant data can be reconstructed into a low-dimensional representation, which retains more than 95% of the original variation, effectively reduces noise interference, improves data quality, and facilitates further clustering.
[0087] For example, cluster analysis can be applied to compressed data to obtain link groups. In environmental monitoring scenarios, the K-means algorithm can be used to cluster compressed data into pollution propagation link groups, such as grouping similar concentration patterns. This reveals the dynamic propagation path of pollutants and supports precise intervention.
[0088] In one possible implementation, global sensing link data is obtained by fusing global data based on link grouping. For example, local groups are fused with global meteorological data to form a complete pollution link view. This enhances the global perspective of monitoring, helps predict large-scale environmental events, and improves response efficiency.
[0089] Step S107: The compressed global sensing link data is analyzed using a classification prediction method. Based on historical patterns, the distribution of potential risk sources is predicted to obtain an optimized layout density configuration scheme. The scheme is used to adjust the layout of the data acquisition device.
[0090] The compressed global sensing link data is processed using a classification and prediction method to obtain the distribution of potential risk sources. High-density areas are identified based on this distribution, and a priority ranking of these areas is obtained. Deployment density parameters are calculated using this priority ranking to obtain a preliminary configuration scheme. If the density parameters in the preliminary configuration scheme exceed a preset threshold, historical pattern data is integrated to adjust the parameters and determine a revised configuration scheme. A random forest algorithm is used to train the revised configuration scheme, resulting in an optimized deployment density model. Adjustment instructions are extracted from the optimized deployment density model to obtain the final layout scheme. The consistency of the link data is verified against the final layout scheme to determine its effectiveness.
[0091] In one embodiment, a classification prediction method is used to process the compressed global perception link data. First, the data is classified using the support vector machine algorithm. The principle is to map high-dimensional data to a higher-dimensional space to achieve linear separability, thereby identifying potential risk patterns.
[0092] For example, in a network monitoring system, compressed data includes link latency and traffic metrics. After classification, the distribution of risk sources can be obtained, such as the proportion of abnormal nodes reaching 30%. This helps to detect potential problems early and improve system stability.
[0093] Specifically, this method reduces the false alarm rate. By training the model with multiple labels, it can support the accuracy of risk assessment from multiple perspectives, such as combining historical intrusion data to verify the reliability of the distribution, and ultimately improve the overall perception efficiency.
[0094] In one embodiment, high-density areas are determined based on the distribution of potential risk sources. The principle is to calculate the spatial density of risk points and use kernel density estimation to identify clustered areas.
[0095] For example, the distribution shows that the risk density of a certain link segment is 0.5 per unit length. The priority ranking of high-density areas can be arranged from high to low density value. For example, the priority of the top three areas is 1, 2, and 3, which supports the targeted allocation of resources.
[0096] Specifically, the sorting process takes into account the size of the area and the severity of the risk, avoids inefficient configuration, and brings about the technical effect of optimizing monitoring coverage, such as reducing blind spots by up to 20%.
[0097] In one embodiment, the layout density parameter is calculated by sorting the regions by priority. The principle is to obtain the parameter value by weighted summation based on the sorting, and thus obtain a preliminary configuration scheme.
[0098] For example, the density parameter for priority 1 area is set to 5 sensors per kilometer, and the scheme covers the entire depth of the monitored objects. This is linked to historical data extraction to ensure the continuity of the link data.
[0099] Specifically, examining the solution from multiple perspectives, such as adjusting parameters by incorporating deep-level data, can enhance its robustness and reduce costs.
[0100] In one embodiment, if the density parameter in the initial configuration scheme exceeds a preset threshold such as 10, the parameter is adjusted by integrating historical pattern data. The principle is to introduce time series analysis to correct the deviation and determine the corrected configuration scheme.
[0101] For example, historical data showed that the peak density needed to be reduced to 8. The adjusted scheme is more realistic, which indirectly verifies the rationality of the parameters.
[0102] Specifically, this integration avoids over-deployment, and the technical benefits include improved solution adaptability and a reduction of resource waste by approximately 15%.
[0103] In one embodiment, a random forest algorithm is used to train and correct the configuration scheme. The principle is to reduce overfitting by integrating multiple decision trees to obtain an optimized layout density model.
[0104] For example, the training input includes corrected parameters and link data, and the model output density prediction accuracy reaches 85%, which supports the logical progression from the core to the extended solution.
[0105] Specifically, parallel processing of the algorithm accelerates training, improves the model's generalization ability, and is beneficial for long-term stability monitoring.
[0106] In one embodiment, adjustment instructions are extracted from the optimized layout density model. The principle is to parse the model nodes to obtain specific instructions and obtain the final layout scheme.
[0107] For example, the instructions suggest adding two sensors in high-density areas, and the final layout covers the entire chain, which is consistent with the data after dimensionality reduction.
[0108] Specifically, multi-faceted support, such as priority matching of instructions, ensures the integrity of the solution, and the technical effect is to improve response speed.
[0109] In one embodiment, the consistency of link data is verified for the final layout scheme. The principle is to compare the deviation between the scheme output and the original sensing data to determine the effectiveness of the scheme.
[0110] For example, a consistency score of 95% is considered valid, which is examined from the perspective of historical global data and supports the continuity of the overall business.
[0111] Specifically, the verification process includes multi-level checks to avoid failure scenarios, resulting in improved reliability, such as a 10% reduction in failure rate.
[0112] The above are merely preferred embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A lake water quality multi-parameter real-time monitoring method based on Internet of Things, characterized in that, The method comprises the following steps: According to the initial environmental indicators and auxiliary parameter data obtained from the data acquisition device, combined with the device position coordinates and installation depth, generate three-dimensional binding relationship data; Adopt distributed computing to clean the three-dimensional binding relationship data of the adjacent area, and if the cleaned data exceeds the preset threshold, mark it as abnormal to obtain an abnormal detection result; According to the abnormal detection result, extract a normal data subset from the multi-level network structure, and report the normal data subset to the center through the transmission link to obtain the converged regional characteristic parameters; Based on the geographical and hydrological characteristics represented by the regional characteristic parameters, use the grouping analysis method to divide the monitoring partition; analyze the change trend of the regional characteristic parameters in each partition to obtain a local dynamic perception model; If the local dynamic perception model detects parameter linkage abnormalities, analyze the influence weight of auxiliary parameters on environmental indicators through fusion learning method to obtain comprehensive monitoring indicators; Extract different levels of data from the comprehensive monitoring indicators, and use data simplification method to reduce the dimension of the data to obtain compressed global perception link data; Use classification prediction method to analyze the compressed global perception link data, predict the potential risk source distribution based on historical patterns, and obtain an optimized arrangement density configuration scheme.
2. The lake water quality multi-parameter real-time monitoring method based on Internet of Things according to claim 1, wherein The process of generating three-dimensional binding relationship data according to the initial environmental indicators and auxiliary parameter data obtained from the data acquisition device, combined with the device position coordinates and installation depth, comprises: Obtain initial environmental indicators and auxiliary parameter data from the data acquisition device to obtain a collection data set; bind the collection data set with device position coordinates and installation depth through a preset association rule to generate initial three-dimensional binding relationship data; use clustering algorithm to group process the initial three-dimensional binding relationship data to determine the spatial distribution group; analyze the change trend of pollutant concentration according to the spatial distribution group, and if the change trend exceeds the preset threshold, mark the abnormal area to obtain an abnormal marker set; obtain the position coordinates and installation depth in the abnormal marker set to generate a depth adjustment instruction; update the installation depth in the initial three-dimensional binding relationship data through the depth adjustment instruction to obtain optimized three-dimensional binding relationship data.
3. The lake water quality multi-parameter real-time monitoring method based on Internet of Things according to claim 1, wherein The process of obtaining an abnormal detection result comprises: According to the adjacent area division rule, group the three-dimensional binding relationship data by region; use distributed computing nodes to perform parallel cleaning operation on each regional data, and if there are continuous zero values or out-of-range readings caused by sensor failure in the data, mark them as invalid values to obtain a preliminary cleaning data set; obtain the numerical range of each monitoring point in the preliminary cleaning data set, and perform numerical verification on the dissolved oxygen content and pollutant concentration according to the preset threshold boundary, if the monitoring value exceeds the normal environmental parameter range, trigger the abnormal marking mechanism to determine the abnormal detection result.
4. The lake water quality multi-parameter real-time monitoring method based on Internet of Things according to claim 1, characterized in that, the process of extracting a normal data subset from the multi-level network structure according to the anomaly detection result and reporting the normal data subset to the center through a transmission link to obtain the converged regional characteristic parameters comprises: based on the anomaly detection result, a normal data subset is extracted from the multi-level network structure, and the integrity of the normal data subset is determined; according to the integrity of the normal data subset, it is reported to the center through a transmission link, and a reporting confirmation signal is obtained; if the reporting confirmation signal indicates success, a convergence index is extracted from the reported normal data subset, and the stability of the convergence index is judged; based on the stable convergence index, a regional characteristic parameter is extracted, and the consistency of the regional characteristic parameter is determined; according to the consistency of the regional characteristic parameter, a k-means clustering algorithm is used to cluster the regional characteristic parameter to obtain a clustering group; if the number of the clustering group exceeds a preset threshold, a dominant feature is selected from each group to obtain an optimized parameter set; based on the optimized parameter set, the multi-level network structure information is integrated to obtain the final regional characteristic parameter.
5. The lake water quality multi-parameter real-time monitoring method based on Internet of Things according to claim 1, characterized in that, the process of analyzing the change trend of the regional characteristic parameters in each partition to obtain a local dynamic perception model comprises: the trend sequence is generated by analyzing the change of the regional characteristic parameters of each partition; the key nodes indicating fluctuations are extracted from the trend sequence, and the spatial distribution of the key nodes is obtained; according to the spatial distribution of the key nodes, the hydrological characteristic data is integrated, and the distribution consistency is judged; based on the consistent distribution, the geographic feature data is fused to obtain a feature fusion set; a random forest algorithm is used to classify the feature fusion set to determine the evolution path of the regional environmental index; and a local dynamic perception model is constructed according to the evolution path.
6. The lake water quality multi-parameter real-time monitoring method based on Internet of Things according to claim 1, characterized in that, the process of obtaining a comprehensive monitoring index comprises: the real-time collected data is input into the local dynamic perception model to determine whether a parameter linkage anomaly occurs; if it is determined that a parameter linkage anomaly occurs, auxiliary parameters are extracted from the real-time data that trigger the anomaly to form an auxiliary parameter set; according to the auxiliary parameter set, a fusion learning method is used to calculate the influence weight of each auxiliary parameter on the environmental index to obtain a weight distribution; based on the weight distribution, the auxiliary parameters and the environmental index are weighted and fused to generate a comprehensive monitoring index.
7. The lake water quality multi-parameter real-time monitoring method based on Internet of Things according to claim 1, characterized in that, the process of extracting different level data from the comprehensive monitoring index and using a data simplification method to reduce the dimension of the data to obtain compressed global perception link data comprises: Extracting different level data covering the full depth of the monitoring object from the comprehensive monitoring index, determining a to-be-reduced dimension data set; adopting a principal component analysis method to perform dimension reduction processing on the to-be-reduced dimension data set, obtaining simplified data; calculating a covariance matrix according to the simplified data, and extracting a feature vector; retaining the first few feature vectors with the highest variance contribution rate to constitute a principal component set; reconstructing data through the principal component set to obtain compressed data; performing clustering analysis on the compressed data to obtain a link grouping; fusing the link grouping with historical monitoring data and characteristic parameters of other monitoring partitions to generate compressed global perception link data.
8. The lake water quality multi-parameter real-time monitoring method based on the Internet of Things according to claim 1, characterized in that, The process of analyzing the compressed global perception link data using a classification prediction method, predicting the potential risk source distribution based on historical patterns, and obtaining an optimized arrangement density configuration scheme includes: processing the compressed global perception link data using a classification prediction method to obtain a potential risk source distribution; determining a high-density risk area based on the potential risk source distribution and prioritizing the high-density risk area; calculating an arrangement density parameter according to the priority ranking to generate a preliminary configuration scheme; if the density parameter in the preliminary configuration scheme exceeds a preset threshold, adjusting the parameter by fusing historical pattern data to obtain a revised configuration scheme; training the revised configuration scheme using a random forest algorithm to obtain an optimized arrangement density configuration scheme.