Water quality pollutant change monitoring method and system for clustering analysis of pollutant transmission track
By arranging dense monitoring points in the water flow, combining improved clustering algorithms and hydrodynamic models, the transmission trajectory of pollutants is simulated, and the problem that traditional water quality monitoring methods cannot reflect the water condition in real time is solved, achieving efficient monitoring and analysis of water pollution.
Patent Information
- Application Number
- CN202411833581.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-13
- Publication Date
- 2025-05-23
AI Technical Summary
Traditional water quality monitoring methods rely on regular sampling and laboratory analysis, resulting in insufficient data timeliness and spatial representation, unable to effectively reflect the real-time conditions of water bodies, and it is difficult to identify pollution sources and their changing trends.
By arranging monitoring points at every 500 meters from the entrance to the outlet of the water flow, water quality monitoring data was collected, and water flow was simulated using improved K-mean clustering algorithm and two-dimensional Navier-Stokes equations, and contaminant migration was simulated by combining particle tracking method to perform clustering analysis of pollutant transmission trajectory.
Real-time monitoring and accurate analysis of water pollution conditions is achieved, and areas with severe pollution can be identified in a timely manner and early warning can be issued, which improves the efficiency of water quality protection and provides a scientific basis for environmental management and the sustainable use of water resources.
Smart Images

Figure CN120030865A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of water quality pollutant changes, and in particular to a method and system for monitoring water quality pollutant changes by cluster analysis of pollutant transmission trajectories. Background Art
[0002] Sources of water pollution include industrial wastewater, agricultural runoff pollution, and urban sewage, which cause serious harm to aquatic ecosystems. Therefore, it is particularly important to develop efficient and accurate water quality monitoring technologies and methods.
[0003] Traditional water quality monitoring relies on regular sampling and laboratory analysis. Although it can provide a certain degree of data, it is often unable to effectively reflect the real-time status of the water body due to the lack of timeliness and spatial representativeness of the samples. The limitations of this method mean that the monitoring results may lag behind the actual situation when facing a dynamically changing water quality environment. Relying on fixed sensors and static data analysis leads to insufficient real-time data acquisition and reduced accuracy, making it difficult to identify pollution sources and their changing trends in a timely manner, and can no longer meet the needs of a rapidly changing environment. In addition, existing technologies have limitations in analyzing the migration patterns of pollutants and the dynamic characteristics of water flow, and cannot fully reflect the actual distribution and behavior of pollutants in water bodies. This makes the formulation of environmental governance measures lack scientific basis, resulting in poor governance effects, and may even aggravate water pollution problems.
[0004] The prior art has the following deficiencies:
[0005] Traditional monitoring methods often rely on manual periodic sampling and laboratory analysis, which is time-consuming and labor-intensive, and is easily restricted by the timeliness and spatial representativeness of samples, resulting in monitoring results that cannot reflect the actual conditions of water bodies in a timely manner. In addition, existing monitoring technologies mostly focus on monitoring a single or a few parameters, and lack a comprehensive analysis of multi-dimensional changes in water quality. This limitation makes it difficult to fully understand the complexity and diversity of water quality changes, especially in terms of pollution source identification and impact assessment, and cannot provide comprehensive insights. The lack of in-depth analysis of dynamic changes often leads to the inability to effectively formulate response strategies, thus affecting the scientificity and effectiveness of water resources management.
[0006] The above information disclosed in this Background section is only for enhancement of understanding of the background of the present disclosure and therefore it may contain information that does not constitute the prior art that is already known to one of ordinary skill in the art. Summary of the invention
[0007] The purpose of the present invention is to provide a method and system for monitoring changes in water quality pollutants by cluster analysis of pollutant transmission trajectories, so as to solve the problems raised in the above-mentioned background technology.
[0008] To achieve the above object, the present invention provides the following technical solutions:
[0009] The method for monitoring changes in water quality pollutants by cluster analysis of pollutant transmission trajectories includes the following specific steps:
[0010] Step 1: From the inlet to the outlet of the water flow, a monitoring point is arranged every 500 meters, and water quality monitoring data of each monitoring point is collected. The water quality monitoring data includes water quality parameters, heavy metal concentration, temperature and flow rate;
[0011] Step 2: Set the sensor to collect water quality monitoring data once a minute, pre-process the water quality monitoring data and form a state vector, and store the state vector as a sample data set in matrix form;
[0012] Step 3: Randomly select k as the initial cluster centers in the state vector of the sample data set, use the improved K-means clustering algorithm to determine the number of clusters, and dynamically adjust the number of clusters based on the comparison between the cost function and the silhouette coefficient and the threshold;
[0013] Step 4: Use the two-dimensional Navier-Stokes equation to simulate water flow and establish a hydrodynamic model. Use the finite difference method to discretize the equation, simulate pollutant migration through the particle tracking method, and obtain the average pollutant concentration of each cluster.
[0014] Step 5: Compare the average pollutant concentration of each cluster with the preset threshold to identify the severely polluted areas, mark the corresponding monitoring points as severely polluted, and issue an early warning.
[0015] Furthermore, it is characterized in that the logic for collecting water quality monitoring data at each monitoring point is:
[0016] Starting from the entrance of the water flow, a monitoring point is arranged every 500 meters, and water quality parameter related sensors, heavy metal sensors, temperature sensors and flow rate sensors are deployed at each monitoring point. The water quality parameter related sensors include ammonia nitrogen sensors, total phosphorus sensors, pH sensors, dissolved oxygen sensors and turbidity sensors to collect ammonia nitrogen concentration, total phosphorus concentration, pH value, dissolved oxygen content and turbidity. The heavy metal sensors include lead sensors and mercury sensors to collect lead concentration and mercury content in the water body. The temperature sensor collects the temperature of the water body, and the flow rate sensor collects the water flow rate.
[0017] Furthermore, it is characterized in that the specific logic for preprocessing the water quality monitoring data is:
[0018] The preprocessing steps include Kalman filter denoising and normalization processing. The water quality monitoring data after Kalman filtering replaces the original collected water quality monitoring data, and then the replaced water quality monitoring data is normalized using min-max. The preprocessed ammonia nitrogen concentration, total phosphorus concentration, pH value, dissolved oxygen concentration, turbidity, lead concentration, mercury content, temperature and water flow velocity are used as state vectors, expressed as:
[0019]
[0020] Among them, X represents the state vector, NH 3 is ammonia nitrogen concentration, TP is total phosphorus concentration, pH is pH value, DO is dissolved oxygen concentration, Tu is turbidity, Pn is lead concentration, Hg is mercury content, Te is temperature, and v is water velocity;
[0021] The state vectors of all monitoring points are stored in the form of a matrix as a sample data set. Each row represents the water quality monitoring data of a monitoring point after preprocessing, which is the transpose of a state vector. The matrix D is expressed as:
[0022]
[0023] Where D represents the matrix formed by the state vector, y is the number of monitoring points, X m represents the state vector of the mth monitoring point, x mo represents the oth feature of the mth monitoring point, where m is the monitoring point index, and m=1, 2, ..., y, o is the feature index, and o=1, 2, ..., 9.
[0024] Furthermore, it is characterized in that the specific logic for dynamically adjusting the number of clusters is:
[0025] In the preprocessed state vector, k are randomly selected as the initial cluster centers, and the formula for calculating the clustering cost function is:
[0026]
[0027] Among them, J is the clustering cost function value, M a represents the ath cluster, including all points in the cluster, μ a is the center of the a-th cluster, x i ∈M a Represents data point x i is the cluster M a A point in ||x i -μ a || 2 is the data point x iThe square of the Euclidean distance to the cluster center, k is the number of initial cluster centers, a is the index of the cluster, and a∈[1, k], i is the index of the data point in the cluster, and a data point in the cluster represents a state vector;
[0028] For each data point, the logic used to calculate the silhouette coefficient is:
[0029] Calculate the data point x i The average distance to other points of the same type is based on the formula:
[0030]
[0031] Where a(i) represents the data point x i The average distance to other points of the same type, |M a | represents the number of data points in the ath cluster, d(x i , x z ) is the data point x i With x z The Euclidean distance between z ∈M a Represents data point x z Belongs to cluster M a ;
[0032] Calculate the data point x i The average distance to the points in each other cluster and the formula for obtaining the minimum average distance are as follows:
[0033]
[0034] Where b(i) represents the data point x i With each other cluster M b The minimum average distance of the inner points, x j ∈M b Represents data point x j Belongs to cluster M b , b is also the index of the cluster, and a≠b, b∈[1,k], |M b | represents the number of data points in the bth cluster, j is also the index of the data point in the cluster, and is used to represent x i and x j Do not belong to the same cluster;
[0035] Substitute a(i) and b(i) into the calculation formula of the silhouette coefficient:
[0036]
[0037] Where s(i) represents the data point x i The silhouette coefficient;
[0038] The average silhouette coefficient of all clusters is the average of the silhouette coefficients of all data points. The formula for obtaining the overall silhouette coefficient is:
[0039]
[0040] Where S is the overall silhouette coefficient, and n is the total number of data points;
[0041] Set the observation to 5 iterations. If the cost function decreases by less than 0.01 in 5 consecutive rounds or the silhouette coefficient decreases by less than 0.01 in 5 consecutive rounds, adjust the number of clusters and update the cluster centers.
[0042] Update the cluster centers using the following formula:
[0043]
[0044] Among them, μ a new represents the updated a-th cluster center, is the sum of all points in the a-th cluster.
[0045] Furthermore, it is characterized in that the specific logic for obtaining the average pollutant concentration of each cluster is:
[0046] The formula for simulating water flow using the two-dimensional Nav ier-Stokes equation is:
[0047]
[0048] in, represents the change of velocity vector v with time t, v represents the velocity vector, ρ is the density of water, p is the pressure field, σ is the kinematic viscosity, f is the volume force, represents the gradient operator;
[0049] The equations are discretized by the finite difference method to convert them into a form that can be processed by a computer to obtain velocity field, pressure field and streamline data;
[0050] The particle migration is calculated by the particle tracking method, and the particle position is updated according to the flow velocity field. The formula is:
[0051]
[0052] Among them, v(m, t) is the flow velocity of the water at the mth monitoring point at time t, and t is the time variable;
[0053] The pollutant concentration at each monitoring point is updated through the propagation model using the convection diffusion equation as follows:
[0054]
[0055] Among them, C m (g) is the concentration of the g-th pollutant at the m-th monitoring point, and D(g) is the diffusion coefficient of the g-th pollutant;
[0056] The formula used to calculate the comprehensive pollutant concentration is:
[0057]
[0058] Among them, C m is the comprehensive pollutant concentration at the mth monitoring point, ω g is the weight of the g-th pollutant, G is the number of pollutant types, and pollutants are divided into four types: chemical pollutants, physical pollutants, biological pollutants, and radioactive pollutants, and the corresponding weights are ω q ,ω 2 ,ω 3 ,ω 4 , 0<ω 4 <ω 2 <ω 3 <ω 1 <1, and ω 1 +ω 2 +ω 3 +ω 4 =1;
[0059] Here, the cluster analysis method in step 3 is implemented again, and the average pollutant concentration of each cluster is calculated based on the formula:
[0060]
[0061] in, is the average pollutant concentration in the ath cluster, N a is the number of monitoring points in the ath cluster, C m is the pollutant concentration at the mth monitoring point, It means summing up the comprehensive pollutant concentrations of the monitoring points belonging to the ath cluster.
[0062] Furthermore, it is characterized in that the specific logic for identifying the seriously polluted areas is:
[0063] The average pollutant concentration of each cluster is compared with the preset threshold. If the average pollutant concentration of the cluster is lower than the preset threshold, the monitoring points of the cluster will continue to be monitored. If the average pollutant concentration of the cluster is higher than the preset threshold, the monitoring points in the cluster will be marked as severely polluted areas and an early warning will be issued.
[0064] The present invention further provides a water quality pollutant change monitoring system for clustering analysis of pollutant transport trajectories, which is used to implement the water quality pollutant change monitoring method for clustering analysis of pollutant transport trajectories as described above, and includes:
[0065] A monitoring point arrangement and data acquisition module, which is used to arrange a monitoring point every 500 meters from the inlet to the outlet of the water flow, and collect water quality monitoring data at each monitoring point. The water quality monitoring data includes water quality parameters, heavy metal concentrations, temperature, and flow velocity;
[0066] A data preprocessing module, which is used to set the sensor to collect water quality monitoring data once every minute, preprocess the water quality monitoring data to form a state vector, and store the state vector as a sample data set in matrix form;
[0067] A clustering center selection and adjustment module, which is used to randomly select k from the state vectors of the sample data set as the initial clustering centers, determine the number of clusters using an improved K-means clustering algorithm, and dynamically adjust the number of clusters based on the comparison of the cost function and the silhouette coefficient with the threshold;
[0068] A hydrodynamic model construction module, which is used to simulate the water flow using the two-dimensional Navier-Stokes equation to establish a hydrodynamic model, discretize the equation using the finite difference method, and simulate the pollutant migration through the particle tracking method to obtain the average pollutant concentration of each cluster;
[0069] A pollution identification and warning module, which is used to compare the average pollutant concentration of each cluster with a preset threshold, identify the areas with serious pollution, mark the corresponding monitoring points as seriously polluted, and issue a warning.
[0070] In the above technical solution, the technical effects and advantages provided by the present invention are:
[0071] By monitoring water quality parameters in real time, the ability to understand and respond to water pollution has been significantly improved. First, the Kalman filter algorithm is used to denoise and normalize sensor data, making the monitoring data more accurate and reliable. This high-precision data processing capability ensures that actual water quality changes can be captured in a timely manner, thus providing a solid data basis for decision-making. Secondly, the improved K-means clustering algorithm dynamically adjusts the number of clusters and optimizes the clustering results based on the cost function and silhouette coefficient, so that the pollution characteristics of different waters can be more accurately identified. In this way, areas with similar pollutant concentrations can be effectively classified into one category, which is convenient for subsequent analysis and management. At the same time, the two-dimensional Navier-Stokes equation simulates water flow, combined with the particle tracking method to calculate pollutant migration, providing a deep insight into the dynamic changes of pollutants. This efficient and accurate monitoring mechanism greatly improves the efficiency of water quality protection and provides a scientific basis for environmental management and sustainable use of water resources. BRIEF DESCRIPTION OF THE DRAWINGS
[0072] Figure 1 It is a schematic diagram of the overall method flow of the present invention;
[0073] Figure 2 Schematic diagram of the system structure of the present invention. DETAILED DESCRIPTION
[0074] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with specific embodiments.
[0075] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in the present invention should be understood by people with ordinary skills in the field to which the present invention belongs. The words "first", "second" and similar words used in the present invention do not indicate any order, quantity or importance, but are only used to distinguish different components. "Include" or "comprise" and similar words mean that the elements or objects appearing before the word include the elements or objects listed after the word and their equivalents, without excluding other elements or objects. "Connect" or "connected" and similar words are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. "Up", "down", "left", "right" and the like are only used to indicate relative positional relationships. When the absolute position of the described object changes, the relative positional relationship may also change accordingly.
[0076] Example:
[0077] See also Figure 1 , the present invention provides a technical solution:
[0078] The method for monitoring changes in water quality pollutants by cluster analysis of pollutant transmission trajectories includes the following specific steps:
[0079] Step 1: From the inlet to the outlet of the water flow, a monitoring point is arranged every 500 meters, and water quality monitoring data of each monitoring point is collected. The water quality monitoring data includes water quality parameters, heavy metal concentration, temperature and flow rate;
[0080] In this embodiment, it is characterized in that the logic for collecting water quality monitoring data at each monitoring point is:
[0081] Starting from the entrance of the water flow, a monitoring point is arranged every 500 meters, and water quality parameter related sensors, heavy metal sensors, temperature sensors and flow rate sensors are deployed at each monitoring point. The water quality parameter related sensors include ammonia nitrogen sensors, total phosphorus sensors, pH sensors, dissolved oxygen sensors and turbidity sensors to collect ammonia nitrogen concentration, total phosphorus concentration, pH value, dissolved oxygen content and turbidity. The heavy metal sensors include lead sensors and mercury sensors to collect lead concentration and mercury content in the water body. The temperature sensor collects the temperature of the water body, and the flow rate sensor collects the water flow rate.
[0082] In the cluster analysis of pollutant transmission trajectories, the implementation of the step of arranging monitoring points every 500 meters from the inlet to the outlet of the water flow has many important significances. First, this arrangement can ensure comprehensive coverage of the water body, so that water quality changes at different locations and different time periods can be captured in time. By setting a fixed interval, sufficient spatial data can be obtained to help analyze the transmission path and concentration changes of water quality pollutants, thereby revealing the characteristics of the pollution source and the dynamic behavior of pollutants in the water flow. Secondly, the various water quality parameters collected by each monitoring point, such as ammonia nitrogen, total phosphorus, pH value, dissolved oxygen and turbidity, combined with heavy metal concentration, temperature and flow rate, constitute a comprehensive water quality data set. This multi-dimensional data collection can provide a rich information basis for cluster analysis, thereby improving the understanding and accuracy of water quality changes. Especially in complex water environments, one or a few parameters alone often cannot fully reflect the true state of water quality, and multi-parameter monitoring can comprehensively consider the interactions and influences between different pollutants, providing a more reliable basis for subsequent analysis.
[0083] In addition, the monitoring of temperature and flow rate is also crucial because they have a direct impact on the physical and chemical properties of water quality. Changes in water flow velocity are related to the diffusion and deposition of pollutants, while temperature affects the metabolism and chemical reaction rates of aquatic organisms. This comprehensive monitoring can help analyze the migration and transformation processes of pollutants under different water flow conditions, thereby providing a scientific basis for pollution control and water resources management. In summary, this monitoring step not only provides solid data support for subsequent cluster analysis, but also helps to comprehensively and systematically evaluate the status of water quality and its changes, providing key support for accurate pollution source identification and effective water quality management.
[0084] Step 2: Set the sensor to collect water quality monitoring data once a minute, pre-process the water quality monitoring data and form a state vector, and store the state vector as a sample data set in matrix form;
[0085] In this embodiment, it is characterized in that the specific logic for preprocessing the water quality monitoring data is:
[0086] The preprocessing steps include Kalman filter denoising and normalization. In the Kalman filter process, the state model of the system needs to be established first to describe the changes of the system over time. The system state is represented as a state vector, which contains the current state information of the system, that is, the state vector constructed by the original collected water quality monitoring data. The first step of filtering is prediction. In this stage, the state at the current moment is predicted by the state estimate of the previous moment and the known system dynamic model. Next, the system will obtain observation data from the sensor, which is usually noisy. The second step of Kalman filtering is updating. In this stage, the predicted state and the actual measurement value are combined to adjust the state estimate by calculating the Kalman gain. The Kalman gain reflects the degree of trust between the predicted result and the measurement result. The higher the gain, the more trusted the measurement result. During the update process, the filter will correct the previous estimate based on the measurement result, and finally obtain a more accurate current state estimate. This process will be iterated every time new data is obtained. Over time, the Kalman filter can effectively suppress the measurement noise, thereby providing more accurate and reliable water quality monitoring data. This dynamic estimation method makes Kalman filtering very useful in real-time monitoring and data analysis.
[0087] The water quality monitoring data after Kalman filtering replaces the original collected water quality monitoring data, and then the replaced water quality monitoring data is normalized using mmin-max. The pre-treated ammonia nitrogen concentration, total phosphorus concentration, pH value, dissolved oxygen concentration, turbidity, lead concentration, mercury content, temperature and water flow velocity are used as state vectors, expressed as:
[0088]
[0089] Among them, X represents the state vector, NH 3 is ammonia nitrogen concentration, TP is total phosphorus concentration, pH is pH value, DO is dissolved oxygen concentration, Tu is turbidity, Pb is lead concentration, Hg is mercury content, Te is temperature, and v is water velocity;
[0090] The state vectors of all monitoring points are stored in the form of a matrix as a sample data set. Each row represents the water quality monitoring data of a monitoring point after preprocessing, which is the transpose of a state vector. The matrix D is expressed as:
[0091]
[0092] Where D represents the matrix formed by the state vector, y is the number of monitoring points, X m represents the state vector of the mth monitoring point, x mo represents the oth feature of the mth monitoring point, where m is the monitoring point index, and m=1, 2, ..., y, o is the feature index, and o=1, 2, ..., 9.
[0093] Kalman filtering is an effective recursive filtering algorithm that can predict the system state in a noisy measurement environment by using previous estimates and current observations to minimize the mean square error. This is particularly important for water quality monitoring data, because the collection of water quality parameters is often affected by environmental noise, equipment errors and other uncertain factors, and the direct use of raw data may lead to inaccurate and misleading analysis results. Secondly, through Kalman filtering denoising, the signal-to-noise ratio of the data can be significantly improved, thereby enhancing the reliability of subsequent data analysis. The filtered data can more accurately reflect the real changes in water quality and provide more realistic information on pollutant concentrations and water body status. This is crucial for cluster analysis, because inaccurate data may lead to erroneous clustering results, which will affect the identification of pollution sources and the formulation of governance strategies. The dynamic characteristics of Kalman filtering enable it to adapt to the timeliness of water quality changes. As water flow and environmental conditions change, water quality parameters may also change. Kalman filtering can update state estimates in real time to ensure a rapid response to water quality changes during monitoring. This real-time nature not only improves the accuracy of monitoring, but also provides the information needed for decision-making in a timely manner when sudden pollution incidents occur, helping managers to respond quickly and effectively.
[0094] After denoising and normalization, water quality monitoring data can be stored in a more standardized form, which is convenient for subsequent data integration and analysis. Data normalization is crucial for cluster analysis because it ensures that data from different monitoring points are compared on the same scale, thereby improving the effectiveness and stability of the clustering algorithm.
[0095] In summary, the use of Kalman filtering for denoising not only improves the accuracy and reliability of water quality monitoring data, but also enhances the adaptability of the monitoring system to dynamic changes, and provides a solid data foundation for the cluster analysis of pollutant transmission trajectories, thereby making water quality management and pollution control more effective.
[0096] Step 3: Randomly select k as the initial cluster centers in the state vector of the sample data set, use the improved K-means clustering algorithm to determine the number of clusters, and dynamically adjust the number of clusters based on the comparison between the cost function and the silhouette coefficient and the threshold;
[0097] In this embodiment, the specific logic for dynamically adjusting the number of clusters is as follows:
[0098] In the preprocessed state vector, k are randomly selected as the initial cluster centers, and the formula for calculating the clustering cost function is:
[0099]
[0100] Among them, J is the clustering cost function value, M a represents the ath cluster, including all points in the cluster, μ a is the center of the a-th cluster, x i ∈M a Represents data point x i is the cluster M a A point in ||x i -μ a || 2 is the data point x i The square of the Euclidean distance to the cluster center, k is the number of initial cluster centers, a is the index of the cluster, and a∈[1, k], i is the index of the data point in the cluster, and a data point in the cluster represents a state vector;
[0101] For each data point, the logic used to calculate the silhouette coefficient is:
[0102] Calculate the data point x i The average distance to other points of the same type is based on the formula:
[0103]
[0104] Where a(i) represents the data point x i The average distance to other points of the same type, |M a | represents the number of data points in the ath cluster, d(x i , x z ) is the data point x i With x z The Euclidean distance betweenz ∈M a Represents data point x z Belongs to cluster M a ;
[0105] Calculate the data point x i The average distance to the points in each other cluster and the formula for obtaining the minimum average distance are as follows:
[0106]
[0107] Where b(i) represents the data point x i With each other cluster M b The minimum average distance of the inner points, x j ∈M b Represents data point x j Belongs to cluster M b , b is also the index of the cluster, and a≠b, b∈[1,k], |M b | represents the number of data points in the bth cluster, j is also the index of the data point in the cluster, and is used to represent x i and x j Do not belong to the same cluster;
[0108] Substitute a(i) and b(i) into the calculation formula of the silhouette coefficient:
[0109]
[0110] Where s(i) represents the data point x i The silhouette coefficient;
[0111] The average silhouette coefficient of all clusters is the average of the silhouette coefficients of all data points. The formula for obtaining the overall silhouette coefficient is:
[0112]
[0113] Where S is the overall silhouette coefficient, and n is the total number of data points;
[0114] Set the observation to 5 iterations. If the cost function decreases by less than 0.01 in 5 consecutive rounds or the silhouette coefficient decreases by less than 0.01 in 5 consecutive rounds, adjust the number of clusters and update the cluster centers.
[0115] Update the cluster centers using the following formula:
[0116]
[0117] Among them, μ a new represents the updated a-th cluster center, is the sum of all points in the a-th cluster.
[0118] The quality of clustering can be effectively evaluated by randomly selecting the initial cluster center and calculating the cost function. Dynamically adjusting the number of clusters helps to find the optimal cluster structure and ensure that the clustering results can more accurately reflect the actual distribution and changes of water quality pollutants. This method enables the clustering algorithm to adapt to different data characteristics and capture potential pollution patterns. Combining the evaluation indicators of cost function and silhouette coefficient, the effectiveness of clustering can be measured in multiple dimensions. The cost function focuses on the distance between the data point and the cluster center, reflecting the compactness of the cluster, while the silhouette coefficient considers the relative distance between the data point and the same and different clusters, which can more comprehensively evaluate the separation of clusters. This comprehensive evaluation mechanism can improve the robustness of the clustering results. Setting a strategy of observing 5 iterations can dynamically adjust the number of clusters according to the convergence of the algorithm. When the change range of the cost function and the silhouette coefficient is less than the set threshold, it means that the improvement of clustering is limited, and the cluster center can be updated in time to avoid unnecessary calculations and improve the efficiency of the algorithm. This dynamic adaptability ensures that the clustering process can always run in the best state. Using the mean update cluster center formula can ensure that after each update, the cluster center can more accurately reflect the data distribution of the cluster. This mean update method based on all points in the current cluster helps to reduce the sensitivity of the cluster center to outliers, thereby improving the stability and accuracy of clustering. Through the above steps, the final clustering results can better depict the trajectory of pollutant transmission and change. This provides a scientific basis for the identification of pollution sources, the prediction of pollutant diffusion, and the formulation of water quality management strategies, which helps to achieve more effective water resource protection and pollution control.
[0119] Step 4: Use the two-dimensional Navier-Stokes equation to simulate water flow and establish a hydrodynamic model. Use the finite difference method to discretize the equation, simulate pollutant migration through the particle tracking method, and obtain the average pollutant concentration of each cluster.
[0120] In this embodiment, it is characterized in that the specific logic for obtaining the average pollutant concentration of each cluster is:
[0121] The formula for simulating water flow using the two-dimensional Navier-Stokes equation is:
[0122]
[0123] in, represents the change of velocity vector v with time t, v represents the velocity vector, ρ is the density of water, p is the pressure field, σ is the kinematic viscosity, f is the volume force, represents the gradient operator;
[0124] The equations are discretized by the finite difference method to convert them into a form that can be processed by a computer to obtain velocity field, pressure field and streamline data;
[0125] The particle migration is calculated by the particle tracking method, and the particle position is updated according to the flow velocity field. The formula is:
[0126]
[0127] Among them, v(m, t) is the flow velocity of the water at the mth monitoring point at time t, and t is the time variable;
[0128] The pollutant concentration at each monitoring point is updated through the propagation model using the convection diffusion equation as follows:
[0129]
[0130] Among them, C m (g) is the concentration of the g-th pollutant at the m-th monitoring point, and D(g) is the diffusion coefficient of the g-th pollutant;
[0131] The formula used to calculate the comprehensive pollutant concentration is:
[0132]
[0133] Among them, C m is the comprehensive pollutant concentration at the mth monitoring point, ω g is the weight of the g-th pollutant, G is the number of pollutant types, and pollutants are divided into four types: chemical pollutants, physical pollutants, biological pollutants, and radioactive pollutants, and the corresponding weights are ω 1 ,ω 2 ,ω 3 ,ω 4 , 0<ω 4 <ω 2 <ω 3 <ω 1 <1, and ω 1 +ω 2 +ω 3 +ω 4 =1;
[0134] Chemical pollutants are usually the direct result of human activities and have the most significant impact on water quality. For example, heavy metals and organic pollutants can accumulate in organisms and have long-term effects on ecosystems and human health. Their concentration changes often directly reflect the degree of water pollution, so chemical pollutants have the highest weight. Biological pollutants such as pathogens (bacteria, viruses, etc.) pose a direct threat to human health, especially in drinking water sources, which may cause waterborne diseases. Because people have high requirements for the safety of water quality, the weight of biological pollutants should also be relatively high, second only to chemical pollutants. Although physical pollutants such as suspended solids also affect water quality and ecology, their impacts often appear in the short term and are relatively easy to remove through physical methods such as sedimentation and filtration. Therefore, the weight of physical pollutants is relatively low. The occurrence of radioactive pollutants is usually rare, and their concentration levels are relatively low in most natural water bodies. Although they may have serious health effects, under normal circumstances, the frequency and demand for their monitoring and management are relatively low, so their weight is the lowest. ω 1 +ω 2 +ω 3 +ω 4 = 1. This constraint ensures that the comprehensive pollutant concentration is a normalized indicator, ensuring that the total contribution of its weighted combination is 100%.
[0135] Here, the cluster analysis method in step 3 is implemented again, and the average pollutant concentration of each cluster is calculated based on the formula:
[0136]
[0137] in, is the average pollutant concentration in the ath cluster, N a is the number of monitoring points in the ath cluster, C m is the pollutant concentration at the mth monitoring point, It means summing up the comprehensive pollutant concentrations of the monitoring points belonging to the ath cluster.
[0138] In this scheme, the density, kinematic viscosity and body force of water are regarded as fixed values, and the density, kinematic viscosity and body force of water at normal temperature (about 20°C) and normal pressure are used. The density of water is usually 1000kg / m 3 , kinematic viscosity is 1.0×10 -6 m 2 / s, body force usually goes to 9.81m / s 2 .
[0139] The two-dimensional Navier-Stokes equations can be used to scientifically and accurately simulate the movement of fluids, taking into account flow rate changes, pressure distribution, and physical properties (such as water density and viscosity). This simulation provides a basis for understanding the spread of pollutants in water bodies, making the results more physically realistic. Therefore, under the influence of water flow dynamics, the migration trajectory of pollutants can be better predicted. Through the particle tracking method, the position of pollutants in the water flow can be updated in real time. This method provides a clearer dynamic view of the migration and diffusion process of pollutants, which can reflect how pollutants migrate in water bodies at different flow rates and flow directions. This feature is of great significance for pollutant source identification and pollution diffusion control. Through the application of the convection-diffusion equation, the migration characteristics of different types of pollutants are integrated, so that the calculated pollutant concentration is not just the result of a single pollutant, but a comprehensive consideration of the impact of multiple pollutants. This comprehensive analysis can provide a more comprehensive water quality assessment and help managers formulate governance measures more effectively. After completing the calculation of the pollutant concentration, combined with the previous cluster analysis steps, the average pollutant concentration can be calculated for each cluster. This process ensures that each cluster not only reflects the aggregation of monitoring points, but also provides specific quantitative information about the pollution status of the area, providing data support for further environmental decision-making.
[0140] Step 5: Compare the average pollutant concentration of each cluster with the preset threshold to identify the severely polluted areas, mark the corresponding monitoring points as severely polluted, and issue an early warning.
[0141] In this embodiment, the specific logic for identifying the seriously polluted area is as follows:
[0142] The average pollutant concentration of each cluster is compared with the preset threshold. If the average pollutant concentration of the cluster is lower than the preset threshold, the monitoring points of the cluster will continue to be monitored. If the average pollutant concentration of the cluster is higher than the preset threshold, the monitoring points in the cluster will be marked as severely polluted areas and an early warning will be issued.
[0143] By comparing the average pollutant concentration of each cluster with the preset threshold, the areas with more serious pollution can be quickly identified. This process ensures that the pollution source and the affected area can be discovered at the first time, which helps the relevant departments to take timely measures to deal with potential environmental hazards. The value of the threshold can be quantified based on the average of the average pollutant concentration in the seriously polluted areas in the historical data. Through the comparison of the quantitative threshold, scientific and objective data support can be provided for environmental management and decision-making. This method makes the decision-making process more transparent, reduces the possibility of relying on subjective judgment, and ensures the effectiveness and pertinence of the governance measures.
[0144] See also Figure 2The present invention further provides a water quality pollutant change monitoring system based on cluster analysis of pollutant transmission trajectories, wherein the water quality pollutant change monitoring system based on cluster analysis of pollutant transmission trajectories is used to implement the water quality pollutant change monitoring method based on cluster analysis of pollutant transmission trajectories, comprising:
[0145] The monitoring point arrangement and data collection module is used to arrange a monitoring point every 500 meters from the inlet to the outlet of the water flow, and collect water quality monitoring data at each monitoring point, wherein the water quality monitoring data includes water quality parameters, heavy metal concentration, temperature and flow rate;
[0146] A data preprocessing module is used to set the sensor to collect water quality monitoring data once a minute, preprocess the water quality monitoring data and form a state vector, and store the state vector as a sample data set in matrix form;
[0147] The cluster center selection and adjustment module is used to randomly select k as the initial cluster centers in the state vector of the sample data set, use the improved K-means clustering algorithm to determine the number of clusters, and dynamically adjust the number of clusters based on the comparison between the cost function and the silhouette coefficient and the threshold;
[0148] The hydrodynamic model building module is used to simulate the water flow using the two-dimensional Navier-Stokes equation to establish a hydrodynamic model, discretize the equation using the finite difference method, simulate the pollutant migration through the particle tracking method, and obtain the average pollutant concentration of each cluster;
[0149] The pollution identification and early warning module is used to compare the average pollutant concentration of each cluster with the preset threshold, identify the severely polluted areas, mark the corresponding monitoring points as severely polluted, and issue an early warning.
[0150] The above formulas are all dimensionless and numerical calculations. The formula is a formula for the most recent real situation obtained by collecting a large amount of data and performing software simulation. The preset parameters in the formula are set by technicians in this field according to actual conditions.
[0151] The above embodiments may be implemented in whole or in part by software, hardware, firmware or any other combination thereof. When implemented by software, the above embodiments may be implemented in whole or in part in the form of a computer program product. Those skilled in the art may appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein may be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed by hardware or software methods depends on the specific application and design constraints of the technical solution.
[0152] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, and may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0153] The above description is only a specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any technician familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application.
Claims
1. A method for monitoring changes in water quality pollutants based on cluster analysis of pollutant transmission trajectories, characterized in that: The specific steps include: Step 1: From the inlet to the outlet of the water flow, a monitoring point is arranged every 500 meters, and water quality monitoring data of each monitoring point is collected. The water quality monitoring data includes water quality parameters, heavy metal concentration, temperature and flow rate; Step 2: Set the sensor to collect water quality monitoring data once a minute, pre-process the water quality monitoring data and form a state vector, and store the state vector as a sample data set in matrix form; Step 3: Randomly select k as the initial cluster centers in the state vector of the sample data set, use the improved K-means clustering algorithm to determine the number of clusters, and dynamically adjust the number of clusters based on the comparison between the cost function and the silhouette coefficient and the threshold; Step 4: Use the two-dimensional Navier-Stokes equation to simulate water flow and establish a hydrodynamic model. Use the finite difference method to discretize the equation, simulate pollutant migration through the particle tracking method, and obtain the average pollutant concentration of each cluster. Step 5: Compare the average pollutant concentration of each cluster with the preset threshold to identify the severely polluted areas, mark the corresponding monitoring points as severely polluted, and issue an early warning.
2. The method for monitoring changes in water quality pollutants based on cluster analysis of pollutant transmission trajectories according to claim 1 is characterized in that: The logic for collecting water quality monitoring data at each monitoring point is: Starting from the entrance of the water flow, a monitoring point is arranged every 500 meters, and water quality parameter related sensors, heavy metal sensors, temperature sensors and flow rate sensors are deployed at each monitoring point. The water quality parameter related sensors include ammonia nitrogen sensors, total phosphorus sensors, pH sensors, dissolved oxygen sensors and turbidity sensors to collect ammonia nitrogen concentration, total phosphorus concentration, pH value, dissolved oxygen content and turbidity. The heavy metal sensors include lead sensors and mercury sensors to collect lead concentration and mercury content in the water body. The temperature sensor collects the temperature of the water body, and the flow rate sensor collects the water flow rate.
3. The method for monitoring changes in water quality pollutants based on cluster analysis of pollutant transmission trajectories according to claim 2 is characterized in that: The specific logic for preprocessing water quality monitoring data is as follows: The preprocessing steps include Kalman filter denoising and normalization processing. The water quality monitoring data after Kalman filtering replaces the original collected water quality monitoring data, and then the replaced water quality monitoring data is normalized using min-max. The preprocessed ammonia nitrogen concentration, total phosphorus concentration, pH value, dissolved oxygen concentration, turbidity, lead concentration, mercury content, temperature and water flow velocity are used as state vectors, expressed as: Where X represents the state vector, NH3 is the ammonia nitrogen concentration, TP is the total phosphorus concentration, pH is the pH value, DO is the dissolved oxygen concentration, Tu is the turbidity, Pb is the lead concentration, Hg is the mercury content, Te is the temperature, and v is the water flow velocity; The state vectors of all monitoring points are stored in the form of a matrix as a sample data set. Each row represents the water quality monitoring data of a monitoring point after preprocessing, which is the transpose of a state vector. The matrix D is expressed as: Where D represents the matrix formed by the state vector, y is the number of monitoring points, X m represents the state vector of the mth monitoring point, x mo Represents the oth feature of the mth monitoring point, where m is the monitoring point index, and m=1, 2, …, y, and o is the feature index, and o=1, 2, …, 9.
4. The method for monitoring changes in water quality pollutants based on cluster analysis of pollutant transmission trajectories according to claim 3 is characterized in that: The specific logic for dynamically adjusting the number of clusters is as follows: In the preprocessed state vector, k are randomly selected as the initial cluster centers, and the formula for calculating the clustering cost function is: Among them, J is the clustering cost function value, M a represents the ath cluster, including all points in the cluster, μ a is the center of the a-th cluster, x i ∈M a Represents data point x i is the cluster M a A point in ||x i -μ a || 2 is the data point x i The square of the Euclidean distance to the cluster center, k is the number of initial cluster centers, a is the index of the cluster, and a∈[1,k], i is the index of the data point in the cluster, and a data point in the cluster represents a state vector; For each data point, the logic used to calculate the silhouette coefficient is: Calculate the data point x i The average distance to other points of the same type is based on the formula: Where a(i) represents the data point x i The average distance to other points of the same type, |M a | represents the number of data points in the ath cluster, d(x i ,x z ) is the data point x i With x z The Euclidean distance between z ∈M a Represents data point x z Belongs to cluster M a ; Calculate the data point x i The average distance to the points in each other cluster and the formula for obtaining the minimum average distance are as follows: Where b(i) represents the data point x i With each other cluster M b The minimum average distance of the inner points, x j ∈M b Represents data point x j Belongs to cluster M b , b is also the index of the cluster, and a≠b, b∈[1,k], |M b | represents the number of data points in the bth cluster, j is also the index of the data point in the cluster, and is used to represent x i and x j Do not belong to the same cluster; Substitute a(i) and b(i) into the calculation formula of the silhouette coefficient: Where s(i) represents the data point x i The silhouette coefficient; The average silhouette coefficient of all clusters is the average of the silhouette coefficients of all data points. The formula for obtaining the overall silhouette coefficient is: Where S is the overall silhouette coefficient, and n is the total number of data points; Set the observation to 5 iterations. If the cost function decreases by less than 0.01 in 5 consecutive rounds or the silhouette coefficient decreases by less than 0.01 in 5 consecutive rounds, adjust the number of clusters and update the cluster centers. Update the cluster centers using the following formula: Among them, μ a new represents the updated a-th cluster center, is the sum of all points in the a-th cluster.
5. The method for monitoring changes in water quality pollutants based on cluster analysis of pollutant transmission trajectories according to claim 4 is characterized in that: The specific logic for obtaining the average pollutant concentration of each cluster is as follows: The formula for simulating water flow using the two-dimensional Navier-Stokes equation is: in, represents the change of velocity vector v with time t, v represents the velocity vector, ρ is the density of water, p is the pressure field, σ is the kinematic viscosity, f is the volume force, represents the gradient operator; The equations are discretized by the finite difference method to convert them into a form that can be processed by a computer to obtain velocity field, pressure field and streamline data; The particle migration is calculated by the particle tracking method, and the particle position is updated according to the flow velocity field. The formula is: Among them, v(m,t) is the flow velocity of the water at the mth monitoring point at time t, and t is the time variable; The pollutant concentration at each monitoring point is updated through the diffusion model using the convection diffusion equation as follows: Among them, C m (g) is the concentration of the g-th pollutant at the m-th monitoring point, and D(g) is the diffusion coefficient of the g-th pollutant; The formula used to calculate the comprehensive pollutant concentration is: Among them, C m is the comprehensive pollutant concentration at the mth monitoring point, ω g is the weight of the g-th pollutant, G is the number of pollutant types, and pollutants are divided into four types, namely chemical pollutants, physical pollutants, biological pollutants and radioactive pollutants, with corresponding weights of ω1, ω2, ω3, ω4, 0<ω4<ω2<ω3<ω1<1, and ω1+ω2+ω3+ω4=1; Here, the cluster analysis method in step 3 is implemented again, and the average pollutant concentration of each cluster is calculated based on the formula: in, is the average pollutant concentration in the ath cluster, N a is the number of monitoring points in the ath cluster, C m is the pollutant concentration at the mth monitoring point, It means summing up the comprehensive pollutant concentrations of the monitoring points belonging to the ath cluster.
6. The method for monitoring changes in water quality pollutants based on cluster analysis of pollutant transmission trajectories according to claim 5, characterized in that: The specific logic used to identify the heavily polluted areas is: The average pollutant concentration of each cluster is compared with the preset threshold. If the average pollutant concentration of the cluster is lower than the preset threshold, the monitoring points of the cluster will continue to be monitored. If the average pollutant concentration of the cluster is higher than the preset threshold, the monitoring points in the cluster will be marked as severely polluted areas and an early warning will be issued.
7. A water quality pollutant change monitoring system based on cluster analysis of pollutant transmission trajectories, characterized in that: The water quality pollutant change monitoring system based on cluster analysis of pollutant transmission trajectories is used to implement the water quality pollutant change monitoring method based on cluster analysis of pollutant transmission trajectories according to any one of claims 1 to 6, comprising: The monitoring point arrangement and data collection module is used to arrange a monitoring point every 500 meters from the inlet to the outlet of the water flow, and collect water quality monitoring data at each monitoring point, wherein the water quality monitoring data includes water quality parameters, heavy metal concentration, temperature and flow rate; A data preprocessing module is used to set the sensor to collect water quality monitoring data once a minute, preprocess the water quality monitoring data and form a state vector, and store the state vector as a sample data set in matrix form; The cluster center selection and adjustment module is used to randomly select k as the initial cluster centers in the state vector of the sample data set, use the improved K-means clustering algorithm to determine the number of clusters, and dynamically adjust the number of clusters based on the comparison between the cost function and the silhouette coefficient and the threshold; The hydrodynamic model building module is used to simulate the water flow using the two-dimensional Navier-Stokes equation to establish a hydrodynamic model, discretize the equation using the finite difference method, simulate the pollutant migration through the particle tracking method, and obtain the average pollutant concentration of each cluster; The pollution identification and early warning module is used to compare the average pollutant concentration of each cluster with the preset threshold, identify the severely polluted areas, mark the corresponding monitoring points as severely polluted, and issue an early warning.
Citation Information
Patent Citations
Device and method for environmental monitoring
CN113935394A
Water quality pollution monitoring method for seasonal drought river and computer program product
CN118503890A
Volatile organic compound monitoring method, system, medium and equipment
CN118671224A
Method for detecting underground water and soil coupling pollution in ocean engineering field under tidal action
CN119001047A