A watershed ecological anomaly monitoring method based on clustering processing
By introducing the time weight coefficient and iteratively adjusting the radius to optimize the cluster center, and using weighted kernel K-means clustering and local outlier factor detection, the problems of poor monitoring accuracy and effect in watershed ecological anomaly monitoring are solved, and efficient identification and reliable monitoring of sudden anomalies are achieved.
Patent Information
- Application Number
- CN202510907946.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-02
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2045-07-02
AI Technical Summary
Existing watershed ecological anomaly monitoring methods ignore the changing characteristics of ecological data in the temporal dimension, making it difficult to distinguish real anomalies from noise, resulting in poor monitoring accuracy. In addition, anomaly monitoring is often ignored among many normal monitoring activities, resulting in poor results.
The time weight coefficient is introduced to construct the time fusion feature vector. The cluster center selection is optimized by iteratively adjusting the radius and introducing the time variation factor. The weighted kernel K-means clustering and local outlier factor detection are used to enhance anomaly monitoring.
It improves the sensitivity and monitoring accuracy of sudden abnormal situations, enhances the recognition probability and effect of abnormal monitoring, prevents excessive attention to minor ecological variables, and enhances the reliability of monitoring.
Smart Images

Figure CN120449054B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of ecological monitoring, and in particular to a watershed ecological anomaly monitoring method based on clustering processing. Background Art
[0002] Watershed ecological anomaly monitoring methods are a series of technologies and methods used to identify deviations from normal conditions in watershed ecosystems. Their purpose is to promptly detect abnormal changes in ecosystems so that appropriate protection and restoration measures can be taken. However, common watershed ecological anomaly monitoring methods ignore the temporal variations in ecological data, poorly capture sudden anomalies, and have difficulty distinguishing true anomalies from noise, resulting in poor anomaly monitoring accuracy. Common watershed ecological anomaly monitoring methods also suffer from the problem of anomaly monitoring being overlooked amidst numerous normal monitoring events and insufficient differentiation of the contributions of different indicators, leading to poor anomaly monitoring effectiveness. Summary of the Invention
[0003] In view of the above situation, in order to overcome the defects of the existing technology, the present invention provides a watershed ecological anomaly monitoring method based on clustering processing. In view of the fact that the general watershed ecological anomaly monitoring method ignores the changing characteristics of ecological data in the time dimension, improperly captures sudden anomalies, and is difficult to distinguish between real anomalies and noise, which leads to poor accuracy of anomaly monitoring, this solution is based on watershed ecological data, introduces a time weight coefficient to construct a time fusion feature vector, highlights the changes in recent ecological data, reduces the interference of long-term historical noise, and improves the sensitivity to sudden situations; by iteratively adjusting the radius, it automatically adapts to the sampling density under different river sections and seasonal changes in flow, and improves monitoring accuracy; introduces a time change factor to construct a Composite indicators select core reference points, objectively emphasize the most representative core positions, and provide reliability for subsequent abnormal monitoring. In view of the problem that abnormal monitoring is ignored in many normal monitoring methods in general watershed ecological abnormal monitoring methods, and the contribution of different indicators is not distinguished enough, which leads to poor abnormal monitoring effect, this scheme introduces sample weights to strengthen abnormal monitoring and increase the probability of abnormal monitoring being identified. Through the guidance of core reference points, the degree of attention paid to different indicators is continuously adjusted, the influence of important ecological variables is enhanced, and excessive attention to secondary ecological variables is prevented. The weights are updated based on the stability of indicators, the contribution of different indicators to abnormal monitoring is distinguished, and the importance of key indicators is highlighted, thereby improving the abnormal monitoring effect.
[0004] The technical solution adopted by the present invention is as follows: The present invention provides a watershed ecological anomaly monitoring method based on clustering processing, which includes the following steps:
[0005] Step S1: data collection;
[0006] Step S2: watershed ecological mapping;
[0007] Step S3: radius optimization;
[0008] Step S4: initial cluster center selection;
[0009] Step S5: clustering processing of watershed ecological data;
[0010] Step S6: Monitoring of watershed ecological anomalies.
[0011] Furthermore, in step S1, the data collection is to collect watershed ecological data; and each dimension of the watershed ecological data is normalized to the maximum and minimum.
[0012] Furthermore, in step S2, the watershed ecological mapping is to set the normalized watershed ecological data of the i-th sampling point at time t as , select the sliding window length T, introduce the time weight coefficient for each delay step, expressed as: ; Control the decay rate of historical information; construct the time fusion feature vector, expressed as: ; Where T is the length of the sliding time window; k is the time point index; and are the ecological data vectors collected by the i-th sampling point at time t-1 and time tT respectively; is the information decay rate parameter; 、 、 and is the time weight coefficient at different moments; is the vector concatenation operator; is the time series fusion feature vector of the i-th sampling point; for any two sampling points and The temporal fusion feature vector of and ,get ; Then construct the distance between sampling points in the whole basin, expressed as: ;in, is the temporal similarity index; is the bandwidth parameter; is the distance between the i-th sampling point and the j-th sampling point; is a mapping function.
[0013] Furthermore, in step S3, the radius optimization is to set the initial radius for ; max(·) and min(·) are the maximum and minimum values respectively; calculate the local density, expressed as: ; ;in, is the local density of the i-th sampling point; N is the total number of sampling points, j is the sampling point index; r is the sampling radius; is the indicator function, x is the function variable; the actual average density is expressed as: ;in, is the actual average density under radius r; p is the ratio of the expected average density to the total number of points; r is adjusted by Newton iteration until .
[0014] Furthermore, in step S4, the initial cluster center selection is to calculate the minimum kernel distance to all higher density sampling points for each sampling point, which is expressed as: ; Introduce time change factors to construct composite indicators , expressed as: ;in, is the shortest distance between sampling points; is the local density of the jth sampling point; is the time trend weight; is the i-th sampling point at time The collected ecological data vector; Sort in descending order, take the first C points as the initial cluster centers, and add the minimum center distance constraint, expressed as: ; where a is the scaling factor; only the distance between any two centers is not less than The sampling point The largest sampling point is used as the core reference point and is denoted by q.
[0015] Furthermore, in step S5, the watershed ecological data clustering process uses weighted kernel K-means clustering on the time fusion feature vector to construct the objective function G, which is expressed as: ; Where c is the cluster; N is the total number of sampling points; D is the total number of dimensions, and d is the dimension index; is the weight of the i-th sampling point; Is whether the i-th sampling point belongs to cluster c; m is the exponential parameter; is the weight of the dth dimension in cluster c, indicating the importance of the ecological variable; is the prototype value; It is The eigenvalue of the dth dimension; and is the item weight; and are the means of the t-th iteration and the t-1-th iteration of the d-th dimension of cluster c respectively; the construction constraints are: ;in, ;in, is the initial cluster center of cluster c; ; and each time the clustering is iterated, the stability of the indicator is updated , expressed as: ,in, is the feature weight before updating; is the mean of the dth dimension of cluster c; is the adjustment coefficient.
[0016] Furthermore, in step S6, the watershed ecological anomaly monitoring is to complete the fusion feature vector using weighted kernel K-means clustering, and then detect whether there is an anomaly in the time fusion feature vector through the local outlier factor, and issue an early warning to the sampling point corresponding to the abnormal time fusion feature vector.
[0017] The beneficial effects achieved by the present invention using the above scheme are as follows:
[0018] (1) In view of the fact that the general watershed ecological anomaly monitoring method ignores the characteristics of ecological data changes in the time dimension, improperly captures sudden anomalies, and has difficulty in distinguishing real anomalies from noise, which leads to poor accuracy of anomaly monitoring, this scheme is based on watershed ecological data and introduces a time weight coefficient to construct a time fusion feature vector, highlighting the changes in recent ecological data, reducing the interference of long-term historical noise, and improving the sensitivity to sudden situations; by iteratively adjusting the radius, it automatically adapts to the sampling density under different river sections and seasonal changes in flow, thereby improving monitoring accuracy; introducing a time change factor, constructing a composite indicator to select core reference points, objectively emphasizing the most representative core position, and providing reliability for subsequent anomaly monitoring.
[0019] (2) In view of the problem that abnormal monitoring is often overlooked in many normal monitoring methods and the contribution of different indicators is not differentiated enough, which leads to poor abnormal monitoring results, this scheme introduces sample weights to strengthen abnormal monitoring and increase the probability of abnormal monitoring being identified. Through the guidance of core reference points, the degree of attention paid to different indicators is continuously adjusted to enhance the influence of important ecological variables and prevent excessive attention to secondary ecological variables. Based on the stability of indicators, the weights are updated to distinguish the contribution of different indicators to abnormal monitoring and highlight the importance of key indicators. This improves the abnormal monitoring effect. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 This is a flow chart of the watershed ecological anomaly monitoring method based on clustering processing provided by the present invention.
[0021] The accompanying drawings are used to provide further understanding of the present invention and constitute a part of the specification. They are used to explain the present invention together with the embodiments of the present invention and do not constitute a limitation of the present invention. DETAILED DESCRIPTION
[0022] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments; based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0023] In the description of the present invention, it should be understood that terms such as "up", "down", "front", "back", "left", "right", "top", "bottom", "inside" and "outside" indicating directions or positional relationships are based on the directions or positional relationships shown in the accompanying drawings. They are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the system or element referred to must have a specific direction, be constructed and operated in a specific direction. Therefore, they should not be understood as limiting the present invention.
[0024] Example 1, see Figure 1 The present invention provides a method for monitoring abnormal watershed ecology based on clustering processing, which includes the following steps:
[0025] Step S1: Data collection: Collect watershed ecological data;
[0026] Step S2: Watershed ecological mapping: Based on the watershed ecological data, a time fusion feature vector is constructed by introducing a time weight coefficient; and the distance between sampling points is obtained;
[0027] Step S3: Radius optimization: setting the initial radius and iteratively adjusting the radius based on the local density and the actual average density;
[0028] Step S4: Initial cluster center selection: Based on the distance and radius between sampling points, the time variation factor is introduced to construct a composite index, and the initial cluster center and core reference point are finally selected;
[0029] Step S5: Clustering of watershed ecological data; performing weighted kernel K-means clustering on the time fusion feature vectors to construct the objective function guided by the core reference points;
[0030] Step S6: watershed ecological anomaly monitoring; using local outlier factors to perform anomaly monitoring on the clustering results.
[0031] Example 2, see Figure 1This embodiment is based on the above embodiment. In step S1, data collection is to collect watershed ecological data; the watershed ecological data includes water quality data, hydrological data, meteorological data and remote sensing indicators; the water quality data includes dissolved oxygen, pH and turbidity; the hydrological data includes flow and flow velocity; the meteorological data includes rainfall and air temperature; the remote sensing indicators include vegetation index and land surface temperature; each dimension of the watershed ecological data is normalized to the maximum and minimum to eliminate the influence of different dimensions and ensure the stability of subsequent calculated values.
[0032] Example 3, see Figure 1 This embodiment is based on the above embodiment. In step S2, the watershed ecological mapping is to highlight the recent ecological data changes, reduce the long-term historical noise interference, and improve the sensitivity to sudden pollution caused by industrial wastewater leakage and abnormal rainfall events; let the normalized watershed ecological data of the i-th sampling point at time t be , select the sliding window length T, introduce the time weight coefficient for each delay step, expressed as: ; Control the decay rate of historical information; construct the time fusion feature vector, expressed as: ; Where T is the length of the sliding time window, which contains T+1 time points in total, capturing the ecological dynamics of the recent period; k is the time point index; and are the ecological data vectors collected by the i-th sampling point at time t-1 and time tT respectively; is the information decay rate parameter, The larger it is, the faster the past data decays; 、 、 and is the time weight coefficient at different moments; It is a vector concatenation operator used to construct temporal embedding; is the time series fusion feature vector of the i-th sampling point, and each sampling point corresponds to a time series fusion feature vector; recent data has a higher weight, which can quickly respond to pollution outbreaks; historical noise is attenuated to enhance stability; based on attenuation embedding, it captures the nonlinear coupling between multi-source ecological variables of rainfall → runoff → turbidity changes and maintains the time attenuation characteristics; for any two sampling points and The temporal fusion feature vector of and ,get ; Then construct the distance between sampling points in the whole basin, expressed as: ;in, It is a temporal similarity index that measures the similarity of sampling points in terms of temporal-ecological characteristics; is the bandwidth parameter; is the distance between the i-th sampling point and the j-th sampling point; is a mapping function.
[0033] Example 4, see Figure 1 This embodiment is based on the above embodiment. In step S3, the radius optimization is that the sampling points are affected by the differences in the upstream, downstream and tributaries of the river, and are often uneven in spatial distribution. A fixed radius is likely to cause distortion of local density estimation. Set the initial radius for ; max(·) and min(·) are the maximum and minimum values respectively; calculate the local density, expressed as: ; ;in, is the local density of the i-th sampling point; N is the total number of sampling points, j is the sampling point index; r is the sampling radius; is the indicator function, x is the function variable; the actual average density is expressed as: ;in, is the actual average density under radius r; p is the ratio of the expected average density to the total number of points; r is adjusted by Newton iteration until Automatically adapt to the sampling density of different river sections and seasonal changes in flow; ensure reasonable density estimation in both sparse and dense areas, and improve monitoring sensitivity.
[0034] Example 5, see Figure 1 This embodiment is based on the above embodiment. In step S4, the initial cluster center is selected to distinguish the true ecological anomaly core from isolated noise; for each sampling point, the minimum kernel distance to all sampling points with higher density is calculated, which is expressed as: The time variation factor is introduced to more accurately distinguish the real abnormal pattern from the observation noise, highlight the sampling points with obvious time trend changes, and better identify the abnormal situation where the concentration of pollutants in the water body continues to increase over time, thereby constructing a composite index. , expressed as: ;in, is the shortest distance between sampling points; is the local density of the jth sampling point; is the time trend weight; is the i-th sampling point at time The collected ecological data vector; Sort in descending order, take the first C points as the initial cluster centers, and add the minimum center distance constraint, expressed as: ; where a is the scaling factor; only the distance between any two centers is not less than The sampling point The largest sampling point is used as the core reference point, denoted by q, to guide the clustering of watershed ecological data; Effectively distinguish real abnormal patterns from observation noise; objectively emphasize the most representative core positions based on core reference points.
[0035] By performing the above operations, we can address the problem that general watershed ecological anomaly monitoring methods ignore the changing characteristics of ecological data in the temporal dimension, improperly capture sudden anomalies, and find it difficult to distinguish between real anomalies and noise, which in turn leads to poor accuracy in anomaly monitoring. Based on watershed ecological data, this scheme introduces a time weight coefficient to construct a time fusion feature vector, highlighting the changes in recent ecological data, reducing the interference of long-term historical noise, and improving sensitivity to sudden situations; by iteratively adjusting the radius, it automatically adapts to the sampling density under different river sections and seasonal changes in flow, thereby improving monitoring accuracy; introduces a time change factor, constructs a composite indicator to select core reference points, objectively emphasizes the most representative core position, and provides reliability for subsequent anomaly monitoring.
[0036] Example 6, see Figure 1 This embodiment is based on the above embodiment. In step S5, the watershed ecological data clustering process is because abnormal monitoring requires soft judgment between normal and abnormal data to reflect the gradual change of ecological status. Different indicators contribute differently to abnormalities, and a small amount of abnormal monitoring needs to be given higher credibility to avoid being submerged in a large amount of normal data. The time fusion feature vector is clustered using weighted kernel K-means to construct the objective function G, which is expressed as: ; Where c is the cluster; N is the total number of sampling points; D is the total number of dimensions, and d is the dimension index; is the weight of the i-th sampling point; Is whether the i-th sampling point belongs to cluster c; m is the exponential parameter; is the weight of the dth dimension in cluster c, indicating the importance of the ecological variable; is the prototype value; It is The eigenvalue of the dth dimension; and is the item weight; and are the means of the t-th iteration and the t-1-th iteration of the d-th dimension of cluster c respectively; the construction constraints are: ;in, ;in, is the initial cluster center of cluster c; ; By sample weight Strengthen possible anomaly monitoring; through feature weights Dynamically highlight key ecological indicators; guide clustering with the most representative anomaly monitoring through the viewpoint mechanism to avoid missing important anomalies due to random initialization; and update the stability of indicators based on each clustering iteration , dynamically enhance the influence of important ecological variables and prevent excessive attention to secondary ecological variables, expressed as: ,in, is the feature weight before updating; is the mean of the dth dimension of cluster c; is the adjustment coefficient.
[0037] By performing the above operations, this scheme introduces sample weights to strengthen anomaly monitoring and increase the probability of anomaly monitoring being identified, addressing the problem that anomaly monitoring is ignored among many normal monitoring methods in general watershed ecological anomaly monitoring methods, and the contribution of different indicators is insufficiently distinguished, which leads to poor anomaly monitoring effects. It also guides the scheme by continuously adjusting the degree of attention paid to different indicators through core reference points, enhancing the influence of important ecological variables, and preventing excessive attention to secondary ecological variables. It also updates weights based on indicator stability, distinguishes the degree of contribution of different indicators to anomaly monitoring, and highlights the importance of key indicators, thereby improving the anomaly monitoring effect.
[0038] Example 7, see Figure 1 This embodiment is based on the above embodiment. In step S6, the watershed ecological anomaly monitoring is to complete the fusion feature vector using weighted kernel K-means clustering, and then detect whether there is an abnormality in the time fusion feature vector through the local outlier factor, and issue an early warning for the sampling point corresponding to the abnormal time fusion feature vector.
[0039] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that includes a list of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus.
[0040] While the embodiments of the present invention have been shown and described, it will be apparent to those skilled in the art that various changes, modifications, substitutions, and alterations can be made to the embodiments without departing from the principles and spirit of the invention.
[0041] The present invention and its embodiments are described above. This description is not restrictive. The drawings show only one embodiment of the present invention, and the actual structure is not limited thereto. In short, if a person skilled in the art is inspired by this and, without departing from the purpose of the present invention, designs structures and embodiments similar to this technical solution without inventiveness, they shall fall within the scope of protection of the present invention.
Claims
1. A watershed ecological anomaly monitoring method based on clustering processing, characterized by: The method comprises the following steps: Step S1: Data collection: Collect watershed ecological data; Step S2: Watershed ecological mapping: Based on the watershed ecological data, a time fusion feature vector is constructed by introducing a time weight coefficient; and the distance between sampling points is obtained; Step S3: Radius optimization: setting the initial radius and iteratively adjusting the radius based on the local density and the actual average density; Step S4: Initial cluster center selection: Based on the distance and radius between sampling points, the time variation factor is introduced to construct a composite index, and the initial cluster center and core reference point are finally selected; Step S5: Clustering of watershed ecological data; performing weighted kernel K-means clustering on the time fusion feature vectors to construct the objective function guided by the core reference points; Step S6: watershed ecological anomaly monitoring: using local outlier factors to perform anomaly monitoring on the clustering results; In step S2, the watershed ecological mapping is to set the normalized watershed ecological data of the i-th sampling point at time t as , select the sliding window length T, introduce the time weight coefficient for each delay step, expressed as: ; Control the decay rate of historical information; construct the time fusion feature vector, expressed as: ; Where T is the length of the sliding time window; k is the time point index; and are the ecological data vectors collected by the i-th sampling point at time t-1 and time tT respectively; is the information decay rate parameter; 、 、 and is the time weight coefficient at different moments; is the vector concatenation operator; is the time series fusion feature vector of the i-th sampling point; for any two sampling points and The temporal fusion feature vector of and ,get ; Then construct the distance between sampling points in the whole basin, expressed as: ;in, is the temporal similarity index; is the bandwidth parameter; is the distance between the i-th sampling point and the j-th sampling point; is a mapping function.
2. The method for monitoring watershed ecological anomalies based on clustering processing according to claim 1, characterized in that: In step S3, the radius optimization is to set the initial radius for ; max(·) and min(·) are the maximum and minimum values respectively; Calculate the local density, expressed as: ; ;in, is the local density of the i-th sampling point; N is the total number of sampling points, j is the sampling point index; r is the sampling radius; is the indicator function, x is the function variable; the actual average density is expressed as: ;in, is the actual average density under radius r; p is the ratio of the expected average density to the total number of points; r is adjusted by Newton iteration until .
3. The method for monitoring watershed ecological anomalies based on clustering processing according to claim 2, characterized in that: In step S4, the initial cluster center selection is to calculate the minimum kernel distance to all higher density sampling points for each sampling point, which is expressed as: ; Introduce time change factors to construct composite indicators , expressed as: ;in, is the shortest distance between sampling points; is the local density of the jth sampling point; is the time trend weight; is the i-th sampling point at time The collected ecological data vector; Sort in descending order, take the first C points as the initial cluster centers, and add the minimum center distance constraint, expressed as: ; where a is the scaling factor; only the distance between any two centers is not less than The sampling point The largest sampling point is used as the core reference point and is denoted by q.
4. The method for monitoring watershed ecological anomalies based on clustering processing according to claim 3, characterized in that: In step S5, the watershed ecological data clustering process uses weighted kernel K-means clustering on the time fusion feature vector to construct the objective function G, which is expressed as: ; Where c is the cluster; N is the total number of sampling points; D is the total number of dimensions, and d is the dimension index; is the weight of the i-th sampling point; Is whether the i-th sampling point belongs to cluster c; m is the exponential parameter; is the weight of the dth dimension in cluster c, indicating the importance of the ecological variable; is the prototype value; It is The eigenvalue of the dth dimension; and is the item weight; and are the means of the t-th iteration and the t-1-th iteration of the d-th dimension of cluster c respectively; the construction constraints are: ;in, ;in, is the initial cluster center of cluster c; ; and each time the clustering is iterated, the stability of the indicator is updated , expressed as: ,in, is the feature weight before updating; is the mean of the dth dimension of cluster c; is the adjustment coefficient.
5. The method for monitoring watershed ecological anomalies based on clustering processing according to claim 4, characterized in that: In step S1, the data collection is to collect watershed ecological data; and perform maximum and minimum normalization on each dimension of the watershed ecological data.
6. The method for monitoring watershed ecological anomalies based on clustering processing according to claim 5, characterized in that: In step S6, the watershed ecological anomaly monitoring is to complete the fusion feature vector using weighted kernel K-means clustering, and then detect whether there is an anomaly in the time fusion feature vector through the local outlier factor, and issue an early warning to the sampling point corresponding to the abnormal time fusion feature vector.
Citation Information
Patent Citations
Method, system and device for clustering analysis of crowd based on spatio-temporal data
CN119513636A
Sewage treatment supervision method and system based on artificial intelligence control
CN119917957A