Water quality range determination method and disturbed water quality monitoring point identification method

Cluster analysis was used to determine the water quality monitoring line range and identify the locations of water quality monitoring points that were disturbed. This solved the problem of human interference with water quality monitoring data in existing technologies and enabled efficient and low-cost identification of water quality monitoring points.

CN119474937BActive Publication Date: 2025-12-09CHINA NAT ENVIRONMENTAL MONITORING CENT
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411820827.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-11
Publication Date
2025-12-09
Estimated Expiration
2044-12-11

AI Technical Summary

Technical Problem

Existing technologies are insufficient to effectively identify water quality monitoring points that have been interfered with by human intervention, resulting in distorted water quality monitoring data and making it impossible to accurately identify the responsible party.

Method used

Cluster analysis was used to determine the water quality boundary range, the number of clusters at the elbow was determined by the sum of cluster errors, the number of clusters was selected as the upper and lower limits of the water quality boundary range, and the locations of water quality monitoring points affected by interference were identified in conjunction with water quality assessment indicators.

Benefits of technology

It improves the accuracy of identifying water quality monitoring points affected by human interference, reduces identification costs, and minimizes regulatory risks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119474937B_ABST
    Figure CN119474937B_ABST
Patent Text Reader

Abstract

The embodiments of the present disclosure provide a water quality fitting line range determination method and a disturbed water quality monitoring point identification method. The water quality fitting line range determination method comprises: performing cluster analysis on water quality monitoring data of existing monitoring points by using each cluster number, respectively, to determine a cluster set of the water quality monitoring data under each cluster number; calculating a corresponding cluster error sum value based on the cluster set under each cluster number, and determining an elbow cluster number based on the corresponding cluster error sum value of each cluster number; selecting a selected cluster number in the cluster number greater than the elbow cluster number, and taking upper and lower limits of a cluster cluster including a water quality evaluation index in the cluster set corresponding to the selected cluster number as upper and lower limits of a water quality fitting line range.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the field of environmental protection screening technology, and in particular to a water quality line range determination method and a disturbed water quality monitoring point identification method. BACKGROUND

[0002] Driven by the motivation of water quality performance assessment compliance, some polluting enterprises or administrative departments with regulatory responsibility for polluted water bodies may achieve water quality compliance by artificially interfering with the water quality of monitoring points. Correspondingly, how to identify the aforementioned artificially disturbed monitoring points, and then exclude the corresponding water quality monitoring data and trace back the relevant responsible subject, is an important work of environmental protection supervision. However, how to identify the water body data characteristics of the aforementioned artificially disturbed monitoring points, and then identify the data disturbed monitoring points according to the water body data characteristics, is a technical problem that needs to be solved first.

[0003] Currently, the environmental protection supervision department still adopts the method of patrolling and checking the water quality monitoring points to avoid the water quality disturbance operation of the relevant responsible subject on the water quality monitoring points. However, the aforementioned method is too costly and cannot effectively identify some organized and targeted numerical disturbance behaviors. In order to avoid the aforementioned problems, the related technology proposes a method of determining the line range according to the water quality monitoring index, and determining the abnormal point according to the line range. However, the aforementioned line range is easy to be identified, and the responsible subject who actively interferes with the water quality will also actively interfere with the water quality in a way that makes the water quality monitoring data outside the aforementioned line range. SUMMARY

[0004] The present disclosure provides a water quality line range determination method and a disturbed water quality monitoring point identification method.

[0005] In a first aspect, the present disclosure provides a water quality line range determination method, comprising:

[0006] The water quality monitoring data of the existing monitoring points is analyzed by clustering with each cluster number respectively, and the cluster sets of the water quality monitoring data under each cluster number are determined;

[0007] The cluster error sum values corresponding to each cluster number are calculated based on the cluster sets under each cluster number, and the elbow cluster number is determined based on the cluster error sum values corresponding to each cluster number;

[0008] A selected cluster number is selected from the cluster numbers greater than the elbow cluster number, and the upper and lower limits of the cluster cluster including the water quality assessment index in the cluster set corresponding to the selected cluster number are taken as the upper and lower limits of the water quality line range.

[0009] Optionally, the selected cluster number is selected from the cluster numbers greater than the elbow cluster number, comprising:

[0010] S1: determine the i-th cluster number and the i+1-th cluster number after the elbow cluster number, wherein the initial value of i is 0;

[0011] S2: calculate the reduction ratio of the cluster error sum values corresponding to the i-th cluster number and the i+1-th cluster number, and determine whether the reduction ratio is less than a preset reduction ratio; if yes, perform S3; if no, perform S4;

[0012] S3: take the i+1-th cluster number as the selected cluster number;

[0013] S4: set i=i+1, and perform S1.

[0014] Optionally, the selected cluster number is selected from the cluster numbers greater than the elbow cluster number, and the method comprises:

[0015] S1: determine the i-th cluster number and the i+1-th cluster number after the elbow cluster number, wherein the initial value of i is 0;

[0016] S2: calculate the difference value of the cluster error sum values corresponding to the i-th cluster number and the i+1-th cluster number, and determine whether the difference value is less than a set difference threshold; if yes, perform S3; if no, perform S4;

[0017] S3: take the i+1-th cluster number as the selected cluster number;

[0018] S4: set i=i+1, and perform S1.

[0019] Optionally, the S3 comprises:

[0020] S31: take the i+1-th cluster number as a to-be-selected cluster number, and determine the upper and lower limits of a cluster cluster including the water quality evaluation index corresponding to the to-be-selected cluster number as to-be-selected upper and lower limits;

[0021] S32: determine whether the water quality evaluation index is in an intermediate region determined based on the to-be-selected upper and lower limits, wherein the upper and lower limits of the intermediate region are located within the to-be-selected upper and lower limits; if yes, perform S33; if no, perform S4;

[0022] S33: take the i+1-th cluster number as the selected cluster number.

[0023] Optionally, the calculating the cluster error sum values corresponding to each cluster number based on the cluster clusters respectively comprises:

[0024] For each cluster cluster under a cluster number, a corresponding cluster center is calculated respectively;

[0025] Calculate the intra-cluster distance sum value based on the intra-cluster center and the intra-cluster water quality monitoring data of each cluster;

[0026] Calculate the cluster error sum value based on the intra-cluster distance sum value of each cluster under a cluster number.

[0027] Optionally, the elbow cluster number is determined based on the cluster error sum value corresponding to each cluster number, comprising:

[0028] Calculate the first-order difference value of the adjacent cluster number, and calculate the difference value ratio of the adjacent next first-order difference value and the previous adjacent first-order difference value;

[0029] Select the cluster number that makes the difference value ratio less than a preset ratio value as the elbow cluster number.

[0030] In a second aspect, the embodiments of the present disclosure provide a method for identifying an interfered water quality monitoring point, comprising:

[0031] Obtain the historical monitoring data sequence of the water quality monitoring point to be identified, and count the number of line-fitting monitoring data in the historical monitoring data sequence; the line-fitting monitoring data is the monitoring data located in the water quality line-fitting range, and the water quality line-fitting range is determined by using the method for determining the water quality line-fitting range as described above;

[0032] Calculate the line-fitting data proportion of the water quality monitoring point to be identified based on the number of line-fitting monitoring data and the total amount of data of the historical monitoring data sequence;

[0033] Determine whether the line-fitting data proportion is greater than an upper limit threshold of the proportion;

[0034] In the case that the line-fitting data proportion is greater than the upper limit threshold of the proportion, determine that the water quality monitoring point to be identified is an interfered water quality monitoring point.

[0035] Optionally, before determining whether the line-fitting data proportion is greater than the upper limit threshold of the proportion, the method further comprises:

[0036] Count the line-fitting data proportion of each monitoring point based on the historical monitoring data sequence of all monitoring points;

[0037] Statistically analyze the line-fitting data proportion of each monitoring point to determine the mean value and the mean square error of the proportion;

[0038] Determine the upper limit threshold of the proportion based on the mean value and the mean square error of the proportion.

[0039] Optionally, the method further includes: obtaining various types of water quality monitoring data of the to-be-identified water quality monitoring point, and respectively counting the number of the line-fitting monitoring data in the various types of water quality monitoring data.

[0040] The method further includes: respectively calculating the line-fitting data proportion of the various types of water quality monitoring data at the to-be-identified water quality monitoring point.

[0041] The method further includes: respectively judging whether the line-fitting data proportion of the various types of water quality monitoring data is greater than a corresponding preset upper limit threshold of the proportion.

[0042] In a case where the line-fitting data proportion is greater than the upper limit threshold of the proportion, determining that the to-be-identified water quality monitoring point is an interfered water quality monitoring point includes: in a case where the line-fitting data proportion of one type of water quality monitoring data is greater than a corresponding preset upper limit threshold of the proportion, determining that the to-be-identified water quality monitoring point is an interfered water quality monitoring point.

[0043] In a third aspect, an embodiment of the present disclosure provides a computing device including a processor and a memory, the memory being configured to store a computer program; the computer program, when loaded by the processor, causes the processor to execute the water quality line-fitting range determination method or the interfered water quality monitoring point identification method.

[0044] In the embodiment of the present disclosure, the cluster sets under each cluster number are determined by the method of cluster analysis, the elbow cluster number is determined according to the cluster error and value of each cluster number cluster set, a larger cluster number is determined as the selected cluster number according to the elbow cluster number, and finally the upper and lower limits of the cluster cluster including the water quality evaluation index in the selected cluster number cluster set are used as the upper and lower limits of the water quality line-fitting range. Because the cluster number greater than the elbow cluster number is used as the selected cluster number, the problem of too large upper and lower limits of the determined line-fitting range can be removed. Because the upper and lower limits of the water quality line-fitting range determined by the embodiment of the present disclosure can be obtained based on the cluster analysis of a large amount of water quality monitoring data and unsupervised learning, it is not easy to predict, and therefore it can better identify whether the line-fitting affects the water quality. BRIEF DESCRIPTION OF DRAWINGS

[0045] The accompanying drawings, which are incorporated herein and form part of the specification, illustrate embodiments consistent with the present disclosure and, together with the description, further serve to explain the principles of the present disclosure.

[0046] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, for those of ordinary skill in the art, other drawings can also be obtained from these drawings without any creative effort, and these drawings are mainly used to provide a better understanding of the present disclosure.

[0047] Figure 1 is a flow chart of a method for determining a water quality range provided by the present disclosure;

[0048] Figure 2 is a flow chart of a method for identifying an interfered water quality monitoring point provided by the present disclosure;

[0049] Figure 3 is a structural schematic diagram of a computing device provided by the present disclosure. DETAILED DESCRIPTION

[0050] The embodiments of the present disclosure will be described in more detail below with reference to the drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms, and should not be interpreted as being limited to the embodiments set forth herein, but rather, these embodiments are provided to make the present disclosure more thorough and complete. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes, and are not intended to limit the scope of protection of the present disclosure.

[0051] The term "comprising" and variations thereof as used herein are open-ended, that is "including but not limited to". The term "based on" is "based, at least in part, on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Related definitions for other terms will be given in the description below. In this document, relational terms such as "first" and "second", and the like, are used solely to distinguish one entity or action from another entity or action, without necessarily requiring or implying any actual such relationship or order between or among the entities or actions.

[0052] The water quality line range determination method provided by the embodiments of the present disclosure is shown in FIG. 1. As shown in FIG. 1, the water quality line range determination method provided by the embodiments of the present disclosure includes S110-S130.

[0053] Figure 1 The water quality line range determination method provided by the embodiments of the present disclosure is shown in FIG. 1. As shown in FIG. 1, the water quality line range determination method provided by the embodiments of the present disclosure includes S110-S130. Figure 1

[0054] S110: Cluster analysis is performed on the water quality monitoring data of the existing monitoring points by using each cluster number respectively, and the cluster sets of the water quality monitoring data under each cluster number are determined.

[0055] Before the execution of the present disclosure scheme, a large amount of water quality monitoring data of long-term water quality monitoring of monitoring points is obtained. The monitoring points in the present disclosure are points set for monitoring the quality of water bodies of specific types or pollution emissions.

[0056] For example, in the case of water body environmental quality evaluation of surface water such as rivers, lakes and seas, the monitoring points are points set at the main positions or key positions of the water bodies. For another example, in the case of monitoring whether the pollution discharged by an enterprise meets the standard, the monitoring points are the positions where the enterprise's sewage culvert is connected with the surface water body, or the positions adjacent to the above-mentioned connected positions.

[0057] ​Correspondingly, the monitoring points in different types of location areas have different types of water quality evaluation indexes. For example, according to the existing water quality standard specification, surface water is divided into five types of water bodies. The water quality evaluation indexes of type I water bodies such as source water and water bodies in national nature reserves include chemical oxygen demand ≤ 15 mg / L, five-day biochemical oxygen demand ≤ 2 mg / L, type II water bodies (mainly applicable to centralized drinking water surface water source area first-class protection zone, rare aquatic habitat, fish and shrimp spawning ground, and fish larvae feeding ground) water quality evaluation indexes include chemical oxygen demand ≤ 30 mg / L, five-day biochemical oxygen demand ≤ 5 mg / L, chemical oxygen demand ≤ 15 mg / L, type III water bodies (mainly applicable to centralized drinking water surface water source area second-class protection zone, fish and shrimp overwintering ground, migration channel, aquaculture area, and fishery water area and swimming area) water quality evaluation indexes include chemical oxygen demand ≤ 40 mg / L, five-day biochemical oxygen demand ≤ 6 mg / L, type IV water bodies (mainly applicable to general industrial water area and human body non-direct contact entertainment water area) water quality evaluation indexes include chemical oxygen demand ≤ 60 mg / L, five-day biochemical oxygen demand ≤ 10 mg / L, and type V water bodies (mainly applicable to agricultural water area and general landscape water area) water quality evaluation indexes include chemical oxygen demand ≤ 100 mg / L, five-day biochemical oxygen demand ≤ 15 mg / L. In addition, various water bodies also have corresponding ammonia nitrogen content indexes, potassium permanganate index, total phosphorus content, and other indexes, which are not listed here.

[0058] It should be noted here that the existing monitoring points include points that may interfere with water quality monitoring data such as subsequent evaluation monitoring points. In specific implementation, the types of points that interfere with water quality monitoring data include spraying, aeration, stirring water, setting up a sewage culvert to dilute the concentration of pollutants, etc.

[0059] More preferably, the aforementioned water quality monitoring data includes water quality monitoring data under various seasons and rainfall conditions (e.g., flood season, dry season) to ensure that a large amount of water quality monitoring data has broad representativeness and has the characteristic of being significantly affected by river water level changes. In addition, in specific implementation, in order to cluster the rationality of clustering itself, water quality monitoring data of rivers with similar water characteristics should be used as much as possible for clustering analysis.

[0060] In the embodiments of the present disclosure, the water quality monitoring data of the existing monitoring points is clustered and analyzed by using each cluster number respectively, that is, each cluster number is used as a hyperparameter, and the water quality monitoring data is clustered and analyzed by using the aforementioned hyperparameter.

[0061] In specific implementation, the clustering analysis method adopted is a K-means clustering analysis method or a method with the same clustering essence. The K-means clustering analysis method first randomly selects K initial data as clustering centers, calculates the distance between each water quality monitoring data and the initial clustering center, constructs a loss function to minimize the distance between all water quality monitoring data and the initial clustering center, and determines K clustering clusters. After determining the K clustering clusters, the K clustering centers are recalculated according to the water quality monitoring data in each clustering cluster, and the foregoing method is repeatedly executed until the clustering center no longer changes or the maximum iteration number is reached. The foregoing distance can be a Euclidean distance or a sum of squares distance.

[0062] After clustering analysis is performed on the water quality monitoring data of the existing monitoring points by using each clustering number, the clustering set under the corresponding clustering number can be obtained. In the embodiment of the present disclosure, considering that a small clustering number does not achieve the significance of clustering, and a large clustering number does not meet the demand in water quality monitoring, the clustering number can be from 3 in actual application, and the value can reach about 20, and a larger or smaller clustering number is not used.

[0063] S120: Calculate the clustering error sum value corresponding to each clustering number based on the clustering set under the clustering number, and determine the elbow clustering number based on the clustering error sum value corresponding to each clustering number.

[0064] After obtaining the clustering set under each clustering number, the clustering error sum value corresponding to each clustering number is calculated by using the clustering set under each clustering number. In specific implementation, the clustering error sum value can be a loss function obtained by performing clustering for the last time, or can be data representing the clustering characteristics of the entire water quality monitoring data.

[0065] In some embodiments, the clustering error sum value can be calculated by using S121-S123 as follows.

[0066] S121: For each clustering cluster under a clustering number, calculate the corresponding cluster center.

[0067] In specific implementation, the foregoing cluster center can be the center determined when the clustering center is determined for the last time during clustering analysis, or can be the centroid of the internal water quality monitoring data of each cluster determined by using a centroid calculation method.

[0068] S122: Calculate the cluster distance sum value based on the cluster center and the cluster water quality monitoring data of each clustering cluster.

[0069] After obtaining the cluster center of each clustering cluster, the distance between each water quality monitoring data and the corresponding cluster center is calculated, and the cluster distance sum value can be obtained by adding the foregoing distances.

[0070] S123: Calculate the cluster error sum value based on the intra-cluster distance sum value of each cluster cluster.

[0071] After obtaining the intra-cluster distance sum value of all cluster clusters, the intra-cluster distance sum value is added to obtain the cluster error sum value.

[0072] After obtaining the cluster error sum value corresponding to each cluster number, the elbow cluster number can be determined based on the cluster error sum value corresponding to each cluster number. The elbow cluster number is the point at which the change characteristic of the cluster error sum value rapidly decreases. In a specific implementation, the elbow cluster number can be determined according to S124-S125 as follows.

[0073] S124: Calculate the first-order difference value of the adjacent cluster number, and calculate the difference value ratio of the next first-order difference value and the previous adjacent first-order difference value.

[0074] Suppose there are now 5 cluster numbers, 5, 6, 7, 8, and 9, corresponding to the cluster error sum values x5, x6, x7, x8, and x9, and corresponding first-order difference values x6-x5, x7-x6, x8-x7, and x9-x8. According to the basic principle of cluster analysis, x5>x6>x7>x8>x9, and the corresponding x6-x5, x7-x6, x8-x7, and x9-x8 are all negative values.

[0075] Then, the next first-order difference value and the previous first-order difference value are divided by each other to obtain the corresponding difference value ratio, which is (x7-x6) / (x6-x5), (x8-x7) / (x7-x6), and (x8-x7 / )(x9-x8).

[0076] S125: Select the first cluster number that makes the difference value ratio less than the preset ratio as the elbow cluster number.

[0077] After calculating each difference ratio, the difference ratios are compared from front to back with the preset ratio. The preset ratio is a value determined in advance according to experience. When the first difference value ratio is less than the preset ratio, the cluster number that makes the corresponding difference ratio less than the preset ratio is determined as the elbow cluster number. For example, if (x7-x6) / (x6-x5) is the first cluster number that makes the difference ratio less than the preset ratio, then the corresponding 6 is the elbow cluster number.

[0078] S130: Select the selected cluster number in the cluster number greater than the elbow cluster number, and take the upper and lower limits of the cluster cluster including the water quality evaluation index in the selected cluster number as the upper and lower limits of the water quality contour line range.

[0079] After the elbow cluster quantity is determined, a selected cluster quantity can be selected from the cluster quantities greater than the elbow cluster quantity.

[0080] In the embodiments of the present disclosure, it has been tried whether the upper and lower limits of the cluster cluster including the water quality evaluation indexes in the corresponding cluster set can be determined as the upper and lower limits of the water quality fitting line range according to the elbow cluster quantity as the selected cluster quantity. However, through a large amount of big data analysis, it is found that the numerical value between the upper and lower limits of the water quality fitting line range determined when the elbow cluster quantity is selected as the selected cluster quantity is large, and the upper and lower limits of the cluster cluster including the water quality evaluation indexes determined directly by the elbow cluster quantity are not good as the upper and lower limits of the water quality fitting line range.

[0081] Therefore, in the present scheme, the selected cluster quantity is selected from the cluster quantities greater than the elbow cluster quantity, and the upper and lower limits of the cluster cluster including the water quality evaluation indexes in the corresponding cluster set based on the selected cluster quantity are taken as the upper and lower limits of the water quality fitting line range.

[0082] In some embodiments, the computing device determines the selected cluster quantity by S131-S134 as follows.

[0083] S131: Determine the i-th cluster quantity and the i+1-th cluster quantity after the elbow cluster quantity is determined.

[0084] Wherein the initial value of i is 0.

[0085] S132: Calculate the reduction ratio of the cluster error sum values corresponding to the i-th cluster quantity and the i+1-th cluster quantity, and determine whether the reduction ratio is less than a preset reduction ratio; if yes, execute S133; if no, execute S134.

[0086] The reduction ratio of the cluster error sum values corresponding to the i-th cluster quantity and the i+1-th cluster quantity is obtained by subtracting the cluster error sum value corresponding to the i+1-th cluster quantity from the cluster error sum value corresponding to the i-th cluster quantity to obtain a difference absolute value, and then comparing the difference absolute value with the cluster error corresponding to the i-th cluster quantity to obtain the reduction ratio.

[0087] S133: Take the i+1-th cluster quantity as the selected cluster quantity.

[0088] S134: Let i=i+1, and execute S131.

[0089] If the reduction ratio corresponding to the i+1-th cluster quantity is already less than the preset reduction ratio, it indicates that the corresponding cluster error sum value will not change much when the cluster quantity changes, so the i+1-th cluster quantity can be taken as the selected cluster quantity.

[0090] In some embodiments, the foregoing S133 can specifically include S1331-S1333.

[0091] S1331: take the (i+1)th cluster number as a candidate cluster number, and determine the upper and lower limits of the cluster cluster including the water quality evaluation index corresponding to the candidate cluster number as candidate upper and lower limits.

[0092] S1332: determine whether the water quality evaluation index is in the middle region determined based on the candidate upper and lower limits, the upper and lower limits of the middle region being within the candidate upper and lower limits; if yes, execute S1333; if no, execute S134.

[0093] S1333: take the (i+1)th cluster number as the selected cluster number.

[0094] In actual application, there may be a problem that the upper and lower limits of the cluster cluster including the water quality evaluation index corresponding to the (i+1)th cluster number are too close to the water quality evaluation index, and it is unreasonable to use this upper and lower limit. To avoid this problem, after obtaining the upper and lower limits including the water quality evaluation index corresponding to the foregoing candidate cluster number (that is, the candidate upper and lower limits), it is first determined whether the water quality evaluation index is in the middle region determined based on the candidate upper and lower limits. In the case where it is determined that the water quality evaluation index is in the middle region determined based on the candidate upper and lower limits, the (i+1)th cluster number is taken as the selected cluster number, and the candidate upper and lower limits are taken as the selected upper and lower limits. If it is determined that the water quality evaluation index is not in the middle region determined based on the candidate upper and lower limits, it is continued to execute S134 to determine whether the cluster number behind can be taken as the selected cluster number.

[0095] The foregoing middle region is a region in which the upper and lower limits are located in the middle of the candidate upper and lower limits. In specific implementation, the middle region can be a middle third region located in the region determined based on the candidate upper and lower limits.

[0096] In some other embodiments, the computing device determines the selected cluster number by using S135-S138 as follows.

[0097] S135: determine the (i)th cluster number and the (i+1)th cluster number after the elbow cluster number.

[0098] wherein the initial value of i is 0.

[0099] S136: calculate the difference value of the cluster error sum values corresponding to the (i)th cluster number and the (i+1)th cluster number, and determine whether the difference value is less than a set difference threshold value; if yes, execute S137; if no, execute S138.

[0100] S137: take the (i+1)th cluster number as the selected cluster number.

[0101] S138: i = i + 1 is made, and S131 is executed.

[0102] Comparing S135-S138 and the foregoing S131-S134, only in determining the difference value of the sum of the cluster errors, it is judged whether the difference value is less than the set difference threshold, which is different from the foregoing scheme. It can be conceived that here the absolute numerical comparison is used to determine whether the i+1 cluster number is selected as the selected cluster number.

[0103] The foregoing mentioned that the selected cluster number is determined by calculation, and then the upper and lower limits of the water quality fitting line range are determined. In other embodiments, the cluster error sum and the cluster number can be plotted into a fold line, and then the experienced technician can observe the fold line to select the selected cluster number in the cluster number greater than the elbow cluster number, and then determine the upper and lower limits of the water quality fitting line range according to the selected cluster number.

[0104] The embodiment scheme of the present disclosure determines the cluster set under each cluster number by the method of cluster analysis, determines the elbow cluster number according to the cluster error sum of the cluster set under each cluster number, then determines a larger cluster number as the selected cluster number according to the elbow cluster number, and finally determines the upper and lower limits of the cluster cluster including the water quality evaluation index in the cluster set corresponding to the selected cluster number as the upper and lower limits of the water quality fitting line range. That is, the embodiment scheme of the present disclosure provides a method for determining the upper and lower limits of the water quality fitting line range by the method of cluster analysis. Because the upper and lower limits of the water quality fitting line range determined by the embodiment of the present disclosure are obtained based on the cluster analysis of unsupervised learning, the water quality fitting line range obtained is not determined by human subjective determination, but is determined according to a large amount of water quality monitoring data, which is not easy to predict, so it can better identify whether the fitting line affects the water quality.

[0105] In addition to providing the foregoing water quality fitting line range determination method, the embodiment of the present disclosure also provides a disturbed water quality monitoring point identification method. The disturbed water quality monitoring point identification method uses the water quality fitting line range determined by the water quality fitting line range determination method determined in the foregoing. Figure 2 The disturbed water quality monitoring point identification method flowchart provided by the embodiment of the present disclosure is shown in FIG. 10. As shown in FIG. 10, the disturbed water quality monitoring point identification method includes S210-S240. Figure 2

[0106] S210: The historical monitoring data sequence of the water quality monitoring point to be identified is obtained, and the number of fitting line monitoring data in the historical monitoring data sequence is counted.

[0107] The fitting line monitoring data is the monitoring data located in the water quality fitting line range.

[0108] ​S220: Calculate the line fitting data proportion of the water quality monitoring point to be identified based on the number of line fitting monitoring data and the total amount of data of the historical monitoring data sequence.

[0109] S230: Determine whether the line fitting data proportion is greater than the upper limit threshold of the proportion; if yes, execute S240.

[0110] S240: Determine that the water quality monitoring point to be identified is an interfered water quality monitoring point.

[0111] The embodiments of the present disclosure are based on a premise: if the proportion of line fitting monitoring data in a certain monitoring point is too large, it may have two reasons: (1) it may indeed be that the water quality meets the standard; (2) the water quality is considered to be interfered. However, the possibility of the former (1) is small, and it is more likely that the latter (2) causes the proportion of line fitting monitoring data to be too large, so the water quality monitoring point to be identified is determined to be an interfered water quality monitoring point.

[0112] In executing the foregoing method, it is also necessary to determine the preset upper limit threshold of the proportion. In some embodiments, the upper limit threshold of the proportion is not a determined upper limit threshold, but is objectively determined. In specific implementation, the upper limit threshold of the proportion can be determined by S250-S270 as follows.

[0113] S250: Based on the historical monitoring data sequence of all monitoring points, the line fitting data proportion of each monitoring point is respectively counted.

[0114] S260: Statistically analyze the line fitting data proportion of each monitoring point to determine the mean and mean square deviation of the proportion;

[0115] S270: Determine the upper limit threshold of the proportion based on the mean and mean square deviation of the proportion.

[0116] According to statistical principles, in the case that the number of interfered water quality monitoring points in the monitoring points is small, the line fitting data proportion thereof does not conform to the characteristics of the line fitting data proportion in a large number of monitoring points. Based on this, the mathematical statistical characteristics of the line fitting data proportion of most monitoring points are determined in the embodiments of the present disclosure, and a reasonable line fitting data proportion range is determined based on the mathematical statistical characteristics. If the line fitting data proportion of a certain monitoring point is located outside the foregoing reasonable line fitting data proportion range, it is determined that this monitoring point is an interfered water quality monitoring point. Accordingly, the upper limit threshold of the proportion needs to be determined based on the mathematical statistical characteristics. In specific implementation, the upper limit threshold of the proportion can be obtained by adding the mean and three times the mean square deviation of the proportion obtained by statistically analyzing the line fitting data proportion of each monitoring point.

[0117] In actual implementation, various types of quality monitoring need to be performed on water quality, such as monitoring of potassium permanganate index, COD content, phosphorus content, and ammonia nitrogen content. Accordingly, S210 can be acquiring various types of water quality monitoring data of a water quality monitoring point to be identified, and respectively counting the number of in-line monitoring data in the various types of water quality monitoring data; S220 can be respectively calculating the in-line data proportion of the various types of water quality monitoring data at the water quality monitoring point to be identified; S230 can be respectively judging whether the in-line data proportion of the various types of water quality monitoring data is greater than a corresponding preset upper proportion threshold; and S240 can be specifically: in a case where the in-line data proportion of one type of water quality monitoring data is greater than the corresponding preset upper proportion threshold, determining that the water quality monitoring point to be identified is an interfered water quality monitoring point.

[0118] In one specific implementation, the upper and lower limits of the in-line range determined by the foregoing method are as shown in Table 1

[0119] Table 1: Upper and lower limits of in-line range

[0120]

[0121] After obtaining the data in Table 1, the upper and lower limits determined according to Table 1 are used to determine whether the water quality monitoring point is interfered. Through on-site investigation, it is found that there is a small amount of white floating material around the sampling port of a monitoring point, and a box containing white unknown objects is placed on the shore, and it is determined that this water quality monitoring point is suspected to be interfered. Backtracking the historical monitoring data sequence of this water quality monitoring point, 30 monitoring data are obtained, of which 21 groups of data are within the in-line range, that is, 70% of the data are within the in-line range, which is not normal, and therefore it is determined that this water quality monitoring point is an interfered monitoring point.

[0122] In specific implementation, the historical monitoring data of each monitoring point can also be statistically analyzed, and when the data evaluation of a monitoring point changes from qualified to 20%-30% unqualified, the foregoing method is used to investigate it, or when the data evaluation changes from qualified to continuous over-standard and then to in-line and meets the data evaluation index, the foregoing method is used to investigate it.

[0123] The embodiments of the present disclosure also provide a computing device for implementing the foregoing method. Figure 3 FIG. 1 is a structural schematic diagram of a computing device provided by the embodiments of the present disclosure. The following specifically refers to Figure 3 FIG. 1 is a structural schematic diagram of a computing device provided by the embodiments of the present disclosure. The following specifically refers to Figure 3 The computing device shown is merely an example, and should not impose any limitation on the functions and use range of the embodiments of the present disclosure.

[0124] As shown in FIG. 1, the computing device 100 includes a processor 101, a memory 102, and a communication interface 103. Figure 3As shown, computing device 300 can include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 301 that can perform various appropriate actions and processes according to programs stored in read-only memory (ROM) 302 or loaded into random access memory (RAM) 303 from storage device 308. Various programs and data required for operation of computing device 300 are also stored in RAM 303. Processing device 301, ROM 302, and RAM 303 are connected to each other by bus 304. An input / output (I / O) interface 305 is also connected to bus 304.

[0125] Generally, the following devices can be connected to I / O interface 305: input devices 306 including, for example, a touch screen, a touch pad, a camera, a microphone, etc.; output devices 307 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; storage devices 308 including, for example, a magnetic tape, a hard disk, etc.; and communication devices 309. Communication devices 309 can allow computing device 300 to communicate with other devices wirelessly or via wires to exchange data. Although Figure 3 Computing device 300 is shown with various devices, but it should be understood that not all of the shown devices are required to be implemented or present. More or fewer devices can alternatively be implemented or present.

[0126] In particular, according to embodiments of the present disclosure, the processes described above with reference to the flowcharts can be implemented as a computer software program. For example, embodiments of the present disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods illustrated by the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication devices 309, or installed from storage devices 308, or installed from ROM 302. When the computer program is executed by processing device 301, the above-described functions defined in the methods of embodiments of the present disclosure are performed.

[0127] It should be noted that the computer-readable medium described above in the present disclosure can be a computer-readable storage medium, a computer-readable signal medium, or, or any combination of the two.

[0128] Computer-readable storage media, for example, can be any media that can be accessed by a computer. Such computer-readable storage media, for example, can include but are not limited to semiconductor memories such as flash memory and electrically programmable read-only memory (EPROM), random access memories (RAMs), CD-ROMs, DVDs, and diskettes. As used herein, the term "memory" includes one or more computer-readable storage media. The use of the word "memory" does not limit the memory to being entirely of one type of medium, or drifting between multiple types of media. The computer-readable storage media, for example, can be volatile or non-volatile memory, removable or non-removable memory, erasable or non-erasable memory, writeable or re-writeable memory, and / or the like.

[0129] The computer-readable signal media can include a computer-readable storage medium that is non-transitory or tangible. The computer-readable storage medium can include a computer-readable storage medium having program code stored therein, which program code is executable by a computer-reading device or computer-reading system to cause the computer-reading device or computer-reading system to carry out a computer-implemented process. The computer-readable signal media can also be any computer-readable medium that is not a computer-readable storage medium and that can communicate, propagate or transport program code such as apparatuses for program code (e.g., electromagnetic radiation).

[0130] In some embodiments, the client, server, and / or other components or processes can communicate information via computer network 130. In some embodiments, the client, server, and / or other components or processes can communicate information via any known or future developed network protocol, including but not limited to HTTP (HyperText Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of a communication network include a local area network ("LAN"), a wide area network ("WAN"), the Internet, and / or peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any current or future developed network. N N In some embodiments, the client, server, and / or other components or processes can communicate information via computer network 130. In some embodiments, the client, server, and / or other components or processes can communicate information via any known or future developed network protocol, including but not limited to HTTP (HyperText Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of a communication network include a local area network ("LAN"), a wide area network ("WAN"), the Internet, and / or peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any current or future developed network.

[0131] The computer-readable medium can be included within the above-described computing device; or can be separate from the computing device and exist outside the computing device.

[0132] ​Computer program code for carrying out operations of the present disclosure can be written in any combination of one or more programming languages, including an object oriented programming language such as Java, Smalltalk, C++ or the like and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the testee's computer, partly on the testee's computer, as a stand-alone software package, partly on the testee's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the testee's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider). N ) or a wide area network (WAN), or can be connected to an external computer (for example, through the Internet using an Internet Service Provider). N

[0133] The flow diagrams and the block diagrams in the drawings are illustrations of architectures, functionalities, and operations of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flow diagrams or block diagrams can represent a module, a segment, or a portion of code, which comprises one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently or the blocks may

[0134] The units described in the embodiments of the present disclosure can be implemented by software, or by hardware. In some cases, the names of the units do not constitute a limitation on the units themselves. The functions described above can be performed at least in part by one or more hardware logic components. For example, non-limiting examples of hardware logic components that can be used include field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SOCs), complex programmable logic devices (CPLDs), etc. I

[0135] ​​The foregoing is merely illustrative of the various ways and specific embodiments in which the disclosure can be carried out. Numerous modifications can be made to these specific embodiments and implementations without departing from the spirit and scope of the disclosure. It is, therefore, intended that this disclosure not be limited to the particular implementation described exclusively hereinabove but that all such alterations and modifications that come within the scope of the present disclosure are covered by the appended claims.

Claims

1. A water quality range determination method characterized by comprising: The method comprises the following steps: The water quality monitoring data of the existing monitoring points are clustered by using different cluster numbers to determine the cluster sets of the water quality monitoring data under different cluster numbers; The cluster error sum values corresponding to the cluster sets under different cluster numbers are calculated respectively, and the elbow cluster number is determined based on the cluster error sum values corresponding to different cluster numbers; The selected cluster number is selected from the cluster numbers greater than the elbow cluster number, and the upper and lower limits of the cluster clusters including the water quality evaluation indexes in the cluster set corresponding to the selected cluster number are taken as the upper and lower limits of the water quality fitting line range, and the selected cluster number is selected from the cluster numbers greater than the elbow cluster number, comprising: S1: determine the i-th cluster number and the i+1-th cluster number after the elbow cluster number, wherein the initial value of i is 0; S2: calculate the reduction ratio of the cluster error sum values corresponding to the i-th cluster number and the i+1-th cluster number, and determine whether the reduction ratio is less than a preset reduction ratio; if yes, execute S3; if no, execute S4; S3: take the i+1-th cluster number as the selected cluster number; S4: make i=i+1, and execute S1, and the S3 comprises: S31: take the i+1-th cluster number as the candidate cluster number, and determine the upper and lower limits of the cluster clusters including the water quality evaluation indexes in the cluster set corresponding to the candidate cluster number as the candidate upper and lower limits; S32: determine whether the water quality evaluation indexes are in the intermediate region determined based on the candidate upper and lower limits, and the upper and lower limits of the intermediate region are located within the candidate upper and lower limits; if yes, execute S33; if no, execute S4; S33: take the i+1-th cluster number as the selected cluster number.

2. The determination method according to claim 1, characterized in that, The selected cluster number is selected from the cluster numbers greater than the elbow cluster number, comprising: S1: determine the i-th cluster number and the i+1-th cluster number after the elbow cluster number, wherein the initial value of i is 0; S2: calculate the difference value of the cluster error sum values corresponding to the i-th cluster number and the i+1-th cluster number, and determine whether the difference value is less than a set difference threshold; if yes, execute S3; if no, execute S4; S3: take the i+1-th cluster number as the selected cluster number; S4: make i=i+1, and execute S1.

3. The determination method according to any one of claims 1-2, characterized in that, The cluster error sum values corresponding to the cluster sets under different cluster numbers are calculated respectively, comprising: For each cluster cluster under a cluster number, the corresponding intra-cluster center is calculated; Based on the intra-cluster center and the intra-cluster water quality monitoring data of each cluster cluster, the intra-cluster distance sum value is calculated; The cluster error sum value is calculated based on the intra-cluster distance sum values of each cluster cluster under a cluster number.

4. The determination method according to any one of claims 1-2, characterized in that, The elbow cluster number is determined based on the cluster error sum values corresponding to different cluster numbers, comprising: The first-order difference value of adjacent cluster numbers is calculated, and the difference value ratio of the next first-order difference value to the previous adjacent first-order difference value is calculated; The first cluster number making the difference value ratio less than a preset ratio is selected as the elbow cluster number.

5. A method for identifying an interfered water quality monitoring point, characterized in that, The method comprises the following steps: acquire a historical monitoring data sequence of a water quality monitoring point to be identified, and count a number of line-fitting monitoring data in the historical monitoring data sequence; the line-fitting monitoring data is monitoring data located in a water quality line-fitting range, and the water quality line-fitting range is determined by the water quality line-fitting range determination method in any one of claims 1-4; calculate a line-fitting data proportion of the water quality monitoring point to be identified based on the number of line-fitting monitoring data and a total amount of data of the historical monitoring data sequence; determine whether the line-fitting data proportion is greater than an upper limit threshold of the proportion; in a case where the line-fitting data proportion is greater than the upper limit threshold of the proportion, determine that the water quality monitoring point to be identified is an interfered water quality monitoring point.

6. The identification method according to claim 5, characterized in that, Before the determination of whether the line-fitting data proportion is greater than the upper limit threshold of the proportion, the method further comprises: count the line-fitting data proportion of each monitoring point based on historical monitoring data sequences of all monitoring points; perform statistical analysis on the line-fitting data proportion of each monitoring point to determine a mean value of the proportion and a mean square deviation of the proportion; determine the upper limit threshold of the proportion based on the mean value of the proportion and the mean square deviation of the proportion.

7. The identification method according to claim 5, characterized in that the acquisition of the historical monitoring data sequence of the water quality monitoring point to be identified and the counting of the number of line-fitting monitoring data in the historical monitoring data sequence comprises: acquisition of various types of water quality monitoring data of the water quality monitoring point to be identified, and counting of the number of line-fitting monitoring data in each type of water quality monitoring data; the calculation of the line-fitting data proportion of the water quality monitoring point to be identified comprises: calculation of the line-fitting data proportion of each type of water quality monitoring data at the water quality monitoring point to be identified; the determination of whether the line-fitting data proportion is greater than the preset upper limit threshold of the proportion comprises: determination of whether the line-fitting data proportion of each type of water quality monitoring data is greater than the corresponding preset upper limit threshold of the proportion; in a case where the line-fitting data proportion is greater than the upper limit threshold of the proportion, the determination that the water quality monitoring point to be identified is the interfered water quality monitoring point comprises: in a case where the line-fitting data proportion of one type of water quality monitoring data is greater than the corresponding preset upper limit threshold of the proportion, the determination that the water quality monitoring point to be identified is the interfered water quality monitoring point.

8. A computing device, comprising: a processor and a memory, the memory being used to store a computer program; the computer program, when loaded by the processor, causes the processor to execute the water quality line-fitting range determination method in any one of claims 1-4 or the interfered water quality monitoring point identification method in any one of claims 5-7.

Citation Information

Patent Citations

  • Data clustering method and device, storage medium and electronic equipment

    CN114330584A

  • Water quality reference value setting method and system

    CN119046705A