Abnormal value identification method and device for monitoring data of underground water and surface water pollutants
By constructing a pollutant network and using the analytic hierarchy process (AHP) and centrality indicators to identify anomalous nodes, the problem of inaccurate identification of outliers in groundwater and surface water pollutant monitoring data in existing technologies has been solved, enabling more accurate pollutant monitoring and timely remedial measures.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-11
- Publication Date
- 2026-04-03
AI Technical Summary
Existing technologies lack effective means to identify outliers in groundwater and surface water pollutant monitoring data, leading to inaccurate analysis results and hindering timely detection of pollution and the implementation of remedial measures.
By constructing a pollutant network and using a visual algorithm, combined with the analytic hierarchy process (AHP) and centrality index, abnormal nodes in the pollutant network can be identified, enabling effective identification of outliers in groundwater and surface water pollutant monitoring data.
It improves the accuracy of pollutant monitoring data analysis, enabling timely detection of pollution and the implementation of remedial measures to protect the aquatic environment.
Smart Images

Figure CN121786678A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of groundwater and surface water pollutant monitoring technology, and in particular to a method and apparatus for identifying outliers in groundwater and surface water pollutant monitoring data. Background Technology
[0002] Groundwater and surface water pollution monitoring is an important part of water resource management. Regularly or in real time, the collection and analysis of groundwater and surface water pollutant monitoring data is beneficial for timely detection of pollution and the implementation of remedial measures, thereby achieving the protection of the water environment.
[0003] However, due to the complex sources and large volume of groundwater and surface water pollutant monitoring data, existing technologies lack effective means to identify outliers in groundwater and surface water pollutant monitoring data (such as outliers caused by drastic fluctuations in the content of pollutants in groundwater). This results in inaccurate analysis results for pollutant monitoring data containing outliers, which is not conducive to timely detection of pollution and the implementation of remedial measures. Summary of the Invention
[0004] This invention proposes a method and apparatus for identifying outliers in groundwater and surface water pollutant monitoring data. By constructing a pollutant network of groundwater and surface water using a visual graph algorithm, and then using the analytic hierarchy process (AHP) and centrality index to identify multiple outlier nodes in the pollutant network, the invention achieves effective identification of outliers in groundwater and surface water pollutant monitoring data. This improves the accuracy of the analysis results of pollutant monitoring data and facilitates timely detection of pollution and the implementation of remedial measures.
[0005] To achieve the above objectives, the present invention adopts the following technical solution: In a first aspect, this invention provides a method for identifying outliers in groundwater and surface water pollutant monitoring data, comprising: acquiring pollutant monitoring data for groundwater and surface water; the pollutant monitoring data being time-series data containing water quality monitoring data at multiple time points over a period of time. Based on the pollutant monitoring data and a visibility chart algorithm, a pollutant network for groundwater and surface water is constructed; the pollutant network includes multiple network layers, each network layer corresponding one-to-one with water quality monitoring data for one type of pollutant contained in the pollutant monitoring data; each network layer includes multiple nodes and multiple edges connecting the multiple nodes; each of the multiple nodes corresponds one-to-one with water quality monitoring data at a specific time point in the water quality monitoring data for one type of pollutant; the principle for establishing each edge is that it does not intersect with other edges in the pollutant network and satisfies the visibility conditions of the visibility chart algorithm. Based on the analytic hierarchy process (AHP) and the centrality index of each node in the pollutant network, multiple outlier nodes in the pollutant network are identified; wherein, the water quality monitoring data at the time point corresponding to each of the multiple outlier nodes is considered an outlier.
[0006] The outlier identification method for groundwater and surface water pollutant monitoring data provided by this invention first uses a visual graph algorithm to combine groundwater and surface water pollutant monitoring data to construct a pollutant network for groundwater and surface water. Then, by combining the analytic hierarchy process (AHP) and the centrality index of each node in the pollutant network, multiple outlier nodes in the pollutant network are identified. This method effectively identifies outliers in groundwater and surface water pollutant monitoring data, thereby improving the accuracy of the analysis results of pollutant monitoring data. It is beneficial for timely detection of pollution and the implementation of remedial measures, thereby achieving the protection of the water environment.
[0007] In one implementation of the first aspect, identifying multiple anomalous nodes in the pollutant network includes: calculating the centrality index of each node in the pollutant network. The analytic hierarchy process (AHP) is used to determine the importance weight of the centrality index corresponding to each node in the pollutant network. For each node in the pollutant network, the centrality index of the node and the importance weight of the corresponding centrality index are weighted and summed to obtain the importance index of each node in the pollutant network. Based on the importance index of each node in the pollutant network, the multiple nodes contained in the pollutant network are ranked to obtain the node importance ranking result of the pollutant network. Based on the ranking of node importance in the pollutant network and the water quality monitoring data at multiple time points corresponding to multiple nodes in the pollutant network, multiple abnormal nodes in the pollutant network are identified.
[0008] In one implementation of the first aspect, the analytic hierarchy process (AHP) is used to determine the importance weight of each node in the pollutant network, including: Step 1: For each network layer in the pollutant network, construct the decision matrix for that network layer. X ,and , x ij Decision matrix X The element in the i-th row and j-th column, m This refers to the number of nodes contained in the network layer. n The number of centrality indicators; Step 2: Use standardized formulas to convert the decision matrix. X Transform into a standardized decision matrix R ,and R = , r ij Standardized decision matrix R The element in the i-th row and j-th column; Step 3: Using the 1-9 scale, evaluate the relative importance of each node in the network layer relative to each centrality metric, and obtain the relative importance matrix of the network layer. A Relative importance matrix , This indicates the relative importance between the first node in the network layer and the second indicator in the centrality index. This indicates the relative importance between the first node in the network layer and the nth centrality metric. This indicates the relative importance between the second node in the network layer and the nth centrality index. Step 4: Calculate the candidate importance weight for each node in the network layer; the formula for calculating the candidate importance weight is as follows: in, This represents the candidate importance weight of the i-th node among multiple nodes contained in the network layer, and ; Represents the relative importance matrix of network layers A The element in the i-th row and j-th column of the array; Step 5: Perform a consistency check on the candidate importance weights of each node among the multiple nodes contained in the network layer, and obtain the consistency check results; If the consistency test result is inconsistent, repeat steps 1 to 4. When the consistency test result is consistent, the standardized decision matrix will be... R The importance weight of each node among the multiple nodes contained in the network layer is determined; Step 6: Combine the importance weights of each node in the multiple nodes contained in the multiple network layers to obtain the importance weight of each node in the pollutant network.
[0009] In one implementation of the first aspect, a consistency check is performed on the candidate importance weights of each node among the multiple nodes contained in the network layer, and the consistency check result is obtained, including: Calculate the consistency ratio of network layers CR Consistency ratio CR Satisfy the following formula: in, It is the consistency index of the network layer, and , , n This refers to the number of nodes contained in the network layer. A This is the relative importance matrix of network layers. The weight vector of the network layer. Relative importance moments of network layers A With the weight vector of the network layer The i-th element in the product result Let be the candidate importance weight of the i-th node among the multiple nodes contained in the network layer. The average random consistency index; When the consistency ratio CR If the consistency test conditions are met, the consistency test result is consistent; otherwise, the consistency test result is inconsistent.
[0010] In one implementation of the first aspect, the consistency check condition is the consistency ratio. CR< 0.1.
[0011] In one implementation of the first aspect, based on the node importance ranking result of the pollutant network and the water quality monitoring data at multiple time points corresponding to multiple nodes contained in the pollutant network, multiple abnormal nodes in the pollutant network are identified, including: identifying multiple nodes whose node importance ranks in the top 10% of the node importance ranking result as multiple abnormal nodes in the pollutant network.
[0012] In one implementation of the first aspect, Centrality metrics include degree centrality, tight centrality, and betweenness centrality.
[0013] The formula for calculating degree centrality is as follows: The formula for calculating compact centrality is as follows: The formula for calculating betweenness centrality is as follows: in, This represents the degree centrality of the i-th node within one of the multiple network layers in a pollutant network. This refers to the number of nodes contained in the network layer. Let be the degree of the i-th node; Indicate the compact centrality of the i-th node. This represents the shortest path length from the i-th node to the j-th node; Indicate the betweenness centrality of the i-th node. This represents the number of shortest paths from the s-th node to the t-th node. This represents the number of shortest paths from the s-th node to the t-th node that pass through the i-th node.
[0014] Secondly, this invention provides an outlier identification device for groundwater and surface water pollutant monitoring data, including an acquisition module, a pollutant network construction module, and an outlier identification module. The acquisition module is used to acquire groundwater and surface water pollutant monitoring data; the pollutant monitoring data is time-series data containing water quality monitoring data at multiple time points over a period of time. The pollutant network construction module is used to construct a pollutant network for groundwater and surface water based on the pollutant monitoring data and a visibility algorithm; the pollutant network includes multiple network layers, each network layer corresponding one-to-one with water quality monitoring data for one type of pollutant contained in the pollutant monitoring data; each network layer includes multiple nodes and multiple edges connecting the multiple nodes; each node corresponds one-to-one with water quality monitoring data at one time point in the water quality monitoring data for one type of pollutant; the principle for establishing each edge is: it does not intersect with other edges in the pollutant network and satisfies the visibility conditions of the visibility algorithm. The outlier identification module is used to identify multiple outlier nodes in the pollutant network based on the analytic hierarchy process and the centrality index of each node in the pollutant network; among these, the water quality monitoring data at the time point corresponding to each outlier node is considered an outlier.
[0015] Thirdly, the present invention provides an electronic device including a processor and a memory coupled to the processor; the memory is used to store computer instructions, and when the electronic device is running, the processor executes the computer instructions stored in the memory to cause the electronic device to perform the method described in the first aspect above or any implementation thereof.
[0016] Fourthly, the present invention provides a computer-readable storage medium including computer program instructions that, when executed by a computer, cause the computer to perform the method described in the first aspect above or any implementation thereof.
[0017] Fifthly, the present invention provides a computer program product, including computer program instructions, which, when executed on a computer, cause the computer to perform the method described in the first aspect above or any implementation thereof.
[0018] The technical effects corresponding to the second to fifth aspects and their possible implementations can be referred to the above description of the technical effects of the first aspect and its possible implementations, and will not be repeated here. Attached Figure Description
[0019] Figure 1 This is one of the schematic diagrams of the outlier identification method for groundwater and surface water pollutant monitoring data provided in the embodiments of this application; Figure 2 This is the second schematic diagram of the outlier identification method for groundwater and surface water pollutant monitoring data provided in the embodiments of this application; Figure 3 This is a schematic diagram illustrating the change of potassium permanganate index over time, provided in an embodiment of this application. Figure 4 This is a schematic diagram of an abnormal time provided in an embodiment of this application; Figure 5 This is a schematic diagram of the outlier identification device for groundwater and surface water pollutant monitoring data provided in this application embodiment. Detailed Implementation
[0020] In the specification and claims of this invention, the terms "first" and "second," etc., are used to distinguish different objects, rather than to describe a specific order of objects.
[0021] In the embodiments of this application, "and / or" indicates a relationship between objects. For example, A and / or B can represent the following three situations: A exists alone, B exists alone, and A and B exist simultaneously.
[0022] In the embodiments of this application, the terms "exemplary" or "for example" are used to indicate that something is an example, illustration, or description. Any embodiment or design that is described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design. Specifically, the use of the terms "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.
[0023] In the description of this invention, unless otherwise stated, "a plurality of" means two or more. For example, a plurality of nodes means two or more nodes.
[0024] The methods and apparatus provided in this application relate to the monitoring of pollutants in groundwater and surface water, and can be used to identify outliers in groundwater and surface water pollutant monitoring data.
[0025] To address the problem that existing technologies lack effective means for identifying outliers in groundwater and surface water pollutant monitoring data, leading to inaccurate analysis results and hindering timely detection of pollution and remedial measures, this application provides an outlier identification method and apparatus for groundwater and surface water pollutant monitoring data. By constructing a pollutant network for groundwater and surface water using a visual graph algorithm, and then identifying multiple outlier nodes in the pollutant network based on the analytic hierarchy process (AHP) and centrality indices, this method effectively identifies outliers in groundwater and surface water pollutant monitoring data, thereby improving the accuracy of pollutant monitoring data analysis and facilitating timely detection of pollution and remedial measures.
[0026] For example, the outlier identification method for groundwater and surface water pollutant monitoring data provided in this embodiment of the invention can be executed by an electronic device with processing capabilities, such as a computer or server. Taking a computer as an example, the hardware components of the computer may include: a processor, memory, network interface, user interface, communication bus, etc.
[0027] The processor controls the electronic equipment to perform related processing and computational tasks, such as acquiring pollutant monitoring data, constructing pollutant networks for groundwater and surface water, and identifying outliers in the pollutant monitoring data. The processor may include a central processing unit (CPU) or other processors, and may be single-core or multi-core; for example, a processor may include multiple CPUs.
[0028] Memory is used to store computer instructions and related data, such as pollutant monitoring data, pollutant networks, and outliers in pollutant monitoring data. Memory can be random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, optical storage, disk storage media, or other magnetic storage devices, or any other medium capable of storing program code or data accessible by a computer. Optionally, memory can be integrated into the processor, or it can be independent of the processor.
[0029] A network interface is used for communication between a computer and other devices or communication networks. A network interface can be a transceiver with transmit and receive capabilities. Optionally, a network interface may include standard wired interfaces or wireless interfaces (such as Wi-Fi interfaces, Bluetooth interfaces, and 5G interfaces).
[0030] The communication bus is used to enable communication between different components. For example, the processor, memory, network interface and user interface mentioned above can be interconnected through the communication bus.
[0031] The user interface may include a display screen and an input unit (such as a keyboard). Optionally, the user interface may also include a standard wired interface or a wireless interface.
[0032] Those skilled in the art will understand that the computer described above may include more or fewer components, or combine certain components, or have different component arrangements; the embodiments of this application do not limit this.
[0033] like Figure 1 As shown in the embodiments of this application, the outlier identification method for groundwater and surface water pollutant monitoring data includes S101-S103.
[0034] S101. Obtain pollutant monitoring data for groundwater and surface water.
[0035] The aforementioned pollutant monitoring data is time-series data containing water quality monitoring data at multiple points within a certain period. This water quality monitoring data may include permanganate index, as well as data related to pollutants such as heavy metals, phenols, cyanides, and organochlorine pesticides. The aforementioned pollutants may include permanganate, heavy metals, phenols, cyanides, and organochlorine pesticides. The aforementioned groundwater and surface water pollutant monitoring data can be obtained directly through monitoring equipment deployed around the groundwater and surface water areas, or through a data center storing the pollutant monitoring data. This application embodiment does not limit the type of pollutant monitoring data or the method of acquisition.
[0036] S102. Based on pollutant monitoring data and visualization algorithms, construct pollutant networks for groundwater and surface water.
[0037] The pollutant network consists of multiple network layers. Each network layer corresponds one-to-one with the water quality monitoring data of a specific pollutant. Each network layer includes multiple nodes and multiple edges connecting these nodes. Each node corresponds one-to-one with the water quality monitoring data of a specific pollutant at a specific time point. The principle for establishing each edge is that it does not intersect with other edges in the pollutant network and satisfies the visibility conditions of the visibility algorithm.
[0038] The aforementioned "no intersection with other edges in the pollutant network" means that the line connecting two nodes (e.g., points i and j in the time sequence, i < j) does not intersect with the line connecting any intermediate node k (i < k < j). The visibility condition of the aforementioned visual algorithm can be that the line connecting two points is unobstructed above / below the sequence curve. Furthermore, since the visual algorithm is a commonly used algorithm in this technical field, the embodiments of this application will not elaborate further on the visibility condition of the aforementioned visual algorithm.
[0039] Understandably, in this application embodiment, before constructing the pollutant network for groundwater and surface water using pollutant monitoring data, the aforementioned pollutant monitoring data was preprocessed. This preprocessing includes missing value imputation and data format rectification. Specifically, missing value imputation uses interpolation to fill in missing high-dimensional time series data, while data format rectification standardizes the format of anomalous time series data from groundwater and surface water pollutant monitoring.
[0040] It should be noted that, since the visual image algorithm is a commonly used technique in this technical field, the construction process of the above-mentioned pollutant network will not be described in detail in the embodiments of this application.
[0041] S103. Based on the analytic hierarchy process and the centrality index of each node in the pollutant network, identify multiple anomalous nodes in the pollutant network.
[0042] In one implementation, combined with Figure 1 ,like Figure 2 As shown, S103 includes S1031-S1035.
[0043] S1031. Calculate the centrality index of each node in the pollutant network.
[0044] In this embodiment of the application, the centrality indices include degree centrality, compact centrality, and betweenness centrality.
[0045] The formula for calculating degree centrality is as follows: The formula for calculating compact centrality is as follows: The formula for calculating betweenness centrality is as follows: in, This represents the degree centrality of the i-th node within one of the multiple network layers in a pollutant network. This refers to the number of nodes contained in the network layer. Let be the degree of the i-th node; Indicate the compact centrality of the i-th node. This represents the shortest path length from the i-th node to the j-th node; Indicate the betweenness centrality of the i-th node. This represents the number of shortest paths from the s-th node to the t-th node. This represents the number of shortest paths from node s to node t that pass through node i. S1032. Using the analytic hierarchy process (AHP), determine the importance weights of the centrality indicators corresponding to each node in the pollutant network.
[0046] Optionally, S1032 above includes the following steps.
[0047] Step 1: For each network layer in the pollutant network, construct the decision matrix for that network layer. X ,and , xij Decision matrix X The element in the i-th row and j-th column, m This refers to the number of nodes contained in the network layer. n The number of centrality indicators; Step 2: Use standardized formulas to convert the decision matrix. X Transform into a standardized decision matrix R ,and R = , r ij Standardized decision matrix R The element in the i-th row and j-th column; Step 3: Using the 1-9 scale, evaluate the relative importance of each node in the network layer relative to each centrality metric, and obtain the relative importance matrix of the network layer. A Relative importance matrix , This indicates the relative importance between the first node in the network layer and the second indicator in the centrality index. This indicates the relative importance between the first node in the network layer and the nth centrality metric. This indicates the relative importance between the second node in the network layer and the nth centrality index. Step 4: Calculate the candidate importance weight for each node in the network layer; the formula for calculating the candidate importance weight is as follows: in, This represents the candidate importance weight of the i-th node among multiple nodes contained in the network layer, and ; Represents the relative importance matrix of network layers A The element in the i-th row and j-th column of the array; Step 5: Perform a consistency check on the candidate importance weights of each node among the multiple nodes contained in the network layer, and obtain the consistency check results; If the consistency test result is inconsistent, repeat steps 1 to 4. When the consistency test result is consistent, the standardized decision matrix will be... R The importance weight of each node among the multiple nodes contained in the network layer is determined.
[0048] In one application scenario, step 5 above includes the following two sub-steps.
[0049] Step 5.1: Calculate the consistency ratio of the network layers. CRConsistency ratio CR Satisfy the following formula: in, It is the consistency index of the network layer, and , , n This refers to the number of nodes contained in the network layer. A This is the relative importance matrix of network layers. The weight vector of the network layer. Relative importance moments of network layers A With the weight vector of the network layer The i-th element in the product result Let be the candidate importance weight of the i-th node among the multiple nodes contained in the network layer. The average random consistency index; Step 5.2, when the consistency ratio CR If the consistency test conditions are met, the consistency test result is consistent; otherwise, the consistency test result is inconsistent.
[0050] For example, the consistency test condition described above can be the consistency ratio. CR< 0.1.
[0051] Step 6: Combine the importance weights of each node in the multiple nodes contained in the multiple network layers to obtain the importance weight of each node in the pollutant network.
[0052] S1033. For each node in the pollutant network, the centrality index of the node and the importance weight of the centrality index corresponding to the node are weighted and summed to obtain the importance index of each node in the pollutant network.
[0053] S1034. Based on the importance index of each node in the pollutant network, sort the multiple nodes contained in the pollutant network to obtain the node importance ranking result of the pollutant network.
[0054] S1035. Based on the ranking of node importance in the pollutant network and the water quality monitoring data at multiple time points corresponding to multiple nodes in the pollutant network, identify multiple abnormal nodes in the pollutant network.
[0055] The water quality monitoring data for each of the above-mentioned abnormal nodes at the corresponding time points are abnormal values.
[0056] In one application scenario, S1035 above includes the following:
[0057] Multiple nodes whose importance ranks in the top 10% of the node importance ranking results are identified as multiple abnormal nodes in the pollutant network.
[0058] It should be noted that in the visual algorithm, the value of the above node importance is determined by the node's connection pattern. This means that the connection patterns of the multiple nodes whose importance ranks in the top 10% of the node importance ranking are significantly different from other nodes. Therefore, it can be determined that the above multiple nodes are abnormal relative to the pollutant network, and thus they can be identified as multiple abnormal nodes in the pollutant network.
[0059] It is understandable that after obtaining multiple outliers in the pollutant monitoring data of groundwater and surface water, the outliers and their corresponding time points can be represented in a table or in an image. This application embodiment does not limit the output and display method of the outliers identified above.
[0060] In one embodiment of this application, the pollutant is permanganate, and the water quality monitoring data at each of the multiple time points included in the pollutant monitoring data is the permanganate index. A schematic diagram illustrating the change of the permanganate index over time is shown below. Figure 3 As shown, the obtained anomaly times and numbers are listed in Table 1 below. Figure 4 As shown.
[0061] Table 1. Schematic diagram of abnormal times
[0062] In summary, the outlier identification method for groundwater and surface water pollutant monitoring data provided in this application first uses a visual graph algorithm to combine groundwater and surface water pollutant monitoring data to construct a pollutant network for groundwater and surface water. Then, by combining the analytic hierarchy process (AHP) and the centrality index of each node in the pollutant network, multiple outlier nodes in the pollutant network are identified. This method effectively identifies outliers in groundwater and surface water pollutant monitoring data, thereby improving the accuracy of the analysis results of pollutant monitoring data. It is beneficial for timely detection of pollution and the implementation of remedial measures, ultimately achieving the protection of the water environment.
[0063] Accordingly, embodiments of this application provide an outlier identification device for groundwater and surface water pollutant monitoring data, such as... Figure 5 As shown, it includes an acquisition module 501, a pollutant network construction module 502, and an outlier identification module 503.
[0064] The acquisition module 501 is used to acquire pollutant monitoring data of groundwater and surface water; the pollutant monitoring data is time-series data containing water quality monitoring data at multiple time points over a period of time. For example, the acquisition module 501 is used to implement S101 of the above method.
[0065] The pollutant network construction module 502 is used to construct pollutant networks for groundwater and surface water based on pollutant monitoring data and a visibility algorithm. The pollutant network includes multiple network layers, each corresponding one-to-one with water quality monitoring data for one type of pollutant included in the pollutant monitoring data. Each network layer includes multiple nodes and multiple edges connecting these nodes. Each node corresponds one-to-one with water quality monitoring data for one pollutant at a specific time point. The principle for establishing each edge is that it does not intersect with other edges in the pollutant network and satisfies the visibility conditions of the visibility algorithm. For example, the pollutant network construction module 502 is used to implement step S102 of the above method.
[0066] The outlier identification module 503 is used to identify multiple outlier nodes in the pollutant network based on the analytic hierarchy process (AHP) and the centrality index of each node in the pollutant network; wherein, the water quality monitoring data corresponding to the time point of each outlier node is an outlier. For example, the outlier identification module 503 is used to implement S103 of the above-mentioned species identification method.
[0067] Each module of the above-mentioned groundwater and surface water pollutant monitoring data anomaly identification device can also be used to perform other steps in the above method embodiments. All relevant content involved in the above method embodiments can be referred to the functional description of the corresponding functional module, and will not be repeated here.
[0068] This application also provides an electronic device, including: a processor and a memory coupled to the processor; the memory is used to store computer instructions, and when the electronic device is running, the processor executes the computer instructions stored in the memory to cause the electronic device to perform the methods described in the above embodiments. The processor can implement the acquisition module 501, the pollutant network construction module 502, and the outlier identification module 503; the memory can also be used for pollutant monitoring data, pollutant networks, and multiple abnormal nodes existing in the pollutant network.
[0069] This application also provides a computer-readable storage medium including a computer program that, when run on a computer, performs the methods described in the above embodiments.
[0070] This application also provides a computer program product, which includes computer program instructions that, when run on a computer, execute the methods described in the above embodiments.
[0071] The various embodiments in this specification are described in a progressive manner. The same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on describing the differences from other embodiments.
[0072] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. A method for identifying outliers in groundwater and surface water pollutant monitoring data, characterized in that, include: Obtain pollutant monitoring data for groundwater and surface water; The pollutant monitoring data is time-series data that includes water quality monitoring data at multiple time points over a period of time. Based on the pollutant monitoring data and the visibility algorithm, a pollutant network for groundwater and surface water is constructed. The pollutant network includes multiple network layers, each of which corresponds one-to-one with the water quality monitoring data of one pollutant included in the pollutant monitoring data. Each network layer includes multiple nodes and multiple edges connecting the multiple nodes. Each of the multiple nodes corresponds one-to-one with the water quality monitoring data of one pollutant at a specific time point. The principle for establishing each edge is that it does not intersect with other edges in the pollutant network and satisfies the visibility condition of the visibility algorithm. Based on the analytic hierarchy process (AHP) and the centrality index of each node in the pollutant network, multiple anomalous nodes in the pollutant network are identified; wherein, the water quality monitoring data at the time point corresponding to each of the multiple anomalous nodes is an outlier.
2. The method as described in claim 1, characterized in that, The determination of multiple abnormal nodes in the pollutant network includes: Calculate the centrality index of each node within the pollutant network; The importance weights of the centrality indicators corresponding to each node in the pollutant network are determined using the analytic hierarchy process (AHP). For each node in the pollutant network, the centrality index of the node and the importance weight of the centrality index corresponding to the node are weighted and summed to obtain the importance index of each node in the pollutant network. Based on the importance index of each node in the pollutant network, the multiple nodes contained in the pollutant network are sorted to obtain the node importance ranking result of the pollutant network. Based on the node importance ranking results of the pollutant network and the water quality monitoring data at multiple time points corresponding to multiple nodes contained in the pollutant network, multiple abnormal nodes in the pollutant network are identified.
3. The method as described in claim 2, characterized in that, The analytic hierarchy process (AHP) is used to determine the importance weight of each node in the pollutant network, including: Step 1: For each of the multiple network layers in the pollutant network, construct the decision matrix of that network layer. X ,and , x ij Decision matrix X The element in the i-th row and j-th column, m This refers to the number of nodes contained in the network layer. n The number of the centrality indices; Step 2: Use a standardized formula to convert the decision matrix into a single matrix. X Transform into a standardized decision matrix R ,and R = , r ij For standardized decision matrix R The element in the i-th row and j-th column; Step 3: Using the 1-9 scaling method, evaluate the relative importance of each node in the network layer to each of the centrality indicators, and obtain the relative importance matrix of the network layer. A The relative importance matrix , This indicates the relative importance between the first node in the network layer and the second indicator in the centrality index. This indicates the relative importance between the first node in the network layer and the nth centrality index. This indicates the relative importance between the second node contained in the network layer and the nth index among the centrality indices; Step 4: Calculate the candidate importance weight for each node among the multiple nodes contained in the network layer; the formula for calculating the candidate importance weight is as follows: in, This represents the candidate importance weight of the i-th node among the multiple nodes contained in the network layer, and ; The matrix representing the relative importance of the network layers A The element in the i-th row and j-th column of the array; Step 5: Perform a consistency check on the candidate importance weights of each node among the multiple nodes contained in the network layer, and obtain the consistency check result; If the result of the consistency test is inconsistent, repeat steps 1 to 4. When the consistency test result is consistent, the standardized decision matrix is... R The importance weight of each node among the multiple nodes contained in the network layer is determined; Step 6: Collect the importance weights of each node in the multiple nodes contained in the multiple network layers to obtain the importance weight of each node in the pollutant network.
4. The method as described in claim 3, characterized in that, The process of performing a consistency check on the candidate importance weights of each node among the multiple nodes contained in the network layer, and obtaining the consistency check result, includes: Calculate the consistency ratio of the network layer. CR The consistency ratio CR Satisfy the following formula: in, Let be the consistency index of the network layer, and , , n This refers to the number of nodes contained in the network layer. A This is the relative importance matrix of the network layers. The weight vector of the network layer. The relative importance moments of the network layers A With the weight vector of the network layer The i-th element in the product result, Let be the candidate importance weight of the i-th node among the multiple nodes contained in the network layer. The average random consistency index; When the consistency ratio CR If the consistency test conditions are met, the consistency test result is consistent; otherwise, the consistency test result is inconsistent.
5. The method as described in claim 4, characterized in that, The consistency test condition is the consistency ratio. CR< 0.
1.
6. The method as described in claim 2, characterized in that, The step of determining multiple abnormal nodes in the pollutant network based on the node importance ranking result of the pollutant network and the water quality monitoring data at multiple time points corresponding to multiple nodes included in the pollutant network includes: Nodes whose importance ranks in the top 10% of the node importance ranking results are identified as multiple abnormal nodes in the pollutant network.
7. The method as described in claim 1 or 2, characterized in that, The centrality indices include degree centrality, compact centrality, and betweenness centrality; The formula for calculating the degree centrality is as follows: The formula for calculating the compact centrality is as follows: The formula for calculating the betweenness centrality is as follows: in, This represents the degree centrality of the i-th node within one of the multiple network layers contained in the pollutant network. This refers to the number of nodes contained in the network layer. Let be the degree of the i-th node; This indicates the compact centrality of the i-th node. This represents the shortest path length from the i-th node to the j-th node; Indicates the betweenness centrality of the i-th node. This represents the number of shortest paths from the s-th node to the t-th node. This represents the number of shortest paths from the s-th node to the t-th node that pass through the i-th node.
8. An outlier identification device for groundwater and surface water pollutant monitoring data, characterized in that, It includes an acquisition module, a pollutant network construction module, and an outlier identification module; The acquisition module is used to acquire pollutant monitoring data of groundwater and surface water; the pollutant monitoring data is time-series data containing water quality monitoring data at multiple time points over a period of time. The pollutant network construction module is used to construct a pollutant network for groundwater and surface water based on the pollutant monitoring data and the visibility algorithm. The pollutant network includes multiple network layers, each of which corresponds one-to-one with the water quality monitoring data of one pollutant included in the pollutant monitoring data. Each network layer includes multiple nodes and multiple edges connecting the multiple nodes. Each of the multiple nodes corresponds one-to-one with the water quality monitoring data of one pollutant at a specific time point. The principle for establishing each edge is that it does not intersect with other edges in the pollutant network and satisfies the visibility condition of the visibility algorithm. The outlier identification module is used to identify multiple outlier nodes in the pollutant network based on the analytic hierarchy process and the centrality index of each node in the pollutant network; wherein, the water quality monitoring data at the time point corresponding to each of the multiple outlier nodes is an outlier.
9. An electronic device, characterized in that, The device includes a processor and a memory coupled to the processor; the memory is used to store computer instructions, which, when the electronic device is running, are executed by the processor to cause the electronic device to perform the method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, It includes computer program instructions that, when executed by a computer, cause the computer to perform the method as described in any one of claims 1 to 7.