A method for tracing and identifying the source of groundwater pollution
By calculating the accumulation coefficient and diffusion impact of groundwater pollutants and combining them with time-delay correlation analysis, the source of pollution can be identified, solving the problem of insufficient real-time identification and early warning in existing technologies for tracing groundwater pollution sources, and achieving accurate location of pollution sources.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-13
- Publication Date
- 2026-04-03
AI Technical Summary
Existing methods for tracing groundwater pollution sources are insufficient to capture the continuous fluctuations and evolution trends of pollutant concentrations over time, and lack the ability to identify pollution sources in real time and provide early warnings.
By acquiring pollutant concentration data at each sampling point, calculating the accumulation coefficient and diffusion impact, screening reference points, and using time-delay correlation analysis to pinpoint the pollution source.
It effectively solves the problem of accurately and timely identifying the sources of groundwater pollution, and provides real-time pollution source identification and early warning capabilities.
Smart Images

Figure CN121434654B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of groundwater pollution source tracing technology, specifically to a groundwater pollution inversion source tracing method. Background Technology
[0002] With industrialization and urbanization, groundwater is facing increasingly severe pollution challenges. Due to the complex sources of pollutants, unclear migration and transformation mechanisms, and difficulty in capturing spatiotemporal evolution trends, accurate source tracing and effective treatment have become extremely difficult.
[0003] Traditional water chemistry diagrams (such as Piper diagrams and Gibbs diagrams) and statistical methods are mostly based on analysis of single or a few sampling data points, essentially a "static snapshot." They are unable to capture the continuous fluctuations, evolution trends, and periodic patterns of pollutant concentrations over time, and cannot reflect the evolutionary behavior of the groundwater system as a dynamic whole; moreover, they are "post-event analysis" after pollution occurs, rather than "process tracking" during the evolution process, lacking the ability to identify and provide early warning of pollution source locations in real time. Summary of the Invention
[0004] To address the technical limitations of existing static analysis methods in tracing and early warning of groundwater pollution sources, this invention aims to provide a groundwater pollution inversion and source tracing method. The specific technical solution adopted is as follows:
[0005] Acquire the concentration data of each pollutant at each sampling point; select pollutants by category as target pollutants;
[0006] Within the current preset sliding time window, based on the distribution of the concentration data of the target pollutant at each sampling point, the accumulation coefficient of the target pollutant at each sampling point is obtained; based on the fluctuation differences of the concentration data of the target pollutant between different sampling points, combined with the differences in the accumulation coefficient, the diffusion influence of the target pollutant at the current time is obtained; the sampling point with the highest overall concentration is selected as the initial sampling point, and based on the distribution of other sampling points relative to the initial sampling point, combined with the diffusion influence and the accumulation coefficient, reference points are selected;
[0007] Based on the time delay performance of the concentration data of the reference point and surrounding sampling points within the current preset time domain interval, the current pollution source is marked.
[0008] Furthermore, the method for obtaining the packing coefficient includes:
[0009] For each sampling point, the difference between each concentration data point and the mean concentration is obtained to form a difference sequence; the difference is classified according to the distribution of the difference in the difference sequence, and the time interval from each type of difference to the current time is obtained;
[0010] In the difference classification where the mean difference is greater than zero, the mean and number of differences for each class are fused to obtain the degree of anomaly; principal component analysis is performed on the two-dimensional dataset consisting of the degree of anomaly and the inverse of the time interval to obtain the variance contribution rate of the first principal component; based on the time interval corresponding to each class of difference, the degree of anomaly is fused, and combined with the variance contribution rate of the first principal component, the packing coefficient of the corresponding sampling point is obtained.
[0011] Furthermore, the method for classifying based on the distribution of differences within the difference sequence includes:
[0012] The difference sequence is clustered using a hierarchical clustering algorithm, with a preset number of clustering layers of 2.
[0013] Furthermore, the method for obtaining the diffusion influence degree includes:
[0014] The diffusion impact is obtained by fusing the DTW distance of the concentration data of the target pollutant between different sampling points and the absolute average deviation of the packing coefficient.
[0015] Furthermore, the method for obtaining the reference point includes:
[0016] Obtain the farthest distance between the initial sampling point and other sampling points, fuse the farthest distance and the diffusion influence to obtain a reference distance, and within the reference distance of the initial sampling point, compare the stacking coefficient and the preset stacking threshold to filter out reference points.
[0017] Furthermore, the method for obtaining the pollution source includes:
[0018] Each reference point is selected as the point to be analyzed, and each surrounding sampling point of the point to be analyzed is selected as a control point; the lag time is obtained within the preset time domain interval with a preset step size;
[0019] Analyze the correlation between the concentration data of the control point and the concentration data of the point to be analyzed under different lag times, obtain the optimal lag time and record the maximum correlation value;
[0020] When the maximum correlation value is greater than the preset correlation threshold, the optimal lag time and the maximum correlation value are combined to obtain the directional score of the control point;
[0021] The probability of a pollution source is obtained by combining all the directional scores of the points to be analyzed; the current pollution source is marked based on the probability of the pollution source.
[0022] Furthermore, the method for obtaining the optimal lag time includes:
[0023] Based on the cross-correlation function, the correlation values between the control point and the point to be analyzed are obtained at different lag times, and the lag time corresponding to the maximum correlation value is selected as the optimal lag time.
[0024] Furthermore, the method for obtaining the probability of the pollution source includes:
[0025] The mean of the directional scores of the point to be analyzed and all the control points is used as the probability of the pollution source.
[0026] Furthermore, the method for marking the pollution source includes:
[0027] When the probability of a pollution source is greater than a preset pollution threshold, the corresponding sampling point is marked as a pollution source.
[0028] Furthermore, the method for obtaining the surrounding sampling points includes:
[0029] A Voronoi diagram is constructed based on the spatial distribution of the sampling points, and other sampling points directly adjacent to the reference point in the diagram are regarded as the surrounding sampling points of the reference point.
[0030] The present invention has the following beneficial effects:
[0031] This invention first acquires the concentration data of various pollutants at each sampling point, providing a basis for subsequent analysis. It then uses a sliding time window to capture and analyze concentration data. The distribution of target pollutant concentration data is analyzed to obtain the accumulation coefficient, which characterizes the current accumulation and anomaly of the pollutant. Furthermore, based on the fluctuation differences in target pollutant concentration data between different sampling points, the diffusion and migration characteristics of the pollutants are inferred. Combined with the differences in the accumulation coefficient, the diffusion characteristics of the pollutants are comprehensively analyzed to obtain the current diffusion influence of the target pollutant, representing the diffusion intensity of the pollutant among these sampling points. Further, based on the distribution of other sampling points relative to the initial sampling point, and combining the diffusion influence and accumulation coefficient, reference points with high correlation to the initial sampling point during the diffusion process and significant pollution accumulation characteristics are selected, providing a reliable basis for locating the pollution source. Finally, based on the time delay performance of the concentration data between the reference point and surrounding sampling points, the current pollution source is marked. This scheme acquires pollutant concentration data through a sliding time window, calculates the accumulation coefficient, and combines the fluctuation differences between sampling points to obtain the diffusion influence. After selecting reference points, time delay correlation analysis is used to locate the pollution source, effectively solving the problem of accurately and promptly identifying the source of groundwater pollution diffusion. Attached Figure Description
[0032] To more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0033] Figure 1 A flowchart illustrating a groundwater pollution inversion and source tracing method provided in one embodiment of the present invention;
[0034] Figure 2 This is a flowchart illustrating a method for obtaining a pollution source according to an embodiment of the present invention. Detailed Implementation
[0035] To further illustrate the technical means and effects adopted by the present invention to achieve its intended purpose, the following, in conjunction with the accompanying drawings and preferred embodiments, details the specific implementation, structure, features, and effects of a groundwater pollution inversion and source tracing method proposed according to the present invention. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. Furthermore, specific features, structures, or characteristics in one or more embodiments can be combined in any suitable form.
[0036] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.
[0037] The following description, in conjunction with the accompanying drawings, details a specific scheme for a groundwater pollution inversion and source tracing method provided by the present invention.
[0038] Please see Figure 1 The diagram illustrates a flowchart of a groundwater pollution inversion and source tracing method according to an embodiment of the present invention, specifically including:
[0039] Step S1: Obtain the concentration data of each pollutant at each sampling point; select pollutants as target pollutants one by one.
[0040] In one embodiment of the present invention, existing civil wells, industrial wells or specially drilled monitoring wells are used as sampling points, and the pollutants monitored include heavy metals, nitrates, organic pollutants, etc.
[0041] The equipment used includes water pumps, Bayler tubing, multi-parameter water quality analyzers, sample bottles, preservatives, and ice packs. All equipment that comes into contact with the water samples must undergo rigorous cleaning to prevent cross-contamination. The cleaning procedure typically includes acid washing and deionized water rinsing. Considering that stagnant water in monitoring wells cannot represent the true condition of the aquifer, before setting up sampling points, operators must use water pumps or Bayler tubing to extract the "dead water" from the well pipes until the water quality parameters of the discharged water stabilize (three consecutive measurements stabilize within ±10%).
[0042] Water quality sensors are installed at each sampling point, and a sampling frequency of 60 minutes per sampling is set synchronously to obtain the concentration data of various pollutants. The data is then transmitted to edge computing nodes or data analysis centers for analysis. Source tracing analysis is performed every 60 minutes, and pollutants are selected as target pollutants for analysis.
[0043] It should be noted that the usage methods of the various instruments are existing technologies. Since the units and dimensions of the concentrations of various pollutants are different, the concentration data can be preprocessed by linear normalization in the data dimension corresponding to each pollutant. In other embodiments of the present invention, the implementer can set the required sampling points, pollutant types and collection frequencies, which will not be elaborated further.
[0044] Step S2: Within the current preset sliding time window, based on the distribution of the target pollutant concentration data at each sampling point, obtain the accumulation coefficient of the target pollutant at each sampling point; based on the fluctuation differences in the target pollutant concentration data between different sampling points, combined with the differences in the accumulation coefficient, obtain the current diffusion influence of the target pollutant; select the sampling point with the highest overall concentration as the initial sampling point, and based on the distribution of other sampling points relative to the initial sampling point, combined with the diffusion influence and accumulation coefficient, select reference points.
[0045] In one embodiment of the present invention, a sliding time window is used to slide along the time axis to capture concentration data for analysis. Considering that groundwater pollution is affected by many factors, such as chemical plant tank leaks, landfill leachate, and sewage pipe ruptures, the degree of groundwater pollution varies at different stages of economic development. Furthermore, the long-term discharge of a certain pollutant can lead to the accumulation of toxicity in groundwater. Therefore, within the current preset sliding time window, the accumulation coefficient of the target pollutant at each sampling point is obtained based on the distribution of the concentration data of the target pollutant at each sampling point. This coefficient is used to characterize the current accumulation and anomaly of the pollutant.
[0046] Preferably, in one embodiment of the present invention, the sliding time window is 7 days in length. Taking the current time point as the end of the sliding time window, within the current preset sliding time window, for each sampling point, the difference between each concentration data and the average concentration is obtained to form a difference sequence. The difference can be positive or negative. A positive difference indicates that the concentration of pollutants is relatively large, and a negative difference indicates that the concentration of pollutants is relatively small.
[0047] To analyze the differences in water pollutants over different time periods, the differences are classified according to their distribution within the difference sequence, and the time interval from each type of difference to the present is obtained.
[0048] As an example, a hierarchical clustering algorithm is used to cluster the difference sequence. The preset number of clustering layers is 2. Each cluster corresponds to one type of difference. The average value of the concentration data corresponding to all differences within the cluster at the time of collection (start) is used to represent the representative time of each type of difference. The time interval between the representative time and the current source tracing analysis (start) time is used as the time interval from each type of difference to the current time, in hours. The minimum time interval is limited to 0.01.
[0049] Considering that in difference classification where the mean difference is greater than zero, the more data points in a class of differences and the larger the mean, the higher the concentration of pollutants in the corresponding time period and the higher the degree of pollution anomaly; and considering that the direction of the principal component represents the trend of pollutant changes, since the input data are all abnormal data within clusters of high concentrations, the larger the variance contribution rate of the first principal component, the more significant or closer a certain anomaly is to the present, and the larger the accumulation coefficient; at the same time, the closer the data is to the current time and the greater the degree of pollutant anomaly, the more continuous the pollution of the target pollutant is and the more serious the pollution situation is, thus the more serious the accumulation situation is.
[0050] Based on this, in the difference classification (clustering) where the mean difference is greater than zero, the mean and number of differences in each class are fused to obtain the degree of anomaly; principal component analysis is performed on the two-dimensional dataset consisting of the degree of anomaly and the inverse of the time interval to obtain the variance contribution rate of the first principal component; based on the time interval corresponding to each class of difference, the degree of anomaly is fused, and combined with the variance contribution rate of the first principal component, the clustering coefficient of the corresponding sampling points is obtained.
[0051] As an example, for any sampling point, in the difference classification where the mean difference is greater than zero, the mean and number of differences in each class are linearly normalized in their respective data dimensions, and the product of the normalization results is used as the degree of abnormality of the corresponding class of differences.
[0052] The sum of the products of the reciprocals of the time intervals corresponding to various differences and the degree of anomaly is used as the anomaly stacking coefficient. The product of the anomaly stacking coefficient and the variance contribution rate is used as the stacking coefficient of the corresponding sampling point, which comprehensively reflects the distribution of concentration data.
[0053] If the mean difference of all clusters is less than or equal to 0, the current clustering coefficient of the sampling point is directly set to 0. It should be noted that, in another embodiment of the invention, the differences of all clusters are analyzed by performing principal component analysis on a two-dimensional dataset composed of anomaly degree and the reciprocal of time intervals to obtain the orientation angle of the first principal component. (Radians), when the anomaly level is high and the most recent events correspond to severe accumulation, a weighting function is set. , , For the ideal trend direction, it is set to π / 4 radians, representing that the two dimensions are equally important. The product of the abnormal stacking coefficient obtained by classifying all differences, the variance contribution rate of the first principal component, and the function value of the weighting function is used as the stacking coefficient of the corresponding sampling point.
[0054] It should be noted that the analysis process is consistent for each sampling point, and only one example is described here; in other embodiments of the present invention, the implementer can adjust the length of the sliding window and use other clustering algorithms. Principal component analysis and hierarchical clustering are existing technologies and will not be described in detail here.
[0055] The groundwater network is intricate, with underground rivers connecting groundwater sources in different areas, which can lead to cross-contamination. Therefore, based on the fluctuation differences in the concentration data of target pollutants between different sampling points, the diffusion and migration characteristics of pollutants can be inferred. Combined with the differences in the accumulation coefficient, the diffusion characteristics of pollutants can be comprehensively analyzed to obtain the diffusion influence of the current target pollutant, which represents the diffusion intensity of pollutants between these sampling points.
[0056] Preferably, in one embodiment of the present invention, considering that the pollutant concentrations of pollution sources such as wastewater discharge at different sampling points are different, the accumulation of target pollutants at different locations varies greatly when the diffusion effect is small, and the fluctuation of concentration data varies greatly.
[0057] Therefore, by integrating the DTW distance of the target pollutant concentration data between different sampling points and the absolute average deviation of the packing coefficient, the diffusion impact can be obtained.
[0058] Among them, the diffusion of groundwater pollution has a time delay, and the DTW algorithm can compare the overall similarity of different sequences in the presence of time misalignment, thus eliminating the interference caused by the time offset between sampling points. The larger the DTW distance, the greater the difference in the concentration evolution process between the two sampling points and the smaller the diffusion impact. The larger the absolute average deviation of the accumulation coefficient, the greater the difference in the degree of pollutant accumulation between sampling points, the more uneven the spatial distribution of pollution, and the lower the degree of diffusion.
[0059] Therefore, the product of the mean DTW distance and the absolute average deviation between different sampling points is used as the independent variable. The negative correlation mapping is performed using the negative correlation function exp(-x) with the natural constant e as the base, and the logical relationship is adjusted. The mapped value is used as the diffusion influence of the current target pollutant. x is the independent variable, with the mean DTW distance representing the fluctuation difference of the concentration data and the absolute average deviation representing the difference of the packing coefficient.
[0060] It should be noted that before using the negative correlation function exp(-x) for negative correlation mapping, the mean and absolute mean deviation of the DTW distance between different sampling points are linearly normalized in their respective data dimensions to eliminate data dimensions. Among them, the data of a certain type of data (such as absolute mean deviation) are collected in each source tracing analysis to form its own data dimension. The DTW algorithm and absolute mean deviation are existing technologies. In other embodiments of this invention, the implementer can also use variance to replace absolute mean deviation to obtain the diffusion influence degree, which will not be elaborated further.
[0061] When tracing the source, it is necessary to select appropriate sampling points for data analysis. The sampling point with the highest overall concentration often represents the location most likely to be close to the pollution source. Therefore, the sampling point with the highest overall concentration is selected as the initial sampling point. Based on the distribution of other sampling points relative to the initial sampling point, the spatial influence range of pollutants can be determined. By combining diffusion influence and accumulation coefficient, reference points with high correlation to the initial sampling point during the diffusion process and significant pollution accumulation characteristics can be effectively screened, providing a reliable basis for locating the pollution source.
[0062] Preferably, in one embodiment of the present invention, the average concentration data of each sampling point within the current sliding window is used as the overall concentration value to reflect the overall concentration characteristics, and the initial sampling point is selected.
[0063] Considering that the greater the diffusion influence, the more sampling points the initial sampling point affects, we obtain the farthest distance between the initial sampling point and other sampling points to determine the maximum selectable range. Then, we fuse the farthest distance and the diffusion influence to obtain the reference distance, and use the diffusion influence to make corrections to obtain an adaptive reference range.
[0064] By comparing the stacking coefficient and the preset stacking threshold within the reference distance of the initial sampling point, reference points with significant stacking characteristics under diffusion can be screened out, avoiding interference from outliers or low-correlation points.
[0065] As an example, the product of the diffusion effect and the farthest distance is used as the reference distance. The preset stacking threshold is 0.6. The stacking coefficient is linearly normalized in the corresponding data dimension. Sampling points with a normalized stacking coefficient greater than the preset stacking threshold and a distance from the initial sampling point less than the reference distance are marked as reference points. The distance between the initial sampling point and itself is 0. When the stacking coefficient meets the condition, it will also be marked as a reference point.
[0066] It should be noted that, in another embodiment of the present invention, considering that a small maximum overall concentration value indicates that the pollution is not serious and source tracing is not required, a corresponding source tracing concentration threshold can be set for each pollutant. When the maximum overall concentration value is less than the corresponding source tracing concentration threshold, source tracing is not performed. In other embodiments of the present invention, sampling points with concentrations greater than the source tracing concentration threshold can be marked as initial sampling points, and each initial sampling point can be analyzed separately. The source tracing concentration threshold can be adjusted by the implementer according to actual needs, and is related to the type of pollutant and the actual water quality standard requirements, and is not limited here.
[0067] Step S3: Based on the time delay performance of the concentration data of the reference point and surrounding sampling points within the current preset time domain interval, mark the current pollution source.
[0068] Considering that when a pollution source affects surrounding sampling points, there is a time delay between the concentration data of the affected sampling points and the source, the current pollution source is marked based on the time delay between the concentration data of the reference point and the surrounding sampling points within the current preset time interval.
[0069] Preferably, in one embodiment of the present invention, a Voronoi diagram is constructed based on the spatial distribution of the sampling points, and other sampling points in the diagram that are directly adjacent to the reference point are regarded as the surrounding sampling points of the reference point.
[0070] It should be noted that the method of constructing Voronoi diagrams is a well-known technology. In other embodiments of the present invention, the implementer may also set a distance threshold, such as 1 kilometer, and select sampling points with a distance of less than 1 kilometer from the reference point, and mark them as the surrounding sampling points of the corresponding reference point; when there are no surrounding sampling points for the reference point, they are skipped directly.
[0071] Preferably, in one embodiment of the present invention, please refer to Figure 2 The diagram illustrates a flowchart of a method for obtaining a pollution source according to an embodiment of the present invention, specifically including:
[0072] Step S301: Select reference points one by one as the points to be analyzed, and select the surrounding sampling points of the points to be analyzed one by one as control points; obtain the lag time with a preset step size within the preset time domain interval.
[0073] First, select the current point to be analyzed and the corresponding control point to facilitate comparative analysis. Considering that the time for pollutants to spread from the source to the surrounding points is affected by various factors such as the interlacing distribution of underground rivers and terrain, the spread time is uncertain. Therefore, the lag time is obtained with a preset step size within the preset time domain interval.
[0074] As an example, the preset time interval length is 10 hours, the preset step size is 1 hour, and the lag time ranges from 0 to 10 hours.
[0075] In other embodiments of the present invention, the implementer may adjust the preset time domain interval and preset step size. The analysis process is the same for each point to be analyzed and the control point. Only one example is described here, and will not be repeated.
[0076] Step S302: Analyze the correlation between the concentration data of the control point and the concentration data of the point to be analyzed under different lag times, obtain the optimal lag time and record the maximum correlation value.
[0077] Considering that the correlation between the concentration data of the control point and the concentration data of the point to be analyzed reflects different lag times, the optimal lag time is obtained and the maximum correlation value is recorded.
[0078] Considering that the cross-correlation function can quantify the similarity between two time series at different lag times, the correlation values between the control point and the point to be analyzed at different lag times are obtained based on the cross-correlation function, reflecting the time delay performance of the concentration data.
[0079] When the concentration sequence of the control point and the concentration sequence of the analysis point show the maximum correlation at a certain lag time, it indicates that the lag time may reflect the actual transmission delay of the pollutant from the analysis point to the control point. Therefore, the lag time corresponding to the maximum correlation value is selected as the optimal lag time.
[0080] It should be noted that when calculating the correlation value, the lag time is limited to the time series range of the control point concentration data, corresponding to the relationship between the point to be analyzed and the control point. The time range of the time series corresponds to the time range of the current sliding window. The cross-correlation function is an existing technology and will not be elaborated further.
[0081] Step S303: When the maximum correlation value is greater than the preset correlation threshold, the optimal lag time and the maximum correlation value are combined to obtain the directional score of the control point.
[0082] Considering that a small maximum correlation value does not necessarily indicate a large impact, but may simply mean that the concentrations are similar, a preset correlation threshold is set. When the maximum correlation value exceeds the preset correlation threshold, it indicates that there is a significant propagation relationship between the point to be analyzed and the control point. Considering that a larger optimal lag time indicates that the point to be analyzed leads the control point more, the lag relationship is more obvious, the higher the maximum correlation value, and the higher the reliability of the lag relationship, the two are combined to obtain the directional score.
[0083] As an example, the product of the optimal lag time and the maximum correlation value is used as the directional score for the control point.
[0084] Specifically, when the maximum correlation value is less than or equal to the preset correlation threshold, the directional score is 0.
[0085] In other embodiments of the present invention, the implementer may first perform linear normalization on the optimal lag time and the maximum correlation value on the corresponding data dimensions, and then fuse them by multiplication or weighted summation to avoid the influence of the value range of a certain element and dominate the directional score.
[0086] Step S304: Combine all directional scores of the points to be analyzed to obtain the probability of pollution sources; mark the current pollution source based on the probability of pollution sources.
[0087] The higher the directional score of all control points of the point to be analyzed, the stronger the directional characteristics of the point to be analyzed to multiple surrounding points, which means that it is likely to be the source of pollution. Therefore, the probability of pollution source is obtained and marked.
[0088] As an example, the mean of the directional scores of the point to be analyzed and all control points is used as the probability of the pollution source.
[0089] The probability of a pollution source is linearly normalized within the corresponding data dimension. The preset pollution threshold is set to 0.5. When the probability of a pollution source is greater than the preset pollution threshold, the corresponding sampling point is marked as the pollution source.
[0090] In summary, to address the technical limitations of existing static analysis methods in tracing and early warning of groundwater pollution sources, this invention provides a groundwater pollution inversion and source tracing method. This invention first acquires the concentration data of various pollutants at each sampling point; then, it analyzes the concentration data distribution of the target pollutant within a sliding time window to obtain the accumulation coefficient; further, it obtains the diffusion impact degree based on the fluctuation differences in the target pollutant concentration data and the differences in the accumulation coefficient; further, it selects reference points based on the distribution of other sampling points relative to the initial sampling point, combined with the diffusion impact degree and the accumulation coefficient; finally, it marks the current pollution source based on the time delay performance of the concentration data between the reference point and surrounding sampling points. This scheme acquires pollutant concentration data through a sliding time window, calculates the accumulation coefficient, and combines the fluctuation differences between sampling points to obtain the diffusion impact degree. After selecting reference points, it uses time delay correlation analysis to pinpoint the pollution source, effectively solving the problem of accurately and promptly identifying groundwater pollution diffusion sources.
[0091] It should be noted that the order of the above embodiments of the present invention is merely for descriptive purposes and does not represent the superiority or inferiority of the embodiments. The processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0092] The various embodiments in this specification are described in a progressive manner. The same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on describing the differences from other embodiments.
Claims
1. A method for tracing and identifying the source of groundwater pollution, characterized in that, The method includes: Acquire the concentration data of each pollutant at each sampling point; select pollutants by category as target pollutants; Within the current preset sliding window, the accumulation coefficient of the target pollutant at each sampling point is obtained based on the distribution of the concentration data of the target pollutant at each sampling point; the diffusion influence of the target pollutant is obtained based on the fluctuation differences of the concentration data of the target pollutant between different sampling points and the differences in the accumulation coefficient; the average concentration data of each sampling point within the current sliding window is used as the overall concentration value, the sampling point with the largest overall concentration is selected as the initial sampling point, and reference points are selected based on the distribution of other sampling points relative to the initial sampling point, combined with the diffusion influence and the accumulation coefficient; Based on the time delay performance of the concentration data of the reference point and surrounding sampling points within the current preset time interval, the current pollution source is marked; The method for obtaining the packing coefficient includes: For each sampling point, the difference between each concentration data point and the mean concentration is obtained to form a difference sequence; the difference is classified according to the distribution of the difference in the difference sequence, and the time interval from each type of difference to the current time is obtained; In the difference classification where the mean difference is greater than zero, the mean and number of differences in each class are fused to obtain the degree of anomaly; principal component analysis is performed on the two-dimensional dataset consisting of the degree of anomaly and the inverse of the time interval to obtain the variance contribution rate of the first principal component; based on the time interval corresponding to each class of difference, the degree of anomaly is fused, and combined with the variance contribution rate of the first principal component, the packing coefficient of the corresponding sampling point is obtained. The methods for obtaining the pollution source include: Each reference point is selected as the point to be analyzed, and each surrounding sampling point of the point to be analyzed is selected as a control point; the lag time is obtained within the preset time domain interval with a preset step size; Analyze the correlation between the concentration data of the control point and the concentration data of the point to be analyzed under different lag times, obtain the optimal lag time and record the maximum correlation value; When the maximum correlation value is greater than the preset correlation threshold, the optimal lag time and the maximum correlation value are combined to obtain the directional score of the control point; The probability of a pollution source is obtained by combining all the directional scores of the points to be analyzed; the current pollution source is marked based on the probability of the pollution source.
2. The groundwater pollution inversion and source tracing method according to claim 1, characterized in that, Methods for classifying differences based on the distribution of differences within the difference sequence include: The difference sequence is clustered using a hierarchical clustering algorithm, with a preset number of clustering layers of 2.
3. The groundwater pollution inversion and source tracing method according to claim 1, characterized in that, The method for obtaining the diffusion influence includes: The diffusion impact is obtained by fusing the DTW distance of the concentration data of the target pollutant between different sampling points and the absolute average deviation of the packing coefficient.
4. The groundwater pollution inversion and source tracing method according to claim 1, characterized in that, The method for obtaining the reference point includes: Obtain the farthest distance between the initial sampling point and other sampling points, fuse the farthest distance and the diffusion influence to obtain a reference distance, and within the reference distance of the initial sampling point, compare the stacking coefficient and the preset stacking threshold to filter out reference points.
5. The groundwater pollution inversion and source tracing method according to claim 1, characterized in that, The method for obtaining the optimal lag time includes: Based on the cross-correlation function, the correlation values between the control point and the point to be analyzed are obtained at different lag times, and the lag time corresponding to the maximum correlation value is selected as the optimal lag time.
6. The groundwater pollution inversion and source tracing method according to claim 1, characterized in that, The methods for obtaining the probability of the pollution source include: The mean of the directional scores of the point to be analyzed and all the control points is used as the probability of the pollution source.
7. The groundwater pollution inversion and source tracing method according to claim 6, characterized in that, The methods for marking the pollution sources include: When the probability of a pollution source is greater than a preset pollution threshold, the corresponding sampling point is marked as a pollution source.
8. The groundwater pollution inversion and source tracing method according to claim 1, characterized in that, The method for obtaining the surrounding sampling points includes: A Voronoi diagram is constructed based on the spatial distribution of the sampling points, and other sampling points directly adjacent to the reference point in the diagram are regarded as the surrounding sampling points of the reference point.
Citation Information
Patent Citations
Pollution traceability system and method based on water quality and water quantity monitoring and analysis of drainage system
CN119959494A
Knowledge graph-based ozone precursor collaborative traceability method and system
CN120932773A