Groundwater pollution degree grading monitoring method based on pH value dynamic perception

By using nonlinear processing and machine learning models of groundwater pH time-series signals, the problem of the inability to identify pollution events in the early stages of existing technologies has been solved, enabling dynamic monitoring and early warning of groundwater pollution levels, and improving the sensitivity and adaptability of monitoring.

CN121659092APending Publication Date: 2026-03-13GANSU GEOLOGICAL ENG SURVEY INST
View PDF 6 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-02-05
Publication Date
2026-03-13

AI Technical Summary

Technical Problem

Existing groundwater monitoring technologies cannot effectively capture sudden or intermittent pollution events, and threshold alarms that rely on hysteresis cannot meet the needs of early warning. Linear analysis methods are difficult to characterize nonlinear dynamic features.

Method used

By acquiring pH time-series signals from groundwater monitoring points, ensemble empirical mode decomposition and wavelet threshold denoising are performed to reconstruct the phase space. Nonlinear dynamic characteristic parameters such as correlation dimension, maximum Lyapunov exponent, and Hurst exponent are extracted and input into an ensemble learning model for pollution level classification.

Benefits of technology

It enables early identification and dynamic monitoring of groundwater pollution levels, improving the foresight and accuracy of monitoring. It can issue early warnings before pollutant concentrations reach thresholds and adapt to dynamic patterns under different hydrogeological backgrounds.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121659092A_ABST
    Figure CN121659092A_ABST
Patent Text Reader

Abstract

The invention discloses an underground water pollution degree grading monitoring method based on pH value dynamic perception, which comprises the following steps: acquiring a high-frequency pH time sequence signal, and preprocessing through ensemble empirical mode decomposition and wavelet threshold denoising to obtain a clean signal; determining an optimal parameter by using an average mutual information method and a pseudo neighbor point method, performing phase-space reconstruction on the clean signal, and recovering a system attractor trajectory; extracting nonlinear dynamic features such as correlation dimensions and maximum Lyapunov exponent based on the attractor trajectory to form a high-dimensional feature vector; and inputting the feature vector into a pre-trained ensemble learning classification model, outputting a pollution level and probability distribution, and realizing dynamic monitoring by sliding a time window. According to the invention, by mining the nonlinear dynamic information of the pH signal, high-sensitivity, dynamic and prospective graded monitoring of the groundwater pollution degree can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of environmental monitoring technology, and in particular to a method for classifying and monitoring the degree of groundwater pollution based on dynamic pH value sensing. Background Technology

[0002] Groundwater quality monitoring is a core component of environmental protection and water resource management. Its technological development has progressed from offline to online, from single-parameter to multi-parameter, and from static characterization to dynamic analysis. Traditional groundwater monitoring relies primarily on periodic on-site sampling and chemical analysis in laboratories. While this method can accurately determine the concentration of various pollutants, it suffers from low monitoring frequency, significant time lag, and high operating costs, making it difficult to detect sudden or intermittent pollution events. To overcome these shortcomings, continuous monitoring technology based on online sensors has emerged. This technology enables high-frequency, real-time data acquisition of key water quality parameters such as pH, conductivity, and dissolved oxygen. Currently, the application of this type of online monitoring data mainly focuses on setting static thresholds for alarms; that is, triggering an alarm when the measured value exceeds a preset normal range. This method improves the timeliness of monitoring to some extent.

[0003] However, these existing technologies still have some limitations. First, the threshold judgment method is a superficial utilization of time-series signals with rich information dimensions, thus ignoring the system state information contained in the dynamic fluctuations of the signal over time. Groundwater systems are typical nonlinear systems, so their response to pollutant input is not only reflected in the shift of parameter mean values, but more importantly in changes in their dynamic behavior patterns. Second, threshold alarms, as a delayed response mechanism, are usually triggered only after pollutant concentrations have reached high levels and have a substantial impact on the aquatic environment, failing to meet the need for early warning of pollution events. Third, while some studies have attempted to introduce time series analysis methods, they are mostly limited to linear methods such as Fourier analysis and autocorrelation analysis. These methods are often based on the assumption of linearity and stationarity, making it difficult to effectively characterize and quantify the non-stationary and nonlinear dynamic characteristics of groundwater systems after pollution disturbances, thus failing to fundamentally improve the sensitivity and foresight of monitoring. Summary of the Invention

[0004] The purpose of this section is to outline some aspects of embodiments of the present invention and to briefly describe some preferred embodiments. Simplifications or omissions may be made in this section, as well as in the abstract and title of this application, to avoid obscuring the purpose of these documents; however, such simplifications or omissions should not be construed as limiting the scope of the invention.

[0005] In view of the aforementioned existing problems, this invention is proposed. Therefore, this invention provides a method for classifying and monitoring groundwater pollution levels based on dynamic pH sensing, to address the problems mentioned in the background art.

[0006] To solve the above-mentioned technical problems, the present invention provides the following technical solution: a method for classifying and monitoring groundwater pollution levels based on dynamic pH sensing, comprising:

[0007] The pH time-series signal from the groundwater monitoring point is acquired, and the pH time-series signal is preprocessed to generate a clean signal for nonlinear dynamic analysis.

[0008] The clean signal is reconstructed in phase space to map the one-dimensional time series into attractor trajectories in a high-dimensional phase space;

[0009] Based on the attractor trajectory, at least one nonlinear dynamic feature parameter is extracted to form a feature vector.

[0010] The feature vector is input into a pre-trained pollution level classification model, which outputs the groundwater pollution level corresponding to the pH time series signal.

[0011] As a preferred embodiment of the groundwater pollution level classification monitoring method based on dynamic pH sensing described in this invention, the pH time-series signal is preprocessed, including:

[0012] The pH time series signal is decomposed into multiple intrinsic mode functions and a residual term by ensemble empirical mode decomposition;

[0013] The first intrinsic mode function component obtained from the decomposition is selected as the high-frequency noise component, and wavelet threshold denoising is performed on it.

[0014] The denoised intrinsic mode functions are reconstructed together with the unprocessed remaining intrinsic mode functions and residual terms to obtain the clean signal.

[0015] As a preferred embodiment of the groundwater pollution level classification monitoring method based on dynamic pH sensing described in this invention, the method includes: performing phase space reconstruction on the cleanliness signal, comprising:

[0016] The average mutual information of the clean signal is calculated as a function of the delay time, and the time corresponding to the first local minimum point on the curve representing the change is determined as the optimal delay time.

[0017] The optimal embedding dimension is determined by calculating the average ratio of the change in distance between all points in the phase space and their nearest neighbors as the embedding dimension increases from m to m+1, and then determining the dimension at which the average ratio tends to saturate and stabilize.

[0018] As a preferred embodiment of the groundwater pollution level classification monitoring method based on dynamic pH sensing described in this invention, after the phase space reconstruction is performed, multiple state vectors are obtained, wherein each state vector consists of m data points from the clean signal, and the time interval between the data points is the optimal delay time.

[0019] As a preferred embodiment of the groundwater pollution level classification monitoring method based on dynamic pH sensing described in this invention, the nonlinear dynamic characteristic parameters included in the feature vector are selected from a combination of at least three of the following:

[0020] The correlation dimension is used to characterize the geometric complexity of the attractor trajectory;

[0021] The maximum Lyapunov exponent is used to characterize the local instability and degree of chaos of the attractor trajectory;

[0022] The Hurst index is used to quantify the long-range correlation of the clean signal through detrended volatility analysis.

[0023] Approximate entropy is used to measure the regularity and complexity of the clean signal.

[0024] As a preferred embodiment of the groundwater pollution level classification monitoring method based on dynamic pH sensing described in this invention, the method for determining the correlation dimension includes:

[0025] In the reconstructed phase space, calculate the probability that the Euclidean distance between any two points on the attractor trajectory is less than a given radius, i.e., the correlation integral;

[0026] In the log-log coordinate graph of the correlation integral and radius, a linear scaling region with stable local slope is identified, and the slope of the linear scaling region is determined as the correlation dimension.

[0027] As a preferred embodiment of the groundwater pollution level classification monitoring method based on dynamic pH sensing described in this invention, the method for determining the maximum Lyapunov index includes:

[0028] In the reconstructed phase space, the average logarithmic separation rate of the initial neighboring orbits on the attractor trajectory over time is tracked, and the slope of the separation rate in the linearly growing region is determined as the maximum Lyapunov exponent.

[0029] As a preferred embodiment of the groundwater pollution level classification monitoring method based on dynamic pH sensing described in this invention, the pre-trained pollution level classification model is an ensemble learning model, which is trained using a nonlinear dynamic feature vector training set containing known pollution level labels.

[0030] As a preferred embodiment of the groundwater pollution level classification monitoring method based on dynamic pH sensing described in this invention, the output further includes a probability distribution of the pH time-series signal belonging to each preset pollution level.

[0031] As a preferred embodiment of the groundwater pollution level classification monitoring method based on dynamic pH value sensing described in this invention, the method is implemented through a sliding time window, which moves on the real-time acquired pH time-series signal with a preset step size to achieve dynamic classification of groundwater pollution level.

[0032] Compared with existing technologies, the beneficial effects of this solution are:

[0033] 1. This invention, by mining the nonlinear dynamic characteristics of pH time-series signals rather than relying on the absolute magnitude or mean shift of pH values, can capture subtle changes in the dynamic behavior patterns of groundwater systems caused by disturbances in the early stages of pollution. Through an identification method based on the "fingerprint" of system state, this invention can issue early warnings before pollutant concentrations reach traditional alarm thresholds, improving the early identification capability and forward-looking nature of pollution events.

[0034] 2. The preprocessing method combining EEMD and wavelet thresholding employed in this invention can adaptively process non-stationary and nonlinear pH signals, effectively separating noise from real dynamic information and ensuring the accuracy of feature extraction. Furthermore, the machine learning-based modeling approach avoids the difficulty of setting empirical thresholds for different monitoring points. By learning from historical data, the model can adapt to dynamic patterns under specific hydrogeological backgrounds, exhibiting stronger universality and environmental adaptability.

[0035] 3. Furthermore, this invention transforms static, single-point judgment into continuous tracking of the pollution state evolution process by introducing a sliding time window mechanism and outputting the probability distribution of pollution levels. This allows environmental managers not only to see the most likely pollution level at present, but also to quantify the uncertainty of the judgment and observe the dynamic trend of pollution level changes, providing them with richer decision-making basis. Attached Figure Description

[0036] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. Wherein:

[0037] Figure 1This is a flowchart illustrating the overall process of a groundwater pollution level classification monitoring method based on dynamic pH sensing, as described in one embodiment of the present invention. Detailed Implementation

[0038] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the protection scope of the present invention.

[0039] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.

[0040] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a single or selective embodiment that is mutually exclusive with other embodiments.

[0041] This invention is described in detail with reference to the schematic diagrams. When detailing the embodiments of this invention, for ease of explanation, the cross-sectional views illustrating the device structure may be partially enlarged, not adhering to the usual scale. Furthermore, the schematic diagrams are merely examples and should not be construed as limiting the scope of protection of this invention. In actual fabrication, the three-dimensional spatial dimensions of length, width, and depth should be included.

[0042] Furthermore, in the description of this invention, it should be noted that the terms "upper," "lower," "inner," and "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. These terms are used solely for the convenience of describing the invention and for simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on the invention. In addition, the terms "first," "second," or "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.

[0043] Unless otherwise explicitly specified and limited, the terms "installation," "connection," and "joining" in this invention should be interpreted broadly. For example, they can refer to fixed connections, detachable connections, or integral connections; similarly, they can refer to mechanical connections, electrical connections, or direct connections, or indirect connections through an intermediate medium, or internal connections between two components. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances.

[0044] Example 1

[0045] Reference Figure 1 This is the first embodiment of the present invention, which provides a method for classifying and monitoring the degree of groundwater pollution based on dynamic pH sensing, including:

[0046] S1. Acquire the pH time series signal from the groundwater monitoring point, and preprocess the pH time series signal to generate a clean signal for nonlinear dynamic analysis.

[0047] It should be noted that, in order to obtain high-fidelity original signals that can truly reflect the dynamic behavior of the groundwater system, the present invention is based on signal processing methods to remove noise and interference from irrelevant trends while preserving the inherent dynamic information of the groundwater system to the maximum extent.

[0048] Furthermore, to ensure the capture of dynamic pH responses caused by transient changes in hydrogeochemical processes (such as pulsed pollution inputs and hydraulic exchanges), a high-frequency sampling strategy is required. Specifically, in this embodiment, an online pH sensor with high resolution (e.g., 0.001 pH units) and fast response time (e.g., less than 10 seconds) is deployed in the groundwater monitoring well, and the data acquisition frequency is set to once every 30 seconds. It should be explained that this sampling frequency needs to be set higher than that of conventional water quality monitoring because the subsequent nonlinear analysis methods mainly use phase space reconstruction, the effectiveness of which is highly dependent on the continuity and density of data points in time. Therefore, only by setting a sampling frequency higher than the conventional sampling frequency can we ensure that the reconstructed attractor trajectory is smooth and undistorted, thereby accurately revealing the true dynamic structure of the groundwater system.

[0049] Furthermore, due to the original pH time-series signal The signal will inevitably be affected by random electrical noise from sensors, environmental electromagnetic interference, and slow seasonal variations unrelated to pollution events. If these interfering components are not removed, they will severely impact the accuracy of subsequent calculations of nonlinear parameters, which are highly sensitive to noise. Therefore, this invention employs a composite preprocessing procedure, which includes signal decomposition based on ensemble empirical mode decomposition, wavelet threshold denoising for high-frequency intrinsic mode function components, and signal reconstruction.

[0050] Furthermore, for signal decomposition based on ensemble empirical mode decomposition, traditional Fourier transform-based filtering methods, which presuppose fixed basis functions, are unsuitable for processing non-stationary and nonlinear signals such as pH time series signals. Therefore, the present invention employs the ensemble empirical mode decomposition (EEMD) method, which adaptively decomposes these non-stationary and nonlinear signals into multiple intrinsic mode functions (IMFs) arranged from high to low frequency and a residual term based on the signal's own time-scale characteristics.

[0051] Specifically, ensemble empirical mode decomposition (EMD) effectively overcomes the mode aliasing problem in the standard EMD method by repeatedly adding Gaussian white noise to the original signal and performing EMD, followed by ensemble averaging of the multiple EMD decomposition results. This results in each decomposed eigenmode function component having a more explicit physical meaning. Its mathematical expression is:

[0052]

[0053] in, This is the original pH time-series signal. For the first Each intrinsic mode function component Highest frequency, following Increasing frequency leads to decreasing frequency. High-frequency IMF (such as...) ) typically contains most of the random noise, and the intermediate frequency IMF (such as ) represents the main dynamic response of the system, low-frequency IMF (such as This reflects the slow-changing behavior of the system. The residual term represents the overall trend of the signal, such as seasonal variations.

[0054] Furthermore, since directly discarding high-frequency IMF components using traditional methods is a simple but crude noise reduction approach, it may lose the true high-frequency fluctuations caused by sudden pollution events. Therefore, to handle noise more conservatively, the present invention adopts the following strategy: calculate the correlation coefficient or energy ratio of each intrinsic mode function component with the original signal, and select the first intrinsic mode function with the lowest energy ratio and highest frequency (…). ) was identified as the main noise carrier, and Perform wavelet thresholding denoising. This process includes:

[0055] Choosing Symlets wavelet systems (such as sym8) as basis functions, for Perform wavelet decomposition at 3-5 levels. The threshold is set using the VisuShrink principle. The calculation formula is: ,in, For signal length, The noise level is estimated using the formula. Estimate, These are the first-level detail coefficients. The wavelet coefficients are processed using a soft thresholding function: when the absolute value of the coefficient is less than... Set to zero when greater than Time-to-zero contraction. Finally, the processed wavelet coefficients are used for reconstruction.

[0056] It should be noted that this method can distinguish between signal and noise in the wavelet domain, effectively suppressing noise while preserving abrupt changes and singularities in the signal to the maximum extent.

[0057] Furthermore, for signal reconstruction, the high-frequency components that have undergone wavelet threshold denoising will be... Compared with unprocessed mid- and low-frequency intrinsic mode function components ( and the residuals representing the trend term. By adding them together, the clean signal can be obtained for nonlinear dynamic analysis. :

[0058]

[0059] It should be noted that the above processing method targets only the frequency band with the most severe noise pollution, while fully preserving the mid-frequency components containing dynamic information and the low-frequency components representing background trends. This achieves the dual goals of signal-to-noise separation and high-fidelity reconstruction. The resulting clean signal... It is a data sequence that can clearly and accurately reflect the real dynamic response of the groundwater system to internal and external disturbances (especially pollution).

[0060] S2. Reconstruct the phase space of the clean signal, mapping the one-dimensional time series to attractor trajectories in a high-dimensional phase space.

[0061] It should be noted that, since the groundwater system is essentially a dynamic system of multiphase fluid moving in a porous medium, according to Takens' embedding theorem, for a deterministic dynamic system, even if only one variable of the system can be observed (such as the pH value here), we can still reconstruct the topological structure of the original system's attractor in a sufficiently high-dimensional space by reconstructing the time series of that variable. The geometry and dynamic characteristics of this reconstructed attractor contain the intrinsic dynamic information of the entire system (in this case, the groundwater system). Therefore, the purpose of this step is to obtain the corresponding cleanliness signal. Through scientific methods, it is mapped to a set of points in a high-dimensional phase space (i.e., attractor trajectories), thus providing an analytical object for subsequent extraction of nonlinear dynamic characteristic parameters (such as fractal dimension and chaotic index) that can quantify the complexity of the system.

[0062] Furthermore, this reconstruction process is determined by two key parameters, one of which is the optimal delay time. The other is the optimal embedding dimension. It is important to emphasize that the appropriateness of the selection of this key parameter directly affects whether the reconstructed attractor can accurately and without distortion reflect the characteristics of the original dynamic system. In this embodiment, we adopt a systematic approach to determine these two parameters.

[0063] Furthermore, due to the delay time This determines the time interval between the coordinate components that make up the state vector. If the delay time is too small, adjacent coordinate components... and Highly correlated, the reconstructed trajectories will cluster near the diagonal of the phase space, making effective expansion impossible; if the delay is too large, then... and The dynamic correlation between variables may have been lost, easily leading to reconstructed trajectories riddled with noise and artifacts. Therefore, this invention employs Average Mutual Information (AMI) to determine the optimal delay time. It is important to emphasize that the advantage of choosing AMI is that, compared to the autocorrelation function method, which can only measure linear correlations, AMI can capture the nonlinear dependencies between variables, making it more suitable for analyzing nonlinear dynamic systems.

[0064] Specifically, calculate the cleanliness signal. Different delay times Average mutual information Here, average mutual information The physical meaning is that, in the known In this case, The average amount of new information that can be provided can be expressed as:

[0065]

[0066] in, The signal value is The probability of. The signal value is The probability, this value is delayed in time. . The signal takes the value at the current moment. And in The value after time is The joint probability. It is in the The clean pH value measured at each sampling time point, after noise reduction and reconstruction. and These are two different clean pH values, such as and .

[0067] Specifically, through drawing Follow For a changing curve, find the first local minimum point on the curve. The local minimum point corresponds to... This value represents the optimal delay time. The reason for this choice is because... The first local minimum means that at this time and The statistical correlation between the two decreased to a local minimum, but they still maintained a dynamic relationship. The value, as a delay time, ensures that the new dimensions in the reconstructed phase space can introduce the most independent information, thereby most effectively unfolding the attractor trajectory in the phase space.

[0068] Furthermore, we can define the embedding dimension. It is the dimension of the reconstructed phase space. Therefore, according to the embedding theorem, this... It must be large enough ( ,in Only when the attractor dimension is specified can the reconstructed attractor be differentially homeomorphic to the original attractor, i.e., completely expanded without intersection. In short, if the... If the value is too small, the trajectories of different parts will be incorrectly folded together due to projection, forming "pseudo-nearest neighbor points," which will seriously mislead subsequent dynamic analysis; if the value is too small... If the value is too large, it will amplify the impact of noise and introduce unnecessary computational burden. Therefore, the present invention employs the False Nearest Neighbors (FNN) method to determine the optimal embedding dimension. The core idea of ​​this method is that if a pair of points is in... In 1-dimensional space, they are nearest neighbors, but when the dimension increases to... Their distance will increase significantly at that time, then they will be in In 3D space, these are a pair of "pseudo-nearest neighbors". The algorithm flow is as follows:

[0069] (i) From Initially, for each state vector in the phase space... Find its nearest neighbor in Euclidean distance. and calculate them in Distance in 3D space .

[0070] (ii) Calculate these two vectors ( and )exist Distance in 3D space .

[0071] (iii) Determine whether the neighbor is a false neighbor. This is usually based on the following two criteria:

[0072] Criterion 1: Calculate the relative increment of the distance. If the following formula is satisfied, it is determined to be a pseudo-nearest neighbor:

[0073]

[0074] in, It is a preset threshold (in this embodiment, it is 10~15), which corresponds to the average rate of distance change.

[0075] Criterion 2: To avoid misjudgments due to small attractor size, this criterion is related to the overall size of the attractor. Specifically, if Larger than a proportion of the attractor size (e.g., the standard deviation of the attractor size) ( ), and was also judged as a false neighbor.

[0076] (iv) Calculate in the current dimension Below, the percentage of pseudo-nearest neighbors among all points.

[0077] (v) Embedding dimension Increase by 1, and repeat steps (i) to (iv) above until the percentage of pseudo-nearest neighbors first drops below zero or a sufficiently small threshold (1% in this embodiment). At this point... The value is the optimal embedding dimension.

[0078] Furthermore, after determining the optimal delay time... and optimal embedding dimension After that, the clean signal can be processed. Reconstruct the state vectors to create multiple state vectors.

[0079] Specifically, the first State vectors Depend on It consists of data points from cleanliness signals, and its form is as follows:

[0080]

[0081] in, , It is the total number of state vectors in the phase space. It is a clean signal relative to Delayed The values ​​of data points at each time step; It is a clean signal relative to Delayed The values ​​of data points at each time step; and so on.

[0082] It should be noted that these state vectors are in Connecting these points in time sequence within Euclidean space constitutes a geometric representation of the dynamic behavior of the groundwater system, namely, the attractor trajectory. The morphology, density distribution, and orbital evolution characteristics of this trajectory become the direct basis for extracting nonlinear dynamic characteristic parameters.

[0083] S3. Based on the attractor trajectory, extract at least one nonlinear dynamic feature parameter to form a feature vector.

[0084] It should be noted that the purpose of this step is to transform the reconstructed high-dimensional attractor trajectory into a set of low-dimensional quantitative indicators that can be understood and processed by machine learning models. These quantitative indicators, namely nonlinear dynamic characteristic parameters, are parameters that characterize the dynamic behavior of the system from different dimensions (geometric structure, stability, long-range correlation, sequence complexity, etc.). By combining these parameters, we can construct a high-information-density feature vector, which serves as the "dynamic fingerprint" of the groundwater system within a specific time window. Then, when the state of the groundwater system (i.e., the degree of pollution) changes, this dynamic fingerprint will undergo identifiable changes. In this embodiment, the high-information-density feature vector can be composed of four nonlinear dynamic characteristic parameters: correlation dimension, maximum Lyapunov exponent, Hurst exponent, and approximate entropy. The input to this feature vector is the generated attractor trajectory or the generated clean signal, and the output is a four-dimensional feature vector.

[0085] Furthermore, the correlation dimension Correlation dimension is a type of fractal dimension used to characterize the geometric complexity and space-filling capacity of attractor trajectories, reflecting the density of the system's state distribution in phase space. In the applied groundwater system, initially a stable, uncontaminated system, its pH dynamic evolution may follow a relatively simple and regular pattern, corresponding to a low-dimensional attractor. When a contamination event occurs, the input of foreign substances acts as an external forcing, increasing the complexity and randomness of the system's internal chemical reactions and physical mixing processes. This leads to a more complex and extended attractor structure, resulting in a significant increase in correlation dimension. Therefore, changes in correlation dimension can serve as a sensitive indicator of a system's migration from an ordered to a complex (or chaotic) state.

[0086] Specifically, for this correlation dimension, the present invention uses the classic Grassberger-Procaccia (GP) algorithm to calculate it, and the calculation process is as follows:

[0087] First, calculate the correlation integral. For a given radius The correlation integral is any two state vectors in the phase space. and The distance between them is less than The probability of is calculated using the following formula:

[0088]

[0089] in, and These are two state vectors in phase space. This represents the Euclidean distance. It is the Heaviside step function, which has a value of 1 when the independent variable is greater than or equal to zero, and 0 otherwise.

[0090] Secondly, for attractors with fractal characteristics, a certain... Within the range, and There exists a power-law relationship, that is... .

[0091] Finally, and double logarithmic coordinate graph ( and In the analysis, a linear region with a relatively constant local slope is identified; this linear region is called the linear scaling region. The slope of this linear scaling region is the correlation dimension to be determined. :

[0092]

[0093] Furthermore, the maximum Lyapunov index It is a core indicator for determining whether a system is in a chaotic state, and this indicator quantifies the local instability of attractor trajectories. It is a clear indicator of chaos, and its numerical magnitude characterizes the average rate at which initial small perturbations in phase space separate exponentially over time, i.e., the system's sensitivity to initial conditions. In groundwater environments, the dynamics of uncontaminated systems may tend to stabilize. The introduction of pollutants, especially intermittent emissions or disturbances in the early stages of diffusion, disrupts the existing water-rock-microbe balance, introducing new nonlinear interactions and pushing the system towards the edge of chaos or a chaotic state. Specifically, this manifests as... A change from a non-positive value to a positive value, or a significant increase in an existing positive value (showing an upward trend), indicates a decrease in the predictability of the system and a more complex and irregular dynamic behavior. Based on this, the present invention employs a small-data-volume algorithm (such as the Rosenstein method) to estimate the maximum Lyapunov exponent. The processing procedure is as follows:

[0094] In phase space, for each state vector Find its nearest neighbor that is not adjacent in time. .

[0095] Track the evolution of these two neighboring orbits over time and calculate their evolution in... Distance after each discrete time step .

[0096] Calculate the average log-separation rate over the distance evolution of all initial point pairs. In short, it involves plotting a... Regarding evolution time The curve, in which Indicates all Find the average. It is the sampling time interval.

[0097] In the initial part of the curve, there exists an approximately linear growth region, and the slope of this linear growth region is the maximum Lyapunov exponent. .

[0098] Furthermore, the Hurst exponent originates from rescaled range analysis and is primarily used to quantify the long-range correlation or "memory" of time series. This indicates that the sequence is a random walk (without memory). This indicates that the sequence has persistence (positive correlation), meaning that past incremental trends predict future incremental trends; This indicates a negative correlation (inverse persistence), meaning that past incremental trends predict future reverse trends. Based on this pattern, the natural evolution of groundwater systems (such as those influenced by seasonal recharge and discharge) typically exhibits strong persistence. Contamination events, especially pulsed or irregular pollution inputs, disrupt this inherent long-term memory, making pH fluctuations more random and causing the Hurst exponent to approach 0.5. Therefore, a decrease in the H value can be seen as a sign that the system's inherent regularity has been disrupted by external random disturbances.

[0099] Specifically, this invention uses Detrended Fluctuation Analysis (DFA) to calculate the Hurst exponent. The steps are as follows:

[0100] The cumulative deviation sequence is segmented, and a local trend is fitted using a polynomial within each segment. The root mean square deviation between the sequence and the local trend is calculated, and the Hurst exponent is finally determined by the slope of the double logarithmic relationship between the root mean square deviation and the segment length.

[0101] Furthermore, approximate entropy Approximate entropy is a nonlinear dynamic parameter used to measure the regularity and complexity of a time series. A higher approximate entropy value indicates a higher probability of new patterns emerging in the series, and thus higher complexity and randomness; conversely, a lower approximate entropy value indicates stronger regularity and higher predictability. Compared to correlation dimension and LLE index, approximate entropy has lower requirements for data length and better noise resistance. In groundwater system monitoring scenarios, the background pH sequence may exhibit periodic or quasi-periodic fluctuations, resulting in a low approximate entropy value. However, the introduction of pollutants introduces non-periodic and complex dynamic changes, causing the series to generate more unpredictable new patterns, thus leading to a significant increase in the approximate entropy value.

[0102] Specifically, this approximate entropy value mainly includes two core parameters: the dimension of the pattern vector and the similarity tolerance. Through comparison... peacekeeping The probability of a given pattern appearing in a vector is used to quantify the complexity of the sequence by calculating the difference between the logarithms of these two probabilities, thus yielding an approximate entropy value. The calculation of this approximate entropy includes the following steps:

[0103] (1) Set the dimension m of the pattern vector (e.g., m=2) and the similarity tolerance q (usually taken as 0.1 to 0.25 times the standard deviation of the original data);

[0104] (2) Reconstruct the clean signal into m-dimensional and m+1-dimensional phase space vectors;

[0105] (3) Calculate the number of vectors whose distance to each m-dimensional vector is less than q from all other m-dimensional vectors, and calculate the logarithmic average of these vectors relative to the total number of vectors, denoted as q. ;

[0106] (4) Repeat step (3) for the m+1 dimension vector to obtain The approximate entropy value is ultimately determined by... We obtain, where M is the number of data points.

[0107] It should be noted that, through the above calculations, we can obtain a four-dimensional feature vector for each analysis time window:

[0108]

[0109] It should be noted that this eigenvector is a comprehensive and quantitative description of the dynamic state of the groundwater system within each time period.

[0110] S4. Input the feature vector into a pre-trained pollution level classification model and output the groundwater pollution level corresponding to the pH time series signal.

[0111] It should be noted that the obtained four-dimensional feature vector exhibits a complex and highly nonlinear correlation between its components and the degree of groundwater pollution, a relationship that is difficult to describe using simple thresholds or linear models. Therefore, this invention employs a supervised learning method to construct a pollution level classification model. By learning from a large number of feature-label sample pairs, this implicit mapping pattern is uncovered, thereby achieving automatic and accurate classification of new data.

[0112] Furthermore, considering that the contribution of each component of the feature vector to the classification result may differ, and that there may be complex interactions between features, a single classifier (such as logistic regression or a single decision tree) may be at risk of overfitting or insufficient generalization ability. To improve the accuracy and robustness of classification, the pollution level classification model in this invention employs an ensemble learning model. This ensemble learning approach effectively reduces variance and overfitting by combining the prediction results of multiple base learners, thus achieving better performance than any single base learner. In addition, in this embodiment, the ensemble learning model preferably uses a classifier based on Gradient Boosting Decision Tree (GBDT). Furthermore, the model's construction and training are implemented using the Scikit-learn library in a Python environment. The model's input layer receives a 4-dimensional feature vector after Z-score standardization, and the output layer outputs a 4-dimensional probability distribution through the Softmax function.

[0113] It should be noted that, in addition to the GBDT classifier, this ensemble learning model can also employ Random Forest, AdaBoost, or XGBoost, or algorithmic models based on Bagging or Boosting strategies.

[0114] Specifically, the input layer of this ensemble learning model corresponds to four standardized nonlinear dynamic feature values: correlation dimension, maximum Lyapunov exponent, Hurst exponent, and approximate entropy. It's important to note that inputting these dynamic feature values ​​into the model establishes a direct mapping between the "fingerprint" information (i.e., feature vector) reflecting the intrinsic dynamic state of the groundwater system and the external environmental state (i.e., pollution level). By learning this mapping, the model can achieve intelligent reasoning from system behavior patterns to pollution levels. The output layer corresponds to four preset pollution levels, and the Softmax function maps the output to a probability distribution. Multi-class log loss is used as the loss function to measure the difference between the predicted probability distribution and the one-hot encoding of the true labels.

[0115] Specifically, this ensemble learning model uses decision trees as base learners, iteratively constructing multiple decision trees in a sequential manner. The learning objective of each new tree is to fit the residuals (i.e., the negative gradient of the difference between the predicted and true values) of the ensemble results of all preceding trees. In this way, the model greedily searches for the optimal solution in the function space, thereby continuously reducing the overall prediction error. Furthermore, this model is an additive model, and its prediction result is a weighted sum of the outputs of all weak learners (decision trees). For an input feature vector... For each pollution level The prediction score can be expressed as:

[0116]

[0117] in, It represents the total number of decision trees. Represents the input feature vector Corresponding to pollution level The total predicted score. Indicated as predicted pollution level And the training of the first Each decision tree is used to process the input feature vector. The output value.

[0118] Furthermore, in order for the model to have classification capabilities, it must be trained using a training set containing known pollution level labels. The steps for constructing the training dataset are as follows:

[0119] Collect long-term historical high-frequency pH time-series data for the target monitoring area.

[0120] Groundwater pollution levels are classified into multiple discrete grades. In this embodiment, based on the "Groundwater Quality Standard" (GB / T 14848-2017) or the pollutant control limits set by the local environmental protection department, the water quality test results corresponding to historical data are marked and classified as follows:

[0121] Grade 0 (Clean): Corresponds to Class I-II water, with all water quality indicators within the natural background value range;

[0122] Level 1 (Slightly Polluted): Corresponds to Class III water, where some indicators (such as ammonia nitrogen and nitrate) exceed the background value but do not exceed the Class III limit;

[0123] Level 2 (Moderately Polluted): Corresponds to Class IV water, with significant pollutant input and some indicators reaching Class IV standards;

[0124] Level 3 (Severely Polluted): Corresponds to Class V water, where water quality indicators are seriously exceeded.

[0125] Alternatively, based on specific monitoring targets (such as heavy metal leaks), specific concentration thresholds can be set as the basis for classification, for example, when the concentration of characteristic pollutants... At level 0, At level 1, where, and The threshold is set, and so on.

[0126] Furthermore, to address the timescale differences between discrete laboratory sampling and continuous online monitoring when constructing training samples, a "point-to-segment" label matching strategy is adopted: a time window length consistent with the sliding window size used in subsequent online monitoring is set, which is 24 hours in this embodiment. Taking the laboratory sampling time as the center, pH time-series data for 12 hours before and after that time, totaling 24 hours, are extracted as a sample segment. Simultaneously, based on the quasi-stationary assumption that the probability of sudden water quality changes in the groundwater system within a short period is low, for each time point with a clear laboratory test result, pH time-series data for 12 hours before and after it are extracted as samples. Its nonlinear dynamic feature vector is calculated, and the corresponding pollution level is used as a label, forming a "feature vector-label" sample pair.

[0127] Before feeding the data into the model for training, all feature vectors in the training set are standardized (e.g., using Z-score standardization) so that the mean of each feature dimension is 0 and the standard deviation is 1, in order to eliminate the influence of differences in the units and numerical ranges of different feature parameters and accelerate model convergence.

[0128] Furthermore, the constructed training dataset is used for model training. The model training steps are as follows:

[0129] Input the preprocessed training dataset.

[0130] Define the loss function that minimizes the model (in this case, the multi-class log loss).

[0131] During model training (with 500 as the initial number of iterations), optimization is performed in the function space using gradient descent. Decision trees are added iteratively to minimize the model's loss function until the model converges (i.e., the maximum number of iterations is reached).

[0132] Furthermore, to achieve optimal model performance, it is usually necessary to fine-tune key hyperparameters. In this embodiment, we randomly divide the constructed training dataset into training and validation sets in an 8:2 ratio and use 5-fold cross-validation to obtain the hyperparameter combination that achieves the optimal classification accuracy on the validation set. Specifically, the number of base learners (n_estimators) is 500, the learning rate is 0.05, the maximum tree depth (max_depth) is 5, and the subsample ratio is 0.8. The ensemble learning model trained using this set of parameters can be embedded and deployed in an online monitoring system for classifying real-time data.

[0133] Specifically, the system achieves dynamic and continuous monitoring through a sliding time window. A window size consistent with the window used when generating training samples is set, i.e., 24 hours, and a step size is set (e.g., 1 hour). Each time the window moves forward one step, the aforementioned complete process S1~S3 is executed on the data within the window, generating a new feature vector. .

[0134] Furthermore, the feature vectors generated in real time and standardized in the same way as during training are... The input is fed into a pre-trained model. This model not only outputs the most probable pollution level, but also a probability distribution vector, calculated by analyzing the prediction scores within the model. Implemented using the Softmax function:

[0135]

[0136] in, Y is the total number of pollution levels (4 in this example). Y is a random variable.

[0137] Specifically, the output of this model is a probability distribution vector. ,in These represent the probabilities that the current dynamic behavior of the pH signal belongs to pollution levels 0, 1, 2, and 3, respectively. For example, the output might be [0.1, 0.2, 0.6, 0.1].

[0138] Furthermore, the category with the highest probability is used as the final pollution level for the current time window. In the example above, the final model output is "Level 2: Moderate Pollution".

[0139] It should be noted that by outputting the probability distribution, the model not only provides the most probable judgment but also quantifies the uncertainty of that judgment. Managers can then set more complex early warning rules based on this. For example, a high-level alarm can be triggered when the sum of the probabilities of "moderate pollution" or "severe pollution" exceeds a certain threshold (e.g., 50%), thereby achieving proactive risk management. Simultaneously, the sliding window mechanism allows the entire methodology to track the dynamic evolution of groundwater pollution in real time.

[0140] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. The solutions in the embodiments of this application can be implemented using various computer languages, such as the object-oriented programming language Java and the interpreted scripting language JavaScript.

[0141] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0142] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0143] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0144] Although preferred embodiments of this application have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of this application.

[0145] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.

Claims

1. A method for classifying and monitoring the degree of groundwater pollution based on dynamic pH sensing, characterized in that, include: The pH time-series signal from the groundwater monitoring point is acquired, and the pH time-series signal is preprocessed to generate a clean signal for nonlinear dynamic analysis. The clean signal is reconstructed in phase space to map the one-dimensional time series into attractor trajectories in a high-dimensional phase space; Based on the attractor trajectory, at least one nonlinear dynamic feature parameter is extracted to form a feature vector. The feature vector is input into a pre-trained pollution level classification model, which outputs the groundwater pollution level corresponding to the pH time series signal.

2. The groundwater pollution level classification monitoring method based on dynamic pH sensing as described in claim 1, characterized in that, The pH time-series signal is preprocessed, including: The pH time series signal is decomposed into multiple intrinsic mode functions and a residual term by ensemble empirical mode decomposition; The first intrinsic mode function component obtained from the decomposition is selected as the high-frequency noise component, and wavelet threshold denoising is performed on it. The denoised intrinsic mode functions are reconstructed together with the unprocessed remaining intrinsic mode functions and residual terms to obtain the clean signal.

3. The groundwater pollution level classification monitoring method based on dynamic pH sensing as described in claim 1, characterized in that, Phase space reconstruction of the cleanliness signal includes: The average mutual information of the clean signal is calculated as a function of the delay time, and the time corresponding to the first local minimum point on the curve representing the change is determined as the optimal delay time. The optimal embedding dimension is determined by calculating the average ratio of the change in distance between all points in the phase space and their nearest neighbors as the embedding dimension increases from m to m+1, and then determining the dimension at which the average ratio tends to saturate and stabilize.

4. The groundwater pollution level classification monitoring method based on dynamic pH sensing as described in claim 1 or 3, characterized in that, After the phase space reconstruction is performed, multiple state vectors are obtained, each of which consists of m data points from the clean signal, and the time interval between the data points is the optimal delay time.

5. The groundwater pollution level classification monitoring method based on dynamic pH sensing as described in claim 1, characterized in that, The nonlinear dynamic characteristic parameters contained in the eigenvector are selected from a combination of at least three of the following: The correlation dimension is used to characterize the geometric complexity of the attractor trajectory; The maximum Lyapunov exponent is used to characterize the local instability and degree of chaos of the attractor trajectory; The Hurst index is used to quantify the long-range correlation of the clean signal through detrended volatility analysis. Approximate entropy is used to measure the regularity and complexity of the clean signal.

6. The groundwater pollution level classification monitoring method based on dynamic pH sensing as described in claim 5, characterized in that, The method for determining the correlation dimension includes: In the reconstructed phase space, calculate the probability that the Euclidean distance between any two points on the attractor trajectory is less than a given radius, i.e., the correlation integral; In the log-log coordinate graph of the correlated integral and radius, a linear scaling region with a stable local slope is identified, and the slope of the linear scaling region is determined as the correlated dimension.

7. The groundwater pollution level classification monitoring method based on dynamic pH sensing as described in claim 5, characterized in that, The method for determining the maximum Lyapunov exponent includes: In the reconstructed phase space, the average logarithmic separation rate of the initial neighboring orbits on the attractor trajectory over time is tracked, and the slope of the separation rate in the linearly growing region is determined as the maximum Lyapunov exponent.

8. The groundwater pollution level classification monitoring method based on dynamic pH sensing as described in claim 1, characterized in that, The pre-trained pollution level classification model is an ensemble learning model, which is trained using a training set of nonlinear dynamic feature vectors containing known pollution level labels.

9. The groundwater pollution level classification monitoring method based on dynamic pH sensing as described in claim 1 or 8, characterized in that, The output also includes a probability distribution of the pH time-series signal belonging to each preset pollution level.

10. The groundwater pollution level classification monitoring method based on dynamic pH sensing as described in claim 1, characterized in that, The method is implemented through a sliding time window, which moves on the real-time acquired pH time-series signal with a preset step size to achieve dynamic classification of the degree of groundwater pollution.

Citation Information

Patent Citations

  • Underground water level dynamic characteristic analysis method based on quantitative indexes

    CN118260644A

  • Water quality change trend rapid prediction method based on multi-source data fusion and physical constraint

    CN120598102A

  • Underground water arsenic pollution range dynamic delimiting method based on multi-source data fusion

    CN120745491A

  • Real-time water quality detection system

    CN121071301A

  • Hexavalent chromium water quality online monitoring system and method based on deep learning

    CN121141564A