Water environment pollution monitoring system based on data analysis

By reconstructing the high-dimensional attractor trajectory matrix using a distributed spectral sensing node array and chaotic mapping theory, and combining it with adaptive fractal dimension calculation, efficient dynamic identification and accurate prediction of water pollution are achieved, solving the problem of nonlinear diffusion process in water pollution monitoring in existing technologies.

CN122451597BActive Publication Date: 2026-08-25NINGBO UNIV +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610913027.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-06-24
Publication Date
2026-08-25
Estimated Expiration
2046-06-24

AI Technical Summary

Technical Problem

Existing water pollution monitoring methods are unable to effectively capture the nonlinear diffusion process and dynamic characteristics of pollutants in water bodies, resulting in high false alarm and false negative rates, and making it difficult to make reliable diffusion path predictions.

Method used

A distributed spectral sensing node array is used to acquire multi-band optical signals of water bodies in real time. The high-dimensional attractor trajectory matrix is ​​reconstructed using chaotic mapping theory, and the singularity index sequence is calculated by adaptive fractal dimension. Combined with a historical pollution sample database, pollution events are identified and diffusion paths are predicted.

Benefits of technology

It improves the ability to detect weak early pollution signals, enhances the accuracy of pollution event identification and the reliability of diffusion path prediction, and reduces false alarm and false negative rates.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122451597B_ABST
    Figure CN122451597B_ABST
Patent Text Reader

Abstract

The application discloses a water environment pollution monitoring system based on data analysis and belongs to the technical field of environment monitoring and data analysis. The system comprises the following modules: a spectrum collection module, which collects multi-band water body optical signals in real time through a distributed spectrum sensing node array to form a primary monitoring data set; a phase space reconstruction module, which reconstructs the phase space of time series spectrum signals by using a chaotic mapping theory to generate a high-dimensional attractor trajectory matrix representing the pollution diffusion state; a fractal feature extraction module, which extracts a singularity index sequence reflecting the pollution source intensity change from the high-dimensional attractor trajectory matrix by using an adaptive fractal dimension calculation method; and a pollution identification module, which dynamically identifies a pollution event triggering threshold and outputs a pollution source positioning and diffusion path prediction result based on the correlation integral analysis of the singularity index sequence and a historical pollution sample database.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of environmental monitoring and data analysis technology, specifically a water pollution monitoring system based on data analysis. Background Technology

[0002] Water pollution monitoring typically relies on fixed-point, timed sampling combined with laboratory physicochemical analysis, or single-point online monitoring using electrochemical and optical probes. The data acquired by these methods are sparse both temporally and spatially, making it difficult to reflect the nonlinear diffusion process and dynamic evolution characteristics of pollutants in water bodies. To improve sensing density, distributed spectral sensing node arrays have been gradually introduced to simultaneously acquire multi-band optical signals from water bodies, including ultraviolet, visible, and near-infrared light. However, conventional analyses of such high-dimensional time-series data often employ statistical regression, principal component analysis, or deep learning-based pattern recognition methods, which essentially still treat spectral signals as linear stationary processes or static patterns, failing to effectively capture the chaotic dynamics inherent in pollution diffusion. When pollutant transport exhibits strong nonlinear fluctuations due to the coupling effects of turbulence, convection, and biochemical reactions, existing linear characteristic quantities and fixed threshold discrimination methods are prone to high false alarms and false negatives, and struggle to provide reliable diffusion path predictions. Chaotic mapping and fractal theory provide tools for describing the state of complex systems. However, in water pollution monitoring, how to construct a high-dimensional attractor that can characterize the pollution diffusion state based on actual multi-band spectral signals and adaptively extract singularity features without prior scale remains a pressing technical challenge. Specifically, the following problems need to be addressed: how to use chaotic mapping theory to reconstruct a high-dimensional attractor trajectory matrix that integrates multi-band information and can characterize the nonlinear state of pollution diffusion from single-node multi-band time-series spectral signals; and how to adaptively calculate the fractal dimension at the local scale in this high-dimensional phase space to obtain a singularity index sequence sensitive to changes in pollution source intensity. Summary of the Invention

[0003] This paper presents a water environment pollution monitoring system based on data analysis. It can reconstruct the high-dimensional attractor trajectory of pollution diffusion state from the chaotic dynamic characteristics of multi-band spectral signals, and adaptively extract the singularity index reflecting the change of pollution source intensity, so as to realize the dynamic identification of pollution events, pollution source location and diffusion path prediction.

[0004] The objective of this invention can be achieved through the following technical solutions:

[0005] This invention provides a data analysis-based water environment pollution monitoring system, comprising a spectral acquisition module, a phase space reconstruction module, a fractal feature extraction module, and a pollution identification module. The spectral acquisition module, through a distributed array of spectral sensing nodes deployed in the target water area, acquires a primary monitoring dataset containing multi-band optical signals of the water body in real time. Preferably, the distributed spectral sensing node array is arranged in a triangular non-uniform grid topology at different depths underwater. Each sensing node synchronously acquires water body optical signals containing at least ultraviolet, visible, and near-infrared bands within a continuous time window, and adds a node identification code, a depth identification code, and an acquisition timestamp to each signal, thereby forming a primary monitoring dataset with spatiotemporal multidimensional attributes, achieving high spatiotemporal resolution three-dimensional perception of pollution plumes.

[0006] The phase space reconstruction module utilizes chaotic mapping theory to reconstruct the phase space of the time-series spectral signals in the original monitoring dataset, generating a high-dimensional attractor trajectory matrix representing the pollution diffusion state. As a preferred embodiment of this invention, the module extracts single-band time-series spectral signal sequences from the original monitoring dataset by node and depth, embeds them into a high-dimensional phase space with a preset embedding dimension using a delayed coordinate state vector reconstruction method, generating initial attractor trajectory point clouds. Then, the initial attractor trajectory point clouds of different bands from the same sensing node are aligned along the time axis and tensor-stitched to generate a multi-band fused high-dimensional attractor trajectory matrix. Particularly during tensor stitching, the initial attractor trajectory point clouds of different bands are weighted and fused according to the signal-to-noise ratio, enhancing the phase space representation weight of the pollution-sensitive bands, making the reconstructed attractor trajectory matrix more significantly carry the nonlinear dynamic characteristics of pollutants. In a further preferred embodiment, during the reconstruction of the delayed coordinate state vector, the time-series spectral signal sequence is first linearly normalized. The delay coordinate step size parameter is determined based on the first zero-crossing point of the autocorrelation function of the normalized sequence, and the embedding dimension parameter is determined based on the pseudo-nearest neighbor point proportional convergence curve. Thus, each sampling point is expanded into a state vector of the corresponding dimension, and arranged in time order to form an initial attractor trajectory point cloud, ensuring the fidelity and dynamic equivalence of the phase space reconstruction.

[0007] The fractal feature extraction module extracts a singularity index sequence reflecting changes in pollution source intensity based on the high-dimensional attractor trajectory matrix using an adaptive fractal dimension calculation method. Specifically, for each state point in the high-dimensional attractor trajectory matrix, a hypersphere neighborhood of different scales is defined in the phase space centered on that point. The number of other state points contained within each hypersphere neighborhood is counted, and a power-law relationship curve between scale and number of points is established. Based on the local slope changes of the power-law relationship curve, a scale-free interval is adaptively selected, and the local fractal dimension of the state points is calculated within this interval. Finally, the local fractal dimensions of all state points are arranged in chronological order to form a singularity index sequence. This adaptive process obtains the slope-scale variation curve by performing first-order numerical differentiation on the power-law relationship curve. It identifies continuous scale intervals where the absolute value of the slope is less than a preset flatness threshold as scale-free intervals. Within these intervals, linear least-squares fitting is used to obtain the slope of a straight line as the local fractal dimension, thereby accurately characterizing the concentration variations of the attractor orbits in the phase space of the pollution plume. This allows the singularity index to sensitively and stably reflect the instantaneous fluctuations in pollution source intensity. To suppress boundary effects, a mirror-extended neighborhood strategy is employed when calculating the edge state points of the high-dimensional attractor trajectory matrix, ensuring the integrity and reliability of the singularity index sequence across the entire space.

[0008] The pollution identification module dynamically identifies pollution event trigger thresholds based on the correlation integral analysis between the singularity index sequence and the historical pollution sample database, and outputs pollution source location and diffusion path prediction results. In a preferred implementation, this module extracts historical singularity index sequences within a preset time window before each pollution event from the historical pollution sample database to construct a historical pollution sample feature library. The database contains pollution source types, pollution intensity levels, and corresponding high-dimensional attractor trajectory templates. A sliding time window is used to segment the real-time calculated singularity index sequence into multiple matching segments, and the correlation integral similarity between each matching segment and each feature sequence in the historical feature library is calculated. When the correlation integral similarity of a matching segment exceeds a preset alarm threshold, it is determined to be a pollution event trigger state. This method fully utilizes the long-term correlation characteristics of chaotic attractors, enabling accurate identification of abnormal patterns in the early stages of pollution diffusion, achieving high-precision pollution source location and dynamic prediction of diffusion paths, and significantly improving the response speed and early warning reliability for sudden water pollution events.

[0009] The beneficial effects of this invention are:

[0010] A delayed coordinate state vector reconstruction method is employed. The delayed coordinate step size parameter is determined based on the first zero-crossing position of the autocorrelation function of a single-band time-series spectral signal, and the embedding dimension parameter is determined based on the pseudo-nearest neighbor proportional convergence curve. The single-band time-series spectral signal sequence is embedded into a high-dimensional phase space to generate an initial attractor trajectory point cloud. The initial attractor trajectory point clouds obtained from different bands of the same sensing node are aligned along the time axis and stitched together using a signal-to-noise ratio (SNR) weighted tensor. During the stitching process, higher weights are assigned to pollution-sensitive bands with higher SNR, forming a multi-band fused high-dimensional attractor trajectory matrix. This matrix can completely preserve the nonlinear chaotic dynamic structure of the spectral signal in the high-dimensional phase space, allowing weak state perturbations caused by pollution source changes to produce discernible geometric deformations in the attractor trajectory, thereby significantly improving the detection capability of early weak pollution signals. On the high-dimensional attractor trajectory matrix, a hypersphere neighborhood is defined for each state point at different scales, and the number of state points within the neighborhood is counted, establishing a power-law relationship curve between scale and the number of points. By obtaining the slope-scale variation curve through the first-order numerical differentiation of the power-law relationship curve, continuous scale intervals with slope absolute values ​​less than a preset flatness threshold are identified as scale-free intervals. Within these intervals, the local fractal dimension is calculated using linear least-squares fitting, forming a singularity index sequence. A mirror-extended neighborhood strategy is employed for the edge state points of the trajectory matrix to suppress the interference of boundary effects on the fractal dimension calculation. The adaptive selection mechanism of scale-free intervals allows the fractal dimension calculation to be unrestricted by a preset fixed scale, enabling dynamic responses to changes in the scale structure within the data. The obtained singularity index sequence accurately reflects the evolution of the local geometric structure of attractors caused by changes in pollution source intensity, improving the accuracy of pollution event trigger identification and providing more reliable dynamic characteristics for diffusion path prediction. Attached Figure Description

[0011] The invention will now be further described with reference to the accompanying drawings.

[0012] Figure 1 This is a schematic diagram of a water environment pollution monitoring system based on data analysis.

[0013] Figure 2 This is a flowchart of the original monitoring dataset acquisition process;

[0014] Figure 3 This is a flowchart of pollution event detection and pollution source localization based on singularity index sequence matching. Detailed Implementation

[0015] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0016] See Figure 1 This invention provides a water environment pollution monitoring system based on data analysis, comprising: a spectral acquisition module for real-time acquisition of a primary monitoring dataset containing multi-band optical signals of the water body through a distributed array of spectral sensing nodes deployed in the target water area; a phase space reconstruction module for reconstructing the phase space of the time-series spectral signals in the primary monitoring dataset using chaotic mapping theory to generate a high-dimensional attractor trajectory matrix characterizing the pollution diffusion state; a fractal feature extraction module for extracting a singularity index sequence reflecting the change in pollution source intensity based on the high-dimensional attractor trajectory matrix using an adaptive fractal dimension calculation method; and a pollution identification module for dynamically identifying pollution event trigger thresholds based on the correlation integral analysis between the singularity index sequence and a historical pollution sample database, and outputting pollution source location and diffusion path prediction results.

[0017] In practice, the spectral acquisition module obtains the initial monitoring dataset in the following manner: A distributed spectral sensor node array is deployed in the underwater space of the target water area using a triangular non-uniform grid topology. (See also...) Figure 2 Based on the shoreline boundary and underwater topography data of the target water area, a triangular grid covering the entire monitoring area is generated on a horizontal projection plane. The vertices of the triangular grid are the projection points of the sensor nodes on the water surface. The non-uniformity of the grid is achieved by controlling the side length of the triangular units: in areas with potential pollution source influence and complex flow fields, the side length of the triangular units is set to a first length value; in open areas with stable flow fields, the side length of the triangular units is set to a second length value, where the first length value is less than the second length value. The first and second length values ​​are determined based on the historical diffusion scale of the target water area and the sensor's sensing radius, and the first length value is not less than half of the sensor's sensing radius. Each triangle vertex corresponds to a horizontal position of a sensor node, and the sensor nodes are deployed at multiple depth levels along the vertical direction. The depth levels are determined using non-uniform intervals, with the first... The depth value of each depth layer is calculated using the following formula:

[0018]

[0019] in, The depth identifier is The depth of the water layer from the water surface. The reference depth value representing the shallowest layer. For depth level ordinal index, Indicates the first Layer and First Vertical spacing between layers. The value is set based on the vertical mixing intensity gradient of the water column: within the depth range where the rate of change of vertical mixing intensity is greater than the preset gradient, Take a smaller interval value; in the depth range where the rate of change of vertical mixing intensity is less than or equal to the preset gradient, A larger interval value is selected. The vertical mixing intensity change rate is calculated based on pre-measured temperature and velocity profiles, and the preset gradient is determined based on the stratification stability threshold. Each sensor node is assigned a unique node identifier, which consists of a numerical code containing location and depth layer information. After the deployment of each sensor node, a continuous time window is set. The duration of the continuous time window is determined based on the pollutant diffusion timescale and monitoring timeliness requirements, for example, a continuous time window of 30 minutes. Within the continuous time window, all sensor nodes synchronize their time with millisecond-level accuracy. Time synchronization is achieved using a combination of network time protocol and satellite time synchronization. Each sensor node has a built-in multi-band miniature spectral sensor, which includes ultraviolet, visible, and near-infrared photosensitive channels, covering wavelength ranges of 200 nm to 400 nm, 400 nm to 700 nm, and 700 nm to 1700 nm, respectively. The ultraviolet (UV), visible, and near-infrared (NII) photosensitive channels acquire water optical signals synchronously at a predetermined sampling frequency of 1 Hz. At each sampling moment, the sensor node simultaneously records the UV, visible, and NII light intensity sequences, forming a single multi-band water optical signal sampling frame. For each multi-band water optical signal sampling frame, the signal processing unit within the sensor node adds a node identifier code, a depth identifier code, and a timestamp in real time. The node identifier code is a unique code fixed at the factory; the depth identifier code is the numerical value corresponding to the depth layer where the sensor node is deployed, with a value of [value missing]. Corresponding depth identifier code The data acquisition timestamps are in Coordinated Universal Time (UTC) format, accurate to milliseconds. The data records with appended identifiers include the following fields: node identifier code, depth identifier code, acquisition timestamp, ultraviolet (UV) light intensity value sequence, visible light intensity value sequence, and near-infrared (NIIR) light intensity value sequence. Within a continuous time window, all sensor nodes continuously execute the above acquisition and identifier appending process, compiling all generated data records into the original monitoring dataset. The original monitoring dataset is stored in chronological order, partitioned and stored on solid-state storage or a remote data server. The storage structure uses a columnar storage format to support efficient retrieval based on node identifier codes and depth identifier codes.

[0020] In practical implementation, the phase space reconstruction module extracts single-band time-series spectral signal sequences of specific sensing nodes from the original monitoring dataset. Node identifiers and depth identifiers are selected as search criteria, and the original monitoring dataset is traversed to filter out all data records where the node identifier equals the target node identifier and the depth identifier equals the target depth identifier. The filtered results are sorted in ascending order by acquisition timestamp to form a time-series dataset. The ultraviolet (UV) light intensity value sequence is extracted from the time-series dataset and denoted as... ,in For sampling sequence number, , This represents the total number of ultraviolet band sampling points within a continuous time window. Similarly, the visible light intensity sequence and near-infrared light intensity sequence for the same sensing node within the same continuous time window are extracted.

[0021] A delayed coordinate state vector reconstruction method is used to embed a single-band time-series spectral signal sequence into a high-dimensional phase space with a preset embedding dimension, generating an initial attractor trajectory point cloud. The preset embedding dimension parameters are set. and delay time parameters Preset embedding dimension parameters The value of is determined by analyzing the convergence curve of the pseudo-nearest neighbor ratio of the ultraviolet band light intensity value sequence: the pseudo-nearest neighbor ratio of the ultraviolet band light intensity value sequence under different embedding dimensions is calculated, the curve of the pseudo-nearest neighbor ratio changing with the embedding dimension is plotted, and the minimum embedding dimension value corresponding to the convergence of the pseudo-nearest neighbor ratio to zero is used as the preset embedding dimension parameter. Delay time parameter The value of is determined by the first zero-crossing point of the autocorrelation function of the ultraviolet band light intensity value sequence: calculate the autocorrelation function of the ultraviolet band light intensity value sequence, identify the delay index when the autocorrelation function value first crosses zero, and use this delay index value as the delay time parameter. .

[0022] The state vector of the ultraviolet band light intensity value sequence is constructed using the delayed coordinate state vector reconstruction method. For each element of the ultraviolet band light intensity value sequence... , build a A state vector is defined as follows:

[0023]

[0024] in, Indicates the ultraviolet band in the 1st The reconstructed state vector at each sampling point, with superscript... Identify the ultraviolet band; This represents the first value in the ultraviolet band light intensity sequence. Each sample value; This is the delay time parameter; Preset embedding dimension parameters; Indicates the transpose operation; The range of values ​​is All valid The state vector corresponding to the value is... Arranged in ascending order, the initial attractor trajectory point cloud in the ultraviolet band is obtained, denoted as... .

[0025] The visible light intensity value sequences and near-infrared light intensity value sequences are processed in the same way. For the visible light intensity value sequences, the same preset embedding dimension parameters are used. and delay time parameters Constructing the state vector yields the initial attractor trajectory point cloud in the visible band. For the near-infrared light intensity sequence, a state vector is constructed to obtain the initial attractor trajectory point cloud in the near-infrared band. The effective sampling point time ranges corresponding to the initial attractor trajectory point clouds of the three bands are consistent, all falling within the original sampling sequence number interval. .

[0026] The initial attractor trajectory point clouds of different bands from the same sensing node are aligned along the time axis and then tensor-stitched to generate a multi-band fused high-dimensional attractor trajectory matrix. Time axis alignment uses sampling sequence alignment: for the same sampling sequence... Ultraviolet band state vector Visible light band state vector and near-infrared band state vector For the same actual sampling time, they are stitched together according to the band dimension. The tensor stitching operation is performed along the newly created band dimension, making all three dimensions equal to... The state vector is expanded into a dimension of Fusion state vector The concatenation calculation expression is: Arrange the fused state vectors corresponding to all sampling numbers in order of sampling number, forming a matrix with [number of rows]. The number of columns is The matrix is ​​the high-dimensional attractor trajectory matrix of multi-band fusion. Each row of the high-dimensional attractor trajectory matrix corresponds to a phase space fusion state point at a sampling time, and the column index of the matrix is ​​determined by the band identifier and the embedding dimension offset.

[0027] In practical implementation, the single-band time-series spectral signal sequence undergoes linear normalization processing, mapping the amplitude range of the single-band time-series spectral signal sequence to a preset standard interval. This is illustrated by the ultraviolet band light intensity value sequence. For example, calculate the ultraviolet band light intensity value sequence. maximum value and minimum value ,in , , For sampling sequence number, , This represents the total number of ultraviolet band sampling points within a continuous time window. The preset standard interval is set to... Interval. Linear normalization is performed by calculating the normalized value point by point according to the following formula. :

[0028]

[0029] in, This represents the first value in the ultraviolet band light intensity sequence. The numerical value obtained after linear normalization of each sample value. This represents the first value in the ultraviolet band light intensity sequence. Each sample value, This represents the maximum value in the ultraviolet band light intensity sequence. This represents the minimum value in the ultraviolet band light intensity sequence. The visible light intensity sequence and the near-infrared light intensity sequence are processed in the same way to obtain the normalized visible light time-series spectral signal sequence and the normalized near-infrared time-series spectral signal sequence.

[0030] The delay coordinate step size parameter is determined based on the first zero-crossing point of the autocorrelation function of the normalized ultraviolet (UV) band time-series spectral signal sequence. The autocorrelation function of the normalized UV band time-series spectral signal sequence is then calculated. , ,in For lazy indexing, , Less than In delayed indexing Depend on Increment to During the process, record the autocorrelation function The numerical change. When When the value is less than or equal to zero for the first time, the index will be delayed. The corresponding numerical value is denoted as Then the delay coordinate step size parameter Set as If the autocorrelation function is in If no value less than or equal to zero is found within the range, then the autocorrelation function is selected to decrease to the initial value. of The corresponding delay index is used as the delay coordinate step size parameter. ,in is the base of the natural logarithm.

[0031] The embedding dimension parameter is determined based on the pseudo-nearest neighbor ratio convergence curve of the normalized ultraviolet band time-series spectral signal sequence. Given the embedding dimension... Below, using the delay coordinate step size parameter Construct state vectors to obtain a point set. For each state vector in the point set, calculate the Euclidean distance between that state vector and the remaining state vectors in the set, and find its nearest neighbor state vector. This is done in the higher-dimensional embedding dimension. Next, the state vector point set is reconstructed, and the distance change ratio of the original nearest neighbor state vector pair under the new embedding dimension is calculated. If the distance change ratio is greater than a preset distance change ratio threshold, the original nearest neighbor state vector pair is determined to be a pseudo-nearest neighbor point vector pair. The proportion of pseudo-nearest neighbor point vector pairs in all state vectors is counted to obtain the embedding dimension. The proportion of pseudo-nearest points. The preset distance change rate threshold is set to... This setting is based on the empirical criterion for distinguishing between true and false neighbors in high-dimensional phase space reconstruction, and is greater than... This indicates that two points appear to be near each other in low-dimensional space due to projection effects, but this is actually pseudo-nearest neighbor. From the embedding dimension... We begin by calculating the proportion of pseudo-nearest neighbors, gradually increasing the embedding dimension, and plotting a curve showing how the proportion of pseudo-nearest neighbors changes with the embedding dimension. When the proportion of pseudo-nearest neighbors converges to less than... Furthermore, when the embedding dimension no longer decreases significantly with increasing embedding dimension, the corresponding minimum embedding dimension is used as the embedding dimension parameter. . The convergence threshold is set based on the separation characteristics of adjacent trajectories in chaotic systems. A false neighbor ratio below this threshold indicates that the attractor structure has been fully unfolded geometrically.

[0032] According to the delay coordinate step size parameter and embedded dimension parameters This involves expanding each sampling point in the normalized ultraviolet (UV) band time-series spectral signal sequence into a state vector with an embedding dimension. For the normalized UV band time-series spectral signal sequence... Construct state vector ,in Indicates the ultraviolet band in the 1st The reconstructed state vector at each sampling point The range of values ​​is All The corresponding state vectors are arranged in chronological order to form the point cloud of the initial attractor trajectory in the ultraviolet band, denoted as the point set. The normalized visible light band time-series spectral signal sequences and the normalized near-infrared band time-series spectral signal sequences were processed using the same method to obtain the initial attractor trajectory point cloud in the visible light band. and near-infrared band initial attractor trajectory point cloud .

[0033] During tensor stitching, point clouds of initial attractor trajectories from different wavelength bands are fused using a signal-to-noise ratio weighting method. The intensity sequence in the ultraviolet band is then calculated. average signal-to-noise ratio The calculation method is as follows: the root mean square value of all sampled values ​​in the ultraviolet band light intensity value sequence is taken as the effective signal value, and the root mean square value of the floor noise recorded by the ultraviolet band photosensitive channel of the sensing node under no-light conditions is taken as the effective noise value. The ratio of the effective signal value to the effective noise value is taken as... Enlarge the base logarithm The average signal-to-noise ratio in the ultraviolet band was obtained by multiplying the signal-to-noise ratio by The unit is decibels. The mean signal-to-noise ratio in the visible light band is calculated in the same way. and the average signal-to-noise ratio in the near-infrared band Then calculate the weighting coefficient for the ultraviolet band. Visible light band weighting coefficient and near-infrared band weighting coefficient The calculation expression is: , , For each sampling number Construct a weighted fusion state vector The splicing method is ,in Represents the point cloud of the initial attractor trajectory in the ultraviolet band. The Middle A state vector, Point cloud representing the initial attractor trajectory in the visible light band The Middle A state vector, Point cloud representing the initial attractor trajectory in the near-infrared band The Middle A number of state vectors. The weighted fusion state vectors corresponding to all sampling numbers are arranged sequentially to obtain a high-dimensional attractor trajectory matrix for multi-band fusion.

[0034] In practical implementation, the fractal feature extraction module delineates hyperspheres of different scales in the phase space for each state point in the high-dimensional attractor trajectory matrix, centered on that state point. Let the high-dimensional attractor trajectory matrix contain... There are n state points, and the dimension of each state point is . , This is equal to the column number of the high-dimensional attractor trajectory matrix obtained through multi-band fusion. The column number is taken from the high-dimensional attractor trajectory matrix. There are n state points, denoted by the state point vector. ,in Define the hypersphere radius sequence. , , Let be the total number of scales in the hypersphere radius sequence. The th in the hypersphere radius sequence... hypersphere radius After statistical analysis of the pairwise Euclidean distances between all state points in the high-dimensional attractor trajectory matrix, the following method was determined: calculate the Euclidean distances between all pairs of state points to obtain a distance set, and use the minimum value of the distance set as the minimum radius. The maximum value of the distance set is taken as the maximum radius. On a logarithmic scale, the interval Evenly divided into Given several intervals, take the anti-corresponding values ​​of the endpoints of each interval to form a sequence of hypersphere radii. .in Set as This setting is based on a balance between the scale resolution requirements and computational complexity of phase space scale analysis. Each scale can finely characterize the power-law relationship curve without significantly increasing the computational burden.

[0035] by state point Centered on the hypersphere radius With radius, in Constructing a hypersphere neighborhood in phase space The hypersphere neighborhood is defined as ,in Indicates that in the high-dimensional attractor trajectory matrix, except for The other A vector of state points, and , This represents the Euclidean distance operator.

[0036] Count the number of other state points contained in the hypersphere neighborhood at each scale. For a fixed number of state points... and a fixed hypersphere radius Calculate the hypersphere neighborhood Number of other state points included , ,in This indicates a counting operation. For all calculate , obtain the state point Corresponding scale-point sequence Establish state points The power-law relationship curve between the scale and the number of points is plotted in a double logarithmic coordinate system, with the x-axis representing the hypersphere radius. by logarithm with base The vertical axis represents the number of state points included. by logarithm with base .

[0037] In the adaptive fractal dimension calculation method, a mirror-image neighborhood expansion strategy is adopted for the edge state points of the high-dimensional attractor trajectory matrix. All state points in the high-dimensional attractor trajectory matrix are traversed, and the edge degree of each state point is determined. The state points are then calculated. The signed distance to the boundary of the point set envelope in phase space is obtained by calculating the coordinate range of the high-dimensional attractor trajectory matrix in each dimension. Dimension, coordinate range is ,in , , State point In the The coordinates of the dimension. If the state point In the 3D coordinates satisfy or Then the state point Points are identified as edge states in this dimension. In the Before constructing a hypersphere neighborhood centered on the boundary, the coordinates of the state points are mirrored along the dimensions spanning the boundary to generate virtual extension points. If the... Wei Shang near On one side of the boundary, the mirror extension point is at the . 3D coordinates ;like near On one side of the boundary, the mirror extension point is at the . 3D coordinates The mirror extension point and the original state point are in addition to the first... The coordinate values ​​are the same in all dimensions other than the first dimension. The mirror extension points are added to the high-dimensional attractor trajectory matrix to form an extension matrix. Based on this, the hypersphere neighborhood is constructed and the number of state points is counted, so that the neighborhood of the edge state points can cross the boundary of the original phase space to obtain a symmetrically extended neighborhood state point distribution.

[0038] Based on the local slope variation of the power-law relationship curve, a scale-free interval is adaptively selected. The first-order numerical derivative of the power-law relationship curve is performed to obtain the slope variation curve with scale. This is then applied to discrete data points in a double logarithmic coordinate system. The central difference method is used to calculate the scale. Local value of slope at ,when The time calculation formula is ,when When forward difference is used, when Back-difference is used to obtain the slope sequence. .

[0039] On the slope variation curve, continuous scale intervals where the absolute value of the slope is less than a preset flatness threshold are identified as scale-free intervals. The preset flatness threshold is set to... This setting is based on the statistical distribution characteristics of the local slope change rate in fractal scaling analysis. The corresponding slope fluctuation range can accommodate conventional noise perturbations in phase space attractor dimension estimation, while excluding significant slope inflections caused by nonlinear effects. (Scanning slope sequence) Find satisfaction And the continuous length is greater than or equal to All subsequences, where This represents the average slope within the subsequence. Among all subsequences that meet the conditions, the subsequence with the most points is selected as the state point. The scale-free interval, the scale-free interval corresponds to the scale index interval. ,in This is the starting scale index for the scale-free interval. This is the terminating scale index for scale-free intervals.

[0040] In scale-free intervals Within this model, a linear least squares fitting method is used to fit a straight line to the double logarithmic relationship between scale and number of points. Let the equation of the fitted line be:

[0041]

[0042] in, The slope parameter represents the fitted line. This represents the intercept parameter of the fitted line. Represented by state point Centered on, with a hypersphere radius of The number of other state points contained within the neighborhood of the time-supersphere. The first in the hypersphere radius sequence The goal of linear least squares fitting is to minimize the sum of squared errors. By solving and Obtain slope parameters and intercept parameter The estimated value. The slope value of the fitted straight line. As a state point The local fractal dimension. For all state points in the high-dimensional attractor trajectory matrix. By performing the above calculation process, a local fractal dimension sequence is obtained. The local fractal dimensions of all state points in the high-dimensional attractor trajectory matrix are arranged in order of the sampling time corresponding to the state points to form a singularity index sequence.

[0043] In practice, the pollution identification module extracts the historical singularity index sequence of each pollution event within a preset time window prior to its occurrence from the historical pollution sample database. (See [reference]) Figure 3 A historical pollution sample feature database was constructed. This database stores records of multiple confirmed pollution events. Each event record includes a pollution source type identifier, a pollution intensity level identifier, a timestamp of the event, and a corresponding high-dimensional attractor trajectory template. The pollution source type identifier includes classification codes for chemical pollution, biological pollution, and suspended solids pollution. The pollution intensity level identifier is divided into three levels: severe pollution, moderate pollution, and light pollution. The high-dimensional attractor trajectory template is a multi-band fused high-dimensional attractor trajectory matrix generated by the phase space reconstruction module within a preset time window before the pollution event.

[0044] The first in the historical pollution sample database For each pollution event record, extract the historical singularity index sequence within a preset time window preceding the pollution event. The duration of the preset time window is set to... This setting, measured in minutes, is based on statistics of the average diffusion time from the onset of an abnormal state to the triggering of an alarm in typical water pollution events. The minute window covers the typical transport period of the pollution plume from the source region to the monitoring node. It is truncated backward from the moment the pollution event occurs. The singularity index sequence corresponding to a duration of minutes is calculated by the fractal feature extraction module within the corresponding time period and stored with associated timestamps. The extracted historical singularity index sequence is denoted as... ,in The length of the singularity index sequence within the preset time window. , This represents the total number of pollution event records in the historical pollution sample database. It is associated with and stored in conjunction with the corresponding pollution source type identifier, pollution intensity level identifier, and high-dimensional attractor trajectory template to form a feature record in the historical pollution sample feature library.

[0045] A sliding time window approach is used to segment the real-time computed singularity index sequence into multiple segments to be matched. The real-time computed singularity index sequence is generated online by the fractal feature extraction module, denoted as . ,in The length of the currently accumulated real-time singularity index sequence. For the first The singularity index value at each moment. Set the length of the sliding window. With preset time window length Consistency, that is The step size of the sliding window is set to... 1 sampling point. Indexed by the start position of the window. from Increment to , segment out the first One segment to be matched All segments to be matched constitute the set of segments to be matched.

[0046] Calculate the integral similarity between each segment to be matched and each feature sequence in the historical contaminated sample feature database. For each segment to be matched... Historical singularity index sequences in historical pollution sample feature library The correlation integral similarity is calculated as follows:

[0047]

[0048] in, Indicates the segment to be matched With historical singularity index sequence The correlation integral similarity value between them; This represents the length of the sliding window, which is the number of singularity index values ​​within the segment to be matched; Indicates the segment to be matched The index of the singularity index value, with a range of values ​​of 1000. arrive ; Represents the historical singularity index sequence The index of the singularity index value, with a range of values ​​of 1000. arrive ; Indicates the segment to be matched The first in A singularity index value; Represents the historical singularity index sequence The first in A singularity index value; This represents the absolute value operation; Indicates the preset distance threshold; Represents the Heaviside step function, when hour Values Otherwise, the value is .

[0049] Preset distance threshold The index is determined based on the statistical distribution characteristics of historical singularity index sequences. The first-order difference absolute value sequence of all historical singularity index sequences in the historical pollution sample feature database is calculated, and the i-th value is obtained by merging all the first-order difference absolute value sequences. Percentile, this percentile is used as the preset distance threshold. The setting value. This setting method makes... It can adapt to the local fluctuations of the singularity index, and Percentiles allow the correlation integral similarity to have a certain tolerance for small numerical perturbations.

[0050] For a given fragment to be matched Traverse all historical pollution sample feature databases 1 feature record, calculated The correlation integral similarity value Obtain the maximum correlation integral similarity value. , .

[0051] When the correlation similarity of the segments to be matched exceeds a preset alarm threshold, a contamination event is triggered. The preset alarm threshold is set as follows: This setpoint was determined through leave-one-out cross-validation on a historical contaminated sample database: each historical contaminated sample feature record was sequentially extracted as a positive example; the correlation integral similarity distribution between the positive example and the remaining feature records in the database, as well as the correlation integral similarity distribution between the positive example and a large number of singularity index sequence fragments from non-contaminated periods, were calculated; and the threshold that maximized the detection accuracy was selected. The optimal cutoff point for balancing the false alarm rate and the missed detection rate is determined by scanning all segments to be matched in the current real-time accumulated singularity index sequence. If at least one segment to be matched is detected... satisfy If so, the system is determined to have entered a contamination event triggered state.

[0052] After determining that a pollution event has been triggered, the system outputs the pollution source location and diffusion path prediction results. Historical pollution sample feature records corresponding to the segments to be matched that maximize the correlation integral similarity are extracted, and the pollution source type identifier, pollution intensity level identifier, and high-dimensional attractor trajectory template are obtained from these records. Trajectory matching is performed between the high-dimensional attractor trajectory template from the historical pollution sample feature records and the currently generated high-dimensional attractor trajectory matrix. Using the spatial topological coordinates of the sensor nodes and the triangular non-uniform grid topology, the most likely release location of the pollution source is inverted. Combining the flow field model and diffusion parameters of the target water area, and starting from the pollution source location, a spatiotemporal evolution sequence of the pollution diffusion path is generated based on the spatial gradient direction of the singularity index sequence and the extrapolated time step, serving as the diffusion path prediction result. The pollution source location result is output in three-dimensional coordinate form, and the diffusion path prediction result is output in the form of a time-series three-dimensional trajectory point set.

[0053] The foregoing has provided a detailed description of one embodiment of the present invention, but this description is merely a preferred embodiment and should not be construed as limiting the scope of the invention. All equivalent variations and modifications made within the scope of the claims of this invention should still fall within the patent coverage of this invention.

Claims

1. A water environment pollution monitoring system based on data analysis, characterized in that, include: The spectral acquisition module is used to acquire raw monitoring datasets containing multi-band optical signals of the water body in real time through a distributed array of spectral sensing nodes deployed in the target water area. The phase space reconstruction module is used to reconstruct the phase space of the time-series spectral signals in the original monitoring dataset using chaotic mapping theory, generating a high-dimensional attractor trajectory matrix characterizing the pollution diffusion state. Specifically, it includes: Extract the single-band time-series spectral signal sequence of a single sensor node within a continuous time window from the original monitoring dataset according to the node identifier code and depth identifier code; The delayed coordinate state vector reconstruction method is used to embed a single-band time-series spectral signal sequence into a high-dimensional phase space with a preset embedding dimension to generate an initial attractor trajectory point cloud. The initial attractor trajectory point clouds of different bands of the same sensing node are aligned along the time axis and tensor stitched together to generate a high-dimensional attractor trajectory matrix fused by multiple bands. During the tensor stitching process, the initial attractor trajectory point clouds of different bands are fused by weighting according to the signal-to-noise ratio to enhance the phase space representation weight of the pollution-sensitive bands. The fractal feature extraction module is used to extract the singularity index sequence reflecting the change in pollution source intensity based on the high-dimensional attractor trajectory matrix and an adaptive fractal dimension calculation method. In the adaptive fractal dimension calculation method, a mirror-extended neighborhood strategy is adopted for the edge state points of the high-dimensional attractor trajectory matrix to suppress the interference of boundary effects on the singularity index calculation. The pollution identification module is used to dynamically identify pollution event trigger thresholds based on the correlation integral analysis between the singularity index sequence and the historical pollution sample database, and output pollution source location and diffusion path prediction results.

2. The water environment pollution monitoring system based on data analysis according to claim 1, characterized in that, The process involves real-time acquisition of a primary monitoring dataset containing multi-band optical signals from the water body through a distributed array of spectral sensing nodes deployed in the target water area. Specifically, this includes: Distributed spectral sensing node arrays are deployed at different underwater depths in the target water area according to a triangular non-uniform grid topology. Each sensing node synchronously acquires water optical signals, including at least the ultraviolet, visible, and near-infrared bands, within a preset continuous time window. The multi-band optical signals collected by each sensor node are supplemented with node identification codes, depth identification codes and collection timestamps to form the original monitoring dataset.

3. The water environment pollution monitoring system based on data analysis according to claim 2, characterized in that, The method employing delayed coordinate state vector reconstruction embeds a single-band time-series spectral signal sequence into a high-dimensional phase space with a preset embedding dimension to generate an initial attractor trajectory point cloud, specifically including: A single-band time-series spectral signal sequence is linearly normalized to map its amplitude range to a preset standard interval; The delay coordinate step size parameter is determined based on the first zero-crossing position of the autocorrelation function of the normalized time-series spectral signal sequence. The embedding dimension parameters are determined based on the pseudo-nearest neighbor ratio convergence curve of the normalized time-series spectral signal sequence. Based on the delay coordinate step size parameter and the embedding dimension parameter, each sampling point in the normalized time-series spectral signal sequence is expanded into a state vector of the embedding dimension, and all state vectors are arranged in time order to form the initial attractor trajectory point cloud.

4. The water environment pollution monitoring system based on data analysis according to claim 1, characterized in that, The step of extracting a singularity index sequence reflecting changes in pollution source intensity based on the high-dimensional attractor trajectory matrix and using an adaptive fractal dimension calculation method specifically includes: For each state point in the high-dimensional attractor trajectory matrix, a hypersphere neighborhood of different scales is delineated in the phase space with that state point as the center; Count the number of other state points contained in the neighborhood of the hypersphere at each scale, and establish a power-law relationship curve between scale and number of points; Based on the local slope variation of the power law relationship curve, an adaptive scale-free interval is selected, and the local fractal dimension of the state point is calculated within this interval. Arrange the local fractal dimensions of all state points in the high-dimensional attractor trajectory matrix in chronological order to form a singularity index sequence.

5. A water environment pollution monitoring system based on data analysis according to claim 4, characterized in that, The step of adaptively selecting a scale-free interval based on the local slope variation of the power-law curve, and calculating the local fractal dimension of the state points within that interval, specifically includes: By performing a first-order numerical differential on the power-law relationship curve, we obtain a curve showing how the slope changes with the scale. On the slope change curve, identify continuous scale intervals where the absolute value of the slope is less than a preset flatness threshold, and use them as scale-free intervals. Within the scale-free interval, a linear least squares fitting method is used to fit the double logarithmic relationship between scale and number of points. The slope value of the fitted line is used as the local fractal dimension of the state point.

6. A water environment pollution monitoring system based on data analysis according to claim 1, characterized in that, The correlation integral analysis based on the singularity index sequence and historical pollution sample database dynamically identifies pollution event trigger thresholds and outputs pollution source location and diffusion path prediction results, specifically including: Extract the historical singularity index sequence of each pollution event within a preset time window before its occurrence from the historical pollution sample database, and construct a historical pollution sample feature library; The singularity index sequence calculated in real time is divided into multiple segments to be matched using a sliding time window method; Calculate the integral similarity between each segment to be matched and each feature sequence in the historical contaminated sample feature library; When the correlation similarity of the fragments to be matched exceeds the preset alarm threshold, it is determined that a pollution event has been triggered.

7. A water environment pollution monitoring system based on data analysis according to claim 6, characterized in that, The historical pollution sample database contains pollution source types, pollution intensity levels, and corresponding high-dimensional attractor trajectory templates.

Citation Information

Patent Citations

  • Water quality pollution detection-based drainage basin water environment monitoring and emergency pollution rapid tracing method and system

    CN120744675A

  • Water environment pollution detection and analysis system and method based on Internet of Things

    CN121299064A