Drainage network pollution source analysis method and system based on deep learning and water quality fluorescence fingerprint spectrum, and storage medium
Patent Information
- Application Number
- CN202610986398.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-03
- Publication Date
- 2026-09-15
- Estimated Expiration
- 2046-07-03
AI Technical Summary
然而,这些现有方法要么直接依赖昂贵且操作繁琐的光谱检测实体设备,难以直接应用于排水管网的在线高频监测;要么依然受限于低频的周期性人工采样模式,无法实现真正意义上的实时敏捷溯源
[0019] The beneficial effects of this invention are as follows: The method of this invention, by collecting basic physicochemical indicators and meteorological data in real time and constructing a multimodal time-series dataset, dynamically removes the influence of meteorological factors on water quality using a dual-track long short-term memory network that integrates meteorological disturbance mechanisms, achieving high-sensitivity identification of sudden pollution events; it does not rely on expensive three-dimensional fluorescence spectroscopy online equipment, but instead uses a cross-modal mapping model to infer and reconstruct fluorescence fingerprint spectra from conventional physicochemical indicators and meteorological data in real time, significantly reducing operation and maintenance costs; furthermore, it adopts a primary and backup dual-mode redundant analytical architecture, with the deep learning model outputting pollution source contribution weights end-to-end, and the statistical inference model intervening to provide fault-tolerant analysis when there is low confidence or high noise, thereby obtaining high-precision multi-source pollution contribution distribution under various operating conditions, overcoming the shortcomings of traditional methods such as delayed early warning, low source tracing accuracy, and high equipment costs.
Smart Images

Figure CN122508440B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of municipal drainage network monitoring and smart water technology, specifically to a method, system, and storage medium for analyzing pollution sources in drainage networks based on deep learning and water quality fluorescence fingerprinting. Background Technology
[0002] Urban development, accompanied by the continuous expansion of municipal drainage networks, inevitably brings problems such as combined sewer overflows, stormwater and sewage mixing, and pipe infiltration, leading to frequent urban water pollution. Drainage network pollution is characterized by its concealment, dispersion, and suddenness, easily causing cumulative water pollution. Pollution sources in municipal drainage networks include: domestic sewage, kitchen wastewater, balcony wastewater, medical wastewater, chemical wastewater, electronic wastewater, food wastewater, and groundwater. Different pollution sources have their own characteristic fingerprints (such as three-dimensional fluorescence spectral characteristics), but in actual network operation, pollution events often manifest as multi-source mixed pollution. Accurately and quantitatively separating the contribution weight of each pollution source from mixed water bodies is the core challenge of refined pollution control in drainage networks.
[0003] Traditional pipeline pollution monitoring relies on a limited number of online monitoring points and blind manual inspections, resulting in a slow response to sudden pollution events. Existing intelligent pipeline early warning technologies, such as the early warning system based on historical pattern statistical baseline comparison disclosed in Chinese patent document CN121808418A, and the monitoring system based on invariant analysis algorithm disclosed in Chinese patent document CN119494043A, while improving data processing efficiency to some extent, still rely on fixed concentration thresholds or simple statistical deviations for alarms in conventional water quality monitoring. This early warning method is lagging and fails to fully quantify the nonlinear disturbances of external environmental factors such as meteorological conditions on the temporal fluctuations of water quality, making it difficult to quickly capture real sudden anomalies. Due to the delayed response, large amounts of untreated pollutants are directly discharged into receiving water bodies, forcing management departments to invest huge amounts of manpower and resources in end-of-pipe river management and ecological restoration.
[0004] In recent years, existing technologies have attempted to introduce relevant methods for pipeline network tracing. For example, Chinese patent document CN121207953A discloses a tracing method based on three-dimensional fluorescence spectral similarity comparison, and Chinese patent document CN119475005A discloses a technology that combines periodic sampling data input into a water gene diagnostic model to achieve pipeline network pollution tracking. However, these existing methods either directly rely on expensive and cumbersome physical spectral detection equipment, making them difficult to apply directly to online high-frequency monitoring of drainage networks; or they remain limited to low-frequency periodic manual sampling, failing to achieve truly real-time and agile tracing. In addition, some existing technologies, such as Chinese patent document CN119006250A, disclose a scheme for tracing using conventional water quality terminal sensors combined with pipeline network topological association clustering, but due to the lack of specific molecular-level fingerprint information, it is difficult to achieve accurate separation and quantification of complex multi-source mixed pollution. Meanwhile, traditional source apportionment models rely heavily on complex statistical deductions, which can easily lead to the loss of feature information and the gradual accumulation of calculation errors, severely limiting the accuracy of source apportionment and making it difficult to meet the current needs of refined pollution control in pipeline networks.
[0005] Therefore, there is an urgent need to develop a pollution source tracing technology that can integrate meteorological disturbances, achieve real-time online monitoring, and quantitatively separate the contribution weights of each pollution source, with high sensitivity, low operation and maintenance costs, and high accuracy. Summary of the Invention
[0006] The technical problem to be solved by the present invention is to provide a highly sensitive, low-maintenance-cost, and high-precision method, system, and storage medium for analyzing pollution sources in drainage pipe networks based on deep learning and water quality fluorescence fingerprinting, which can integrate meteorological disturbances, achieve real-time online monitoring, and quantitatively separate the contribution weights of each pollution source.
[0007] The technical solution adopted by this invention to solve its technical problem is: a method for analyzing pollution sources in drainage pipe networks based on deep learning and water quality fluorescence fingerprinting, comprising the following steps: S1. Collect basic physicochemical index data of water bodies in drainage pipe network in real time, and simultaneously acquire meteorological characteristic data; align the basic physicochemical index data and the meteorological characteristic data on the time scale to construct a standardized multimodal time series dataset; S2. Input the constructed multimodal time-series dataset into a dual-track long short-term memory network that integrates meteorological disturbance mechanisms to construct a dynamic evolution baseline for pipeline water quality, and deduce and output the dynamic baseline prediction values of each basic physicochemical index data; calculate the deviation between the actual monitored values of each basic physicochemical index data and the corresponding dynamic baseline prediction values at the same time to obtain a dynamic residual vector; input the dynamic residual vector into a pre-trained unsupervised anomaly detection and evaluation model to identify abnormal time periods; S3. Based on the identified abnormal time period, extract the basic physicochemical index data and the synchronous meteorological characteristic data within the abnormal time period and concatenate them into a cross-modal input vector; input the cross-modal input vector into a pre-trained cross-modal mapping model to deduce the predicted values of each characteristic parameter of fluorescence in the pipeline water body, and reconstruct the corresponding fluorescence fingerprint spectrum matrix based on the predicted values. S4. The reconstructed fluorescent fingerprint matrix and the predicted values of each feature parameter are used as input feature data and input to a pre-constructed redundant analysis architecture. The redundant analysis architecture adopts a primary and backup dual-mode computing mechanism, wherein the primary computing unit performs quantitative analysis based on a deep learning model, and the backup computing unit intervenes based on a statistical inference model under preset conditions. Finally, the contribution weight distribution of each pollution source to the current abnormal water body in the pipeline network is calculated and output.
[0008] Furthermore, the basic physicochemical index data include water temperature, pH value, electrical conductivity, total dissolved solids, chemical oxygen demand, total nitrogen, ammonia nitrogen, and total phosphorus; the meteorological characteristic data include air temperature, rainfall, and humidity.
[0009] Furthermore, the constructed multimodal time-series dataset is input into a dual-track long short-term memory network that integrates meteorological disturbance mechanisms to construct a baseline for the dynamic evolution of water quality in the pipeline network, and the dynamic baseline prediction values of various basic physicochemical indicators are derived and output, including: The basic physicochemical index data sequences in the multimodal time series dataset are positionally encoded to obtain encoded time series feature vectors; The temporal feature vector is input into a dual-track long short-term memory network for parallel computation, wherein: The first calculation trajectory is the meteorological sensitivity feature derivation trajectory: the time-series feature vector is input into the first long short-term memory network to extract the time-series state and calculate the basic meteorological sensitivity prediction value; at the same time, the meteorological feature data in the multimodal time-series dataset is mapped to the meteorological hidden state; the hidden state corresponding to the basic meteorological sensitivity prediction value is concatenated with the meteorological hidden state, and the nonlinear influence residual is calculated through a response gating network composed of a linear layer and a sigmoid activation function, and the first baseline prediction value after superimposing meteorological disturbances is output; The second calculation track is an independent derivation track for non-meteorologically sensitive features: the time-series feature vector is input into the second long short-term memory network, and after random deactivation, the second baseline prediction value, which is not affected by meteorological factors, is directly output. A preset expert mask matrix is introduced, and the first baseline prediction value and the second baseline prediction value are weighted and fused according to the meteorological influence attributes of each basic physicochemical index data, and the final dynamic baseline prediction vector is output, which is the dynamic baseline prediction value of each basic physicochemical index data.
[0010] Furthermore, the unsupervised anomaly detection and evaluation model is an isolated forest feature evaluation model; The step of inputting the dynamic residual vector into a pre-trained unsupervised anomaly detection and evaluation model to identify abnormal time periods includes: The dynamic residual vector is input into a pre-trained isolated forest feature evaluation model to calculate the multidimensional joint anomaly score, and the residuals of each basic physicochemical index data in the dynamic residual vector are subjected to Z-score standardization. When the multidimensional joint anomaly score is greater than the preset anomaly score threshold, and the residual standardized value of at least one basic physicochemical indicator data is greater than the preset positive deviation standard deviation threshold, the current time period is determined to be an abnormal time period.
[0011] Furthermore, the characteristic parameters of the fluorescence include the absolute intensity of specific fluorescent component peaks and the comprehensive fluorescence index characterizing the biochemical properties of the water body. Specifically, the specific fluorescent component peaks include fulvic acid-like peaks, humic acid-like peaks, tyrosine-like peaks, and tryptophan-like peaks. The comprehensive fluorescence index specifically includes the fluorescence index, the biological index, and the humification index.
[0012] Furthermore, the cross-modal mapping model is a multi-output regression model array based on the LightGBM algorithm; the multi-output regression model array contains multiple regressors corresponding to different fluorescence feature parameters, each regressor is trained independently and all use the LightGBM algorithm; The step of inputting the cross-modal input vector into a pre-trained cross-modal mapping model to deduce the predicted values of each feature parameter of water fluorescence, and reconstructing the corresponding fluorescence fingerprint matrix based on the predicted values, includes: The cross-modal input vector is standardized to obtain a standardized input vector; The standardized input vector is input in parallel into each regressor in the multi-output regression model array to deduce the predicted values of each fluorescence feature parameter; A two-dimensional spatial grid is set for the excitation wavelength and the emission wavelength. The theoretical excitation center and emission center coordinates of each specific fluorescence component peak are used as reference points. The corresponding absolute intensity prediction values are used as peak weights. The Gaussian kernel density distribution of each component peak is calculated in the two-dimensional spatial grid. The Gaussian kernel density distributions of all component peaks are superimposed to obtain the basic reconstruction matrix; Gaussian white noise is introduced into the basic reconstruction matrix to simulate background physical disturbances. The basic reconstruction matrix is then superimposed with the Gaussian white noise and subjected to non-negative truncation to generate the final fluorescence fingerprint matrix.
[0013] Furthermore, the deep learning model is a pre-trained convolutional neural network.
[0014] Furthermore, the pre-trained convolutional neural network is obtained in the following way: End-member water samples from each pollution source were collected, and mixed water samples were prepared according to a set volume ratio gradient. The true three-dimensional fluorescence fingerprint matrix was then measured. Using the real three-dimensional fluorescence fingerprint matrix as input samples and the known actual volume ratio of each pollution source as supervised learning labels, the convolutional neural network is subjected to parameter iteration and weight optimization to obtain the pre-trained convolutional neural network.
[0015] Furthermore, when the main computing unit performs quantitative analysis based on the pre-trained convolutional neural network, it adopts the following strategy: The reconstructed fluorescence fingerprint matrix is used as the spatial input feature and passed through five cascaded residual convolution modules to autonomously peel off and extract the deep global spatial features of various pollutants overlapping and merging under multi-source mixed conditions. Each residual convolution module includes convolution operation, batch normalization, linear rectification activation and skip connection, progressively expanding the number of channels. The first three modules are followed by a max pooling layer for dimensionality reduction. The acquired high-dimensional feature map is flattened into a one-dimensional vector and input into a multi-layer fully connected network containing a random deactivation mechanism. The probability is normalized by the Softmax function, and the initial contribution weights of various pollution sources are output end-to-end.
[0016] Furthermore, the statistical inference model is a Bayesian statistical inference model; The operating mechanism of the backup computing unit includes: When the noise level of the input feature data exceeds the preset noise threshold, or the confidence level of the output of the pre-trained convolutional neural network is lower than the preset confidence threshold, the parsing task switches to the Bayesian statistical inference model. The humification index, biological index, and intensity ratio of humic acid-like peak to fulvic acid-like peak are extracted from the predicted values of each characteristic parameter and used as the observation vector of the mixed water sample. By combining the prior endmember distribution matrix of the characteristic parameters corresponding to various pollution sources, a Bayesian mixture model based on fluorescence features is constructed. Multi-chain iterative sampling is performed using the Markov chain Monte Carlo algorithm to construct the posterior probability distribution; After the calculation results meet the set statistical convergence criteria, the expected posterior probability of each pollution source contribution weight component is solved and output as a backup source tracing weight.
[0017] A system for analyzing pollution sources in drainage pipe networks based on deep learning and water quality fluorescence fingerprinting is provided to implement the aforementioned method for analyzing pollution sources in drainage pipe networks based on deep learning and water quality fluorescence fingerprinting. The system includes: The data acquisition and processing module is configured to perform step S1; A baseline construction module is configured to perform step S2; The cross-modal mapping and graph reconstruction module is configured to perform step S3; and A redundant parsing module is configured to perform step S4.
[0018] A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the above-described method for analyzing pollution sources in drainage pipe networks based on deep learning and water quality fluorescence fingerprinting.
[0019] The beneficial effects of this invention are as follows: The method of this invention, by collecting basic physicochemical indicators and meteorological data in real time and constructing a multimodal time-series dataset, dynamically removes the influence of meteorological factors on water quality using a dual-track long short-term memory network that integrates meteorological disturbance mechanisms, achieving high-sensitivity identification of sudden pollution events; it does not rely on expensive three-dimensional fluorescence spectroscopy online equipment, but instead uses a cross-modal mapping model to infer and reconstruct fluorescence fingerprint spectra from conventional physicochemical indicators and meteorological data in real time, significantly reducing operation and maintenance costs; furthermore, it adopts a primary and backup dual-mode redundant analytical architecture, with the deep learning model outputting pollution source contribution weights end-to-end, and the statistical inference model intervening to provide fault-tolerant analysis when there is low confidence or high noise, thereby obtaining high-precision multi-source pollution contribution distribution under various operating conditions, overcoming the shortcomings of traditional methods such as delayed early warning, low source tracing accuracy, and high equipment costs. Attached Figure Description
[0020] Figure 1 This is a flowchart of the method of the present invention; Figure 2 This is a schematic diagram of the timing prediction results of the embodiment; Figure 3 This is a schematic diagram of the fluorescent fingerprint prediction results in the example; Figure 4 This is a schematic diagram showing the loss curve and test set accuracy curve of the convolutional neural network during the training process in an example embodiment. Detailed Implementation
[0021] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0022] like Figure 1 As shown, the method for analyzing pollution sources in drainage pipe networks based on deep learning and water quality fluorescence fingerprinting of the present invention includes the following steps: S1. Collect basic physicochemical index data of water bodies in drainage pipe network in real time, and simultaneously acquire meteorological characteristic data; align the basic physicochemical index data and the meteorological characteristic data on the time scale to eliminate the sampling frequency differences of multi-source heterogeneous data and construct a standardized multimodal time series dataset. S2. Input the constructed multimodal time-series dataset into a dual-track long short-term memory network that integrates meteorological disturbance mechanisms to construct a dynamic evolution baseline for pipeline water quality, and deduce and output the dynamic baseline prediction values of each basic physicochemical index data; calculate the deviation between the actual monitored values of each basic physicochemical index data and their corresponding dynamic baseline prediction values at the same time to obtain a dynamic residual vector; input the dynamic residual vector into a pre-trained unsupervised anomaly detection and evaluation model to identify abnormal time periods, wherein the abnormal time period is the current time that triggers the anomaly judgment and the time period formed by tracing back a preset time window length from that time as the endpoint; S3. Based on the identified abnormal time period, extract the basic physicochemical index data and the synchronous meteorological characteristic data within the abnormal time period and concatenate them into a cross-modal input vector; input the cross-modal input vector into a pre-trained cross-modal mapping model to deduce the predicted values of each characteristic parameter of fluorescence in the pipeline water body, and reconstruct the corresponding fluorescence fingerprint spectrum matrix based on the predicted values. S4. The reconstructed fluorescent fingerprint matrix and the predicted values of each feature parameter are used as input feature data and input to a pre-constructed redundant analysis architecture. The redundant analysis architecture adopts a primary and backup dual-mode computing mechanism, wherein the primary computing unit performs quantitative analysis based on a deep learning model, and the backup computing unit intervenes based on a statistical inference model under preset conditions. Finally, the contribution weight distribution of each pollution source to the current abnormal water body in the pipeline network is calculated and output.
[0023] Specifically, when the noise level of the input feature data exceeds a preset noise threshold, or when the confidence level output by the main computing unit is lower than a preset confidence threshold, the backup computing unit intervenes and takes over the current parsing task.
[0024] The above method achieves high-sensitivity identification of sudden pollution events by collecting basic physicochemical indicators and meteorological data in real time and constructing a multimodal time-series dataset. It utilizes a dual-track long short-term memory network that integrates meteorological disturbance mechanisms to dynamically remove the influence of meteorological factors on water quality. Instead of relying on expensive three-dimensional fluorescence spectroscopy online equipment, it uses a cross-modal mapping model to infer and reconstruct fluorescence fingerprints from conventional physicochemical indicators and meteorological data in real time, significantly reducing operation and maintenance costs. Furthermore, it adopts a primary and backup dual-mode redundant analysis architecture, with a deep learning model outputting pollution source contribution weights end-to-end, and a statistical inference model intervening to provide fault-tolerant analysis under low confidence or high noise conditions. This enables the acquisition of high-precision multi-source pollution contribution distributions under various operating conditions, overcoming the shortcomings of traditional methods such as delayed early warning, low source tracing accuracy, and high equipment costs.
[0025] In some embodiments, the basic physicochemical index data include water temperature, pH value, electrical conductivity (EC), total dissolved solids (TDS), chemical oxygen demand (COD), total nitrogen (TN), ammonia nitrogen (NH3-N), and total phosphorus (TP); the meteorological characteristic data include air temperature, rainfall, and humidity.
[0026] In some embodiments, the types of basic physicochemical index data can be selectively increased or decreased according to the actual monitoring needs of the target drainage network. For example, turbidity or dissolved oxygen indexes can be increased, or meteorological characteristic data such as wind speed and air pressure can be increased according to regional climate characteristics.
[0027] In some embodiments, time-scale alignment is performed using forward imputation and linear interpolation algorithms to construct a standardized multimodal time-series dataset. Cubic spline interpolation can be used for imputation when there are few missing data points; cubic interpolation can be used when there are many missing data points or the data is unevenly distributed. Furthermore, polynomial interpolation, Lagrange interpolation, or Kalman filtering can also be used to achieve time alignment of multimodal data.
[0028] In some embodiments, step S2 involves inputting the constructed multimodal time-series dataset into a dual-track long short-term memory (LSTM) network that integrates meteorological disturbance mechanisms to construct a baseline for the dynamic evolution of water quality in the pipeline network, and then extrapolating and outputting the dynamic baseline prediction values for each basic physicochemical index data. Specifically, this includes: (1) The basic physicochemical index data sequence in the multimodal time series dataset is positionally encoded to retain the absolute position information of the time series and obtain the encoded time series feature vector.
[0029] The time step is set to t, and the input hidden layer dimension is... The formula for calculating the elements of its position coding matrix is: ; ; In the formula, j is the feature dimension index. The basic physicochemical index data at time t are input into the vector. After performing a linear mapping, the result is added to the positional encoding to obtain the encoded temporal feature vector. : ; In the formula, and These are the weight matrix and bias vector of the initial linear projection layer, respectively. Subsequently, the data enters a dual-track long short-term memory network for parallel computation.
[0030] (2) Input the time-series feature vector into a dual-track long short-term memory network for parallel computation, wherein the first computation track is the meteorological sensitive feature derivation track; the second computation track is the non-meteorological sensitive feature independent derivation track.
[0031] First calculation track: The time-series feature vector Input the first Long Short-Term Memory network to extract the temporal state and obtain the hidden state at the end of the observation window. And calculate the meteorological sensitive basic forecast value. : In the formula, and These are the weight matrix and bias vector of the first basic fully connected layer.
[0032] Simultaneously, the meteorological feature data in the multimodal time-series dataset is mapped to a meteorological hidden state. .
[0033] The hidden state corresponding to the meteorological sensitive basic prediction value With the aforementioned weather concealment state Concatenate the vectors to obtain the context vector C=[ , ].
[0034] The gating coefficient G and the nonlinear effect residual are calculated using a response-gated network consisting of a linear layer and a sigmoid activation function. : ; ; In the formula, , , , and , , , Here are the weight matrix and bias vector for the corresponding linear layer; ReLU and sigma are the linear rectified and sigmoid nonlinear activation functions, respectively.
[0035] Output the first baseline forecast value after superimposing meteorological disturbances. : .
[0036] Second calculation track: The encoded time-series feature vector Input the second-layer Long Short-Term Memory network to obtain the hidden state at the end of the observation window. After random inactivation processing, the second baseline forecast value, unaffected by meteorological factors, is directly output. :
[0037] In the formula, and The weight matrix and bias vector are for the second basic fully connected layer.
[0038] (3) Introduce a pre-set expert mask matrix M , M The elements are determined based on the significance of the correlation coefficient between water quality indicators and meteorological factors. If the correlation coefficient p < 0.05, it is set to 1; otherwise, it is set to 0. The first baseline predicted value and the second baseline predicted value are weighted and fused according to the meteorological influence attributes of each basic physicochemical indicator data, outputting the final dynamic baseline prediction vector for the eight conventional basic physicochemical indicators. That is, the dynamic baseline predicted values of each basic physicochemical indicator data: .
[0039] The deviation between the actual monitored values of each basic physicochemical index at the same time and their corresponding dynamic baseline predicted values is calculated, resulting in a detailed explanation of the dynamic residual vector: The actual monitoring data of the pipeline water body at the current time t will be collected in real time to form an observation vector: ; The data for the basic physicochemical properties of the nth phase at the current time t; At the same time, the dynamic baseline prediction vector for the current moment is simultaneously derived by a dual-track long short-term memory network that integrates meteorological perturbation mechanisms: ; This represents the expected value of the nth basic physicochemical index under normal evolutionary conditions (considering meteorological disturbances).
[0040] For each basic physicochemical index data n, calculate the deviation between its actual monitored value and the corresponding dynamic baseline predicted value to obtain the dynamic residual: ; Then, all individual residuals are combined to form the dynamic residual vector at the current time t.
[0041] In some embodiments, the unsupervised anomaly detection evaluation model is an isolated forest feature evaluation model. In some embodiments, the unsupervised anomaly detection evaluation model may also employ a single-class support vector machine (OC-SVM), an autoencoder, or a clustering-based anomaly detection algorithm.
[0042] When the unsupervised anomaly detection evaluation model is an isolated forest feature evaluation model, in some embodiments, step S2, which involves inputting the dynamic residual vector into the pre-trained unsupervised anomaly detection evaluation model to identify abnormal time periods, specifically includes: The dynamic residual vector is input into a pre-trained isolated forest feature evaluation model to calculate the multidimensional joint anomaly score, and the residuals of each basic physicochemical index data in the dynamic residual vector are subjected to Z-score standardization. When the multidimensional joint anomaly score is greater than the preset anomaly score threshold, and the Z-score verification finds that the residual standardized value of at least one basic physicochemical indicator data is greater than the preset positive deviation standard deviation threshold, it is determined that the current pipeline network has encountered a real sudden pollution discharge, and the current time period is determined to be an abnormal time period.
[0043] Anomaly score thresholds are typically determined using quantile methods (e.g., set according to the expected anomaly rate of 1% to 5%), inflection point methods, or cost optimization methods. Common values for the positive deviation standard deviation threshold are 2.0, 2.5, or 3.0, and the choice should be made based on data distribution and business sensitivity. In some embodiments, the anomaly score threshold is set based on the 2% quantile of the expected anomaly rate using historical data from the training set; the positive deviation standard deviation threshold is fixed at 2.0 (one-sided), representing the criterion for determining significant positive deviation.
[0044] In some embodiments, step S2 further includes triggering an early warning signal when the current time period is determined to be an abnormal time period, in order to remind the inspection personnel.
[0045] In some embodiments, the fluorescence characteristic parameters include the absolute intensity of specific fluorescent component peaks and a comprehensive fluorescence index characterizing the biochemical properties of the water body. Specifically, the specific fluorescent component peaks include a fulvic acid-like peak (Peak A), a humic acid-like peak (Peak C), a tyrosine-like peak (Peak B), and a tryptophan-like peak (Peak T). The comprehensive fluorescence index specifically includes the fluorescence index (FI), the biological index (BIX), and the humification index (HIX). In some embodiments, the specific fluorescent component peaks may also include a marine humic acid-like peak (Peak M) or a fulvic acid-like peak (Peak D); the comprehensive fluorescence index may also include a freshness index (β:α) or a tyrosine / tryptophan ratio (Tyr / Trp). The composition of the characteristic parameters can be adaptively adjusted according to the main pollution source types within the target pipe network catchment area.
[0046] The cross-modal mapping model can employ algorithms such as random forest, XGBoost, support vector regression, or multilayer perceptron to establish a nonlinear mapping relationship between basic physicochemical index data, meteorological characteristic data, and fluorescence characteristic parameters. To reduce hardware costs and meet online real-time requirements while ensuring inference accuracy, the cross-modal mapping model is preferably a multi-output regression model array based on the LightGBM algorithm. This multi-output regression model array contains multiple regressors corresponding to different fluorescence characteristic parameters (such as absolute intensity of each component peak, fluorescence index, bioindices, humification index, etc.), and each regressor is trained independently using the LightGBM algorithm. This preferred scheme exhibits excellent performance in terms of inference accuracy, training speed, and memory usage, making it particularly suitable for online deployment scenarios involving multi-parameter parallel inference.
[0047] When the cross-modal mapping model is a multi-output regression model array based on the LightGBM algorithm, the step of inputting the cross-modal input vector into the pre-trained cross-modal mapping model, deducing the predicted values of each feature parameter of water fluorescence, and reconstructing the corresponding fluorescence fingerprint matrix based on the predicted values specifically includes: (1) For the cross-modal input vector After standardization, the standardized input vector is obtained: ; In the formula, Standardize the input vector; and These are the historical mean and standard deviation of the cross-modal vector of the input feature obtained during the model training phase.
[0048] (2) The standardized input vector is input in parallel into each regressor in the multi-output regression model array to deduce the predicted values of each fluorescence feature parameter: ; In the formula, q A feature index for a specific fluorescence characteristic parameter; The first result obtained from the deduction q Predicted values of fluorescence characteristic parameters; This sets the total number of regression trees in the LightGBM model. Let be the output mapping function of the d-th decision tree for the feature parameter q.
[0049] (3) Set a two-dimensional spatial grid for the excitation wavelength ex and the emission wavelength em to determine the wavelength range and resolution of the grid; use the theoretical excitation center and emission center coordinates of the peak c of each specific fluorescence component ( Using as a reference point, the absolute intensity of the peaks in this component obtained through deduction will be... I c (in I c ∈P q Using the corresponding Gaussian kernel function peak weights, the Gaussian kernel density distribution of the c-th component peak is calculated within the defined two-dimensional spatial grid. K c (ex,em) : ; In the formula, and The first c The theoretical excitation wavelength center and emission wavelength center of each component peak; and These are the Gaussian distribution broadening constants for the preset excitation and emission directions, respectively.
[0050] (4) Superimpose the Gaussian kernel density distributions of all component peaks to obtain the basic reconstruction matrix. : .
[0051] (5) Introduce Gaussian white noise into the basic reconstruction matrix. To simulate the background physical perturbations of a real spectroscopic instrument, the basic reconstruction matrix is superimposed with Gaussian white noise and then subjected to non-negative truncation to generate the final fluorescence fingerprint matrix. : ; In the formula, The variance of the introduced Gaussian white noise variable varies with The mean intensity is scaled proportionally; This is a function for maximizing values, used to ensure that the output fluorescence spectral intensity matrix does not contain any physically meaningless negative values.
[0052] In some embodiments, the deep learning model in step S4 is a pre-trained convolutional neural network; in other embodiments, the deep learning model is a pre-trained residual network or graph convolutional network, etc. The present invention preferably uses a convolutional neural network, which has advantages in feature extraction capability, computational efficiency, data adaptability, and system integration.
[0053] When the deep learning model is a pre-trained convolutional neural network, in order to establish an end-to-end mapping relationship between the fluorescence fingerprint matrix and the contribution weights of each pollution source, and to achieve accurate separation and quantification of mixed pollutants, the pre-trained convolutional neural network is preferably obtained in the following way: Collect data from various pollution sources (domestic sewage, kitchen wastewater, balcony wastewater, medical wastewater, chemical wastewater, electronic wastewater, food wastewater, and groundwater, etc.), assuming a total number of pollution source categories. N src In a laboratory setting, a large number of mixed water samples were prepared from end-member water samples according to a set gradient volume ratio, and their true three-dimensional fluorescence fingerprint matrix was measured. Using the real three-dimensional fluorescence fingerprint matrix as input samples and the known actual volume ratio of each pollution source as supervised learning labels, the convolutional neural network is subjected to parameter iteration and weight optimization to obtain the pre-trained convolutional neural network.
[0054] In some embodiments, when the main computing unit performs quantitative analysis based on the pre-trained convolutional neural network, it adopts the following strategy: the reconstructed fluorescence fingerprint matrix is used as the spatial input feature, and sequentially passed through five cascaded residual convolutional modules (ResNet blocks). Each residual module includes skip connections and batch normalization processing to alleviate the gradient vanishing problem in deep networks. After the last residual module, a global average pooling layer is used instead of the flattening operation to reduce the dimensionality of each feature channel to a scalar, which is then input into a single-layer fully connected network. The contribution weights of various pollution sources are output through the Softmax function. This strategy is suitable for scenarios with many types of pollution sources, highly complex spectral features, and sufficient training data.
[0055] In other embodiments, when the main computing unit performs quantitative analysis based on the pre-trained convolutional neural network, it employs the following strategy: keeping the weights of the pre-trained convolutional neural network unchanged, enabling the random deactivation mechanism (Dropout) in the fully connected layers during the inference phase, performing multiple (e.g., 20-50) forward propagations on the same input fluorescence spectrum matrix, recording the contribution weights of each pollution source in each output, calculating their mean as the final analysis result, and simultaneously calculating the standard deviation to characterize the prediction uncertainty. This strategy can improve the stability of the analysis without retraining the model, and is particularly suitable for online monitoring scenarios with high noise in single measurements.
[0056] In other embodiments, when the main computing unit performs quantitative analysis based on the pre-trained convolutional neural network, it employs the following strategy: The reconstructed fluorescence fingerprint matrix is used as spatial input features and sequentially passed through a three-level convolutional pooling module to autonomously peel off and extract deep global spatial features of various pollutants overlapping and fused under multi-source mixed conditions. The first-level module includes two-dimensional convolution operation, batch normalization processing, linear rectification activation and max pooling dimensionality reduction. The second and third-level modules have similar structures to the first level and gradually expand the number of channels. The acquired high-dimensional feature map is flattened into a one-dimensional vector and input into a multi-layer fully connected network containing a random deactivation mechanism. The probability is normalized by the Softmax function, and the initial contribution weights of various pollution sources are output end-to-end.
[0057] The above strategies provide an excellent starting point and practical baseline for multi-source pollution analysis tasks with moderate complexity, low deployment cost, and reliable generalization performance.
[0058] In some embodiments, the statistical inference model is a Bayesian statistical inference model; in other embodiments, the statistical inference model is any one of a variational inference model, an empirical Bayesian model, a frequentist model based on maximum likelihood estimation, or a resampling inference model based on bootstrapping. In this invention, the statistical inference model is preferably a Bayesian statistical inference model to compensate for the shortcomings of pure deep learning models in uncertainty quantification, prior knowledge fusion, and small-sample robustness, thereby improving the reliability of the overall redundant parsing architecture.
[0059] When the statistical inference model is a Bayesian statistical inference model, in some embodiments, the operation mechanism of the backup computing unit includes: when the noise level of the input feature data exceeds a preset noise threshold, or when the confidence level of the output of the pre-trained convolutional neural network is lower than a preset confidence threshold, the parsing task is switched to the Bayesian statistical inference model; the humification index, biological index, and the intensity ratio of humic acid-like peaks to fulvic acid-like peaks in the predicted values of each feature parameter are extracted as the observation vector of the mixed water sample; a Bayesian mixture model based on fluorescence features is constructed by combining the preset prior endmember distribution matrix of the feature parameters corresponding to various pollution sources; multi-chain iterative sampling is performed using the Markov chain Monte Carlo algorithm to construct the posterior probability distribution; after the calculation results meet the set statistical convergence benchmark, the posterior probability expectation value of the contribution weight components of each pollution source is solved and output as the backup source tracing weight.
[0060] Specifically, the preset confidence threshold is determined based on the output probability distribution of the pre-divided validation set samples in the deep learning model after training convergence: the maximum probability component of Softmax under known matching labels is statistically analyzed, and the 5th quantile of the corresponding output probability with correct parsing and error within ±5% is selected, with a value range of 0.75~0.85, and 0.80 is used in the example; the preset noise threshold is jointly established based on the standard deviation of water background fluctuations during historical periods without anomalies and the relative root mean square error of 5-fold cross-validation of the cross-modal mapping model, and twice the boundary of the mean of the cross-validation relative error (usually 8%~12%) is used as the high noise interception line, and 10% is used in the example.
[0061] This invention also provides a drainage network pollution source analysis system based on deep learning and water quality fluorescence fingerprinting, used to implement the above-mentioned drainage network pollution source analysis method based on deep learning and water quality fluorescence fingerprinting, including: The data acquisition and processing module is configured to perform step S1; A baseline construction module is configured to perform step S2; The cross-modal mapping and graph reconstruction module is configured to perform step S3; and A redundant parsing module is configured to perform step S4.
[0062] The present invention also provides a computer-readable storage medium storing a computer program thereon. When the computer program is executed by a processor, it implements the steps of the above-described method for analyzing pollution sources in drainage pipe networks based on deep learning and water quality fluorescence fingerprinting.
[0063] Example: (1) Obtain samples from 12,000 consecutive time steps in the historical monitoring data of the target drainage network as the baseline training set. To eliminate the influence of different modal data dimensions on gradient optimization, the eight basic physicochemical index data and three meteorological characteristic data are subjected to maximum and minimum normalization processing. The specific calculation formula is as follows:
[0064] In the formula, These are the raw monitoring values of basic physicochemical indicators or meteorological characteristic data; This represents the minimum value of the corresponding feature indicator in the historical monitoring dataset. This represents the maximum value of the corresponding feature indicator in the historical monitoring dataset; These are the standard feature values mapped to the interval [0, 1] after normalization.
[0065] After preprocessing, the standardized multimodal time-series dataset is divided into basic physicochemical index data sequences and meteorological characteristic data sequences. This embodiment employs a sliding window strategy, setting the time step window to 7 (i.e., using historical data from the past 7 time steps to predict water quality at the next moment). The basic physicochemical index data sequences undergo positional encoding. In this embodiment, the specific hyperparameter configuration of the dual-track long short-term memory network is as follows: input hidden layer dimension... The training batch size is 64, the number of LSTM cascaded layers is 2, the number of linear layer nodes in the response gating network is 32, the random inactivation rate is 0.2, the training batch size is 64, and the initial learning rate is 0.001.
[0066] The formula for calculating the elements of the position coding matrix is:
[0067]
[0068] In the formula, j Indexed by feature dimensions. t Input vector of basic physicochemical index data at time point After performing a linear mapping, the result is added to the positional encoding to obtain the encoded temporal feature vector. E t :
[0069] In the formula, and These are the weight matrix and bias vector of the initial linear projection layer, respectively. Subsequently, the data enters a dual-track long short-term memory network for parallel computation.
[0070] The first calculation orbit is the meteorological sensitivity feature extrapolation orbit. The encoded time-series feature vector... Input the first-layer Long Short-Term Memory network to extract the temporal state and obtain the hidden state at the end of the observation window. And calculate the meteorological sensitive basic forecast value. :
[0071] In the formula, and These are the weight matrix and bias vector of the first basic fully connected layer.
[0072] At the same time, the meteorological input vector at the end of the observation window X met The mechanism perception module maps the meteorological hidden state. The context vector C is obtained by concatenating the output states of the two entities. This concatenation operation is represented as C = [ , The gating coefficient G and the nonlinear effect residuals are calculated using a response-gated network consisting of a linear layer and a sigmoid activation function. :
[0073]
[0074] In the formula, , , , and , , , These are the weight matrix and bias vector for the corresponding linear layer; ReLU and sigma These are linear rectification and Sigmoid nonlinear activation functions, respectively. The result is the output of the first baseline prediction value after superimposed meteorological disturbances. Y 1 :
[0075] The second calculated orbit is an independently derived orbit based on non-meteorologically sensitive features. The encoded time-series feature vectors... Input the second-layer Long Short-Term Memory network to obtain the hidden state at the end of the observation window. H 2 After random inactivation processing, the second baseline forecast value, unaffected by meteorological factors, is directly output. Y 2 :
[0076] In the formula, and The weight matrix and bias vector are for the second basic fully connected layer.
[0077] (2) Introduce a pre-set expert mask matrix M This matrix assigns weighted variables based on the physicochemical properties of each water quality indicator affected by meteorological conditions, with each element taking a value of 0 or 1. The first and second baseline predicted values are then weighted and fused using this mask matrix to output the final dynamic baseline prediction vector for the eight basic physicochemical indicators. Y f :
[0078] During network training, the Huber loss function is used to reduce the interference of extreme water quality anomalies in historical data on the network gradient, and a cosine annealing strategy is configured to dynamically optimize the learning rate.
[0079] The model's convergence criteria are set as follows: a maximum of 150 training epochs are set, and an early stopping mechanism is introduced on the validation set. The model is considered convergent when the Huber loss value on the validation set no longer decreases for 15 consecutive training epochs, or when the fluctuation of the loss value in a single epoch is less than [a certain value]. When the model reaches convergence, training is prematurely terminated, and the current weights are saved as the final inference model. The time series prediction results are attached. Figure 2 .
[0080] (3) Calculate the deviation between the actual monitored value and the dynamic baseline predicted value and trigger a multi-dimensional joint anomaly early warning. The system collects the observation vector of the basic physicochemical index data at the current moment in real time. Y o and compared with the synchronously derived dynamic baseline prediction vector. Y f Perform the difference calculation to obtain the dynamic residual vector containing eight indicators. R :
[0081] In this embodiment, baseline extrapolation and subtraction are performed using the aforementioned dual-track long short-term memory network to obtain 12,000 sets of historical residual vectors for continuous time steps, constructing an offline training residual dataset. The mean and standard deviation of the residuals for each water quality indicator in this training residual dataset are calculated as the benchmark statistics for Z-score standardization during subsequent online evaluation. The number of decision trees in the base classifier of the isolated forest model is set to 200, and the threshold for abnormal data contamination rate is set to 0.05. The 12,000 sets of historical residual vectors are input into the isolated forest model for unsupervised fitting training. By constructing a random hyperplane to stagger the high-dimensional residual space, potential abnormal isolated points generate shorter average path lengths in the decision trees, thereby completing the integrated learning of spatial topological features and exporting and saving the trained model structure.
[0082] During online evaluation and inference, the dynamic residual vector calculated online at the current moment is used. R The input is fed into the trained Isolation Forest feature evaluation model. Based on the average path length of the current residual vector across 200 decision trees, the model calculates and outputs a multidimensional joint anomaly score. S IF Simultaneously, using the mean and standard deviation of the residuals of each water quality indicator obtained during the training phase, the residuals were analyzed. R k Perform Z-score standardization.
[0083] The joint early warning judgment logic set in this embodiment is as follows: when the anomaly score output by the isolated forest feature evaluation model is... S IF >0, and Z-score verification shows that at least one core water quality indicator's residual satisfies the condition. R k When the value is >2.0 (i.e., a positive deviation from the mean of more than two standard deviations), the system determines that the current pipeline network has encountered a real sudden pollution discharge and immediately triggers an abnormal discharge alarm command. If determined... S IF >0 but all core indicators meet the requirements. R k < If the value is 2.0, it is determined to be a natural dilution fluctuation caused by factors such as rainfall, and no pollution alarm will be triggered.
[0084] Step 3: Fluorescent fingerprint inference based on cross-modal mapping. After triggering the anomaly alarm in Step 2, the system extracts the basic physicochemical index data observations and synchronous meteorological characteristic data observations within the current anomaly period, and concatenates them into a cross-modal input vector. V in The vector is then input into a pre-trained cross-modal mapping model based on the LightGBM algorithm.
[0085] The training sample construction method and validation strategy of the cross-modal mapping model are as follows: (1) Extract the time nodes of the periodic manual water sample collection within the historical operation cycle, and use the eight basic physicochemical index data and three meteorological characteristic data collected by the online water quality sensor corresponding to the timestamp as the input feature vector; at the same time, use the core feature parameters of the three-dimensional fluorescence spectrum obtained by sending the water sample to the laboratory as the supervised learning label, so as to realize the cross-modal nonlinear pairing of conventional water quality index meteorological data and molecular-level fluorescence characteristic parameters.
[0086] In this embodiment, a total of 850 sets of synchronously paired cross-modal training samples were constructed. During the training phase, a 5-fold cross-validation strategy was adopted, in which the 850 sets of samples were randomly divided into 5 mutually exclusive subsets. Each time, 4 subsets were selected as the training set and the remaining subset was selected as the validation set. Five rounds of training and validation were performed, and the average performance of the 5 rounds of validation set was taken as the evaluation benchmark for hyperparameter tuning to prevent the model from overfitting.
[0087] (2) Construct a multi-output regression model array. First, standardize the input vectors:
[0088] In the formula, The input vector is the standardized form. and These are the historical mean and standard deviation of the cross-modal vector of the input feature obtained during the model training phase.
[0089] For the core feature parameters in the three-dimensional fluorescence fingerprint matrix, including the absolute intensities of multiple specific fluorescence component peaks (Peak A, Peak C, Peak B, Peak T) and multiple comprehensive fluorescence indices characterizing the biochemical properties of water bodies (FI, BIX, HIX), the corresponding LightGBM regressors are trained independently.
[0090] The target loss function for model training is the mean squared error loss function, which is calculated as follows:
[0091] In the formula, This represents the number of samples in the current training batch. These are the true values of the fluorescence characteristic parameters measured in the laboratory; These are the predicted values derived from the LightGBM regressor array. In this embodiment, each LightGBM model in the regressor array uses an independent weak decision tree iterative fitting method. The complete structural hyperparameter configuration specifically refers to the total number of decision trees in the LightGBM regressor. The system is configured with 150 trees, a learning rate of 0.05, a maximum weak tree depth of 6 layers, a maximum of 31 leaf nodes per tree, a minimum number of samples per leaf node of 10, a feature sampling ratio of 0.8, a sample sampling ratio of 0.9, and a random seed of 42. Using the aforementioned regressor array, based on the standardized input vector... In parallel, the predicted values of various core fluorescence characteristic parameters are derived. P q :
[0092] In the formula, qA feature index for a specific fluorescence characteristic parameter; The first result obtained from the deduction q Predicted values of fluorescence characteristic parameters; This sets the total number of regression trees in the LightGBM model. For the first d Decision trees for feature parameters q The output mapping function.
[0093] (3) Based on the predicted values of the derived fluorescence characteristic parameters, the fluorescence fingerprint matrix is reconstructed. Specifically, the excitation wavelength variable is set as... ex The emission wavelength variable is em This is used to define the spatial grid range of the two-dimensional map. This embodiment sets... ex and em The range is from 200.0 nanometers to 550.0 nanometers.
[0094] Based on the peaks of each specific fluorescent component c Theoretical excitation-launch center coordinates Using this as a benchmark, the absolute intensity of the peaks in this component obtained through deduction will be... I c (in I c ∈P q The peak weights of the corresponding Gaussian kernel function are used as the values. The Gaussian kernel density distribution of the c-th component peak is calculated within the defined spatial grid. K c (ex,em) In this embodiment, the preset excitation direction broadening constant and Set to 5.0:
[0095] In the formula, and The first c The theoretical excitation wavelength center and emission wavelength center of each component peak; and These are the Gaussian distribution broadening constants for the preset excitation and emission directions, respectively.
[0096] The Gaussian kernel density distributions of all extracted specific fluorescence component peaks are superimposed within a spatial grid to obtain the basic reconstruction matrix. :
[0097] (4) Introduce Gaussian white noise into the superimposed spatial grid. This simulates the background physical disturbances of a real spectroscopic instrument. In this embodiment, the mean of the Gaussian white noise is set to 0, and the standard deviation is set as the basic reconstruction matrix. The mean of all elements is 0.02 times. The basic reconstruction matrix is superimposed with Gaussian white noise, and a final high-fidelity two-dimensional fluorescence fingerprint matrix covering the entire band is generated by passing a non-negative truncation function. :
[0098] In the formula, The variance of the introduced Gaussian white noise variable varies with The mean intensity is scaled proportionally; This is a maximization function used to ensure that the output fluorescence spectral intensity matrix does not contain physically meaningless negative values. Prediction performance is shown in the appendix. Figure 3 .
[0099] Step 4: End-to-end quantitative source analysis based on deep feature extraction
[0100] The fluorescent fingerprint matrix generated in step three is reconstructed. and predicted values of fluorescence characteristic parameters The system uses feature data as input for source tracing analysis. Internally, it employs a redundant analysis architecture with a convolutional neural network as the main engine and a Bayesian mixture model as a backup algorithm.
[0101] For deep convolutional neural networks, which serve as the main parsing engine, the complete training and working mechanism is as follows: In the model training phase, the first step is to construct the basic dataset. Pollution sources in the pipe network are collected (domestic sewage, kitchen wastewater, balcony wastewater, medical wastewater, chemical wastewater, electronics manufacturing wastewater, food manufacturing wastewater, and furniture manufacturing wastewater, etc., assuming a total number of pollution source categories). N src In this embodiment N src =8) end-member water samples. Mixed water samples were prepared in a laboratory environment according to the following multi-element gradient volume ratio scheme: For cases where pollution sources are mixed in pairs, the number of combinations is calculated by randomly selecting any two types from the eight types of pollution sources. Two types of endmembers were paired according to volume ratios of 1:9, 2:8, 3:7, 4:6, 5:5, 6:4, 7:3, 8:2, and 9:1 to construct 252 dual-source hybrid configuration schemes. For mixed pollution sources (3 to 7 types), the proportioning method adopts a 10th-order simplex mixing uniform design table based on the deviation minimization criterion. For each specific combination, the spatially constrained domain is uniformly divided and 10 sets of the most spatially representative volume proportion vectors are extracted. The volume proportions of each participating source in each scheme are different and sum to 1.
[0102] The number of combinations of three sources, arbitrarily choosing any three types, is: There are 10 proportion schemes for each type of source, totaling 560 schemes. Examples of specific volume percentage gradients for each participating source are: [10%, 20%, 70%], [15%, 45%, 40%], [30%, 35%, 35%], [50%, 10%, 40%], [25%, 60%, 15%], [40%, 40%, 20%], [65%, 25%, 10%], [10%, 55%, 35%], [20%, 30%, 50%], [5%, 15%, 80%].
[0103] The number of combinations of four sources, arbitrarily choosing four types, is: There are 700 sets of 10 proportion schemes for each combination. Examples of specific volume percentage gradients for each participating source are: [10%, 10%, 20%, 60%], [15%, 25%, 40%, 20%], [30%, 20%, 25%, 25%], [45%, 15%, 10%, 30%], [5%, 40%, 50%, 5%], [20%, 35%, 15%, 30%], [40%, 10%, 30%, 20%], [10%, 50%, 15%, 25%], [25%, 25%, 25%, 25%], [50%, 20%, 10%, 20%].
[0104] The number of combinations of five elements, arbitrarily choosing five categories, is: There are 10 proportion schemes for each type of source, totaling 560 schemes. Examples of specific volume percentage gradients for each participating source are: [10%, 10%, 10%, 20%, 50%], [15%, 20%, 30%, 25%, 10%], [5%, 45%, 15%, 15%, 20%], [25%, 10%, 35%, 20%, 10%], [40%, 15%, 10%, 10%, 25%], [20%, 20%, 20%, 20%, 20%], [30%, 5%, 40%, 10%, 15%], [10%, 35%, 15%, 30%, 10%], [15%, 15%, 25%, 15%, 30%], [50%, 10%, 10%, 15%, 15%].
[0105] With six sources mixed, the number of combinations of any six categories is: There are 10 sets of proportion schemes for each type, totaling 280 sets. Examples of specific volume percentage gradients for each participating source are as follows: [10%, 10%, 10%, 10%, 20%, 40%], [15%, 20%, 15%, 25%, 15%, 10%], [5%, 35%, 15%, 15%, 15%, 15%], [20%, 10%, 30%, 15%, 15%, 10%], [30%, 15%, 10%, 10%, 20%, 15%], [16.6%, 16.6%, 16.6%, 16.6%, 16.6%, 17.0%], [25%, 5%, 35%, 10%, 15%, 10%], [10%, 30%, 10%, 25%, 15%, 10%], [15%, 15%, 20%, 15%] [20%, 15%], [40%, 10%, 10%, 15%, 15%, 10%].
[0106] The number of combinations of seven sources, arbitrarily choosing 7 types, is: There are 10 sets of proportion schemes for each type, totaling 80 sets. Examples of specific volume percentage gradients for each participating source are as follows: [10%, 10%, 10%, 10%, 10%, 20%, 30%], [15%, 15%, 20%, 15%, 15%, 10%, 10%], [5%, 25%, 15%, 15%, 15%, 15%, 10%], [20%, 10%, 25%, 10%, 15%, 10%, 10%], [25%, 10%, 10%, 10%, 20%, 15%, 10%], [14.2%, 14.2%, 14.2%, 14.2%, 14.2%, 14.2%, 14.2%, 14.8%], [20%, 5%, 30%, 10%, 15%, 10%, 10%], [10%, 25%, 10%] [20%, 15%, 10%, 10%], [15%, 15%, 15%, 15%, 15%, 15%, 10%], [30%, 10%, 10%, 10%, 15%, 15%, 10%].
[0107] In the case of mixing all pollution sources, where all eight types of pollution sources participate in the mixing simultaneously, the number of combinations is: The mixing method adopted an 18-order simplex mixing uniform design table. Within the seven-dimensional simplex manifold space, 18 representative volume ratio schemes covering different concentrations with overlapping areas were selected, resulting in a total of 18 schemes. Examples of specific volume percentage gradients for each participating source are as follows: [5%, 5%, 10%, 10%, 15%, 15%, 20%, 20%], [10%, 15%, 5%, 20%, 10%, 5%, 15%, 20%], [20%, 10%, 15%, 5%, 5%, 20%, 10%, 15%], [15%, 20%, 20%, 15%, 5%, 10%, 10%, 5%], [5%, 10%, 5%, 15%, 25%, 20%, 10%, 10%], [12.5%, 12.5%, 12.5%, 12.5%, 12.5%, 12.5%, 12.5%, 12.5%], [25%, 5%, 10%, 10%, 15%] [15%, 10%], [10%, 30%, 5%, 5%, 15%, 10%, 20%, 5%], [15%, 5%, 25%, 15%, 10%, 15%, 5%, 10%], [5%, 15%, 15%, 25%, 10%, 10%, 15%, 5%], [30%, 10%, 10%, 5%, 15%, 5%, 15%, 10%], [10%, 10%, 30%, 10%, 5%, 15%, 10%, 10%], [15%, 15%, 10%, 15%, 20%, 5%, 15%, 5%], [5%, 20%, 15%, 10%, 15%, 25%, 5%, 5%], [20%, 5%, 5%, 30%, [10%, 10%, 10%, 10%], [10%, 15%, 20%, 5%, 5%, 10%, 25%, 10%], [15%, 10%, 10%, 10%, 20%, 20%, 5%, 10%], [8%, 12%, 15%, 10%, 10%, 15%, 10%, 20%].
[0108] A total of 2,450 independent and unique volume ratio configuration schemes were generated. Six parallel samples were prepared independently for each scheme, resulting in a total of 14,700 mixed water samples. The true three-dimensional fluorescence fingerprint matrix of each sample was then measured.
[0109] These 14,700 three-dimensional fluorescent fingerprint matrices were used as input feature samples, and their corresponding known volume percentages were used as supervised learning labels to construct a complete traceability database. The 14,700 samples were randomly divided into three groups according to the proportions of 70%, 15%, and 15% of the sample size, with 10,290 samples assigned to the training set, 2,205 samples assigned to the validation set, and the remaining 2,205 samples assigned to the independent test set.
[0110] The specific training hyperparameters and overfitting control strategy configuration for the convolutional neural network are as follows: training batch size 64, basic optimizer type Adam optimizer, initial learning rate 0.0005, learning rate decay strategy cosine annealing dynamic adjustment mechanism, maximum number of training epochs 100, and regularization penalty coefficient. .
[0111] During model training, the optimizer minimizes the cross-entropy loss between the initial weights and the true labels. In addition to configuring a random deactivation mechanism in the fully connected layers, constraints are imposed on the high-dimensional parameter space through weight decay coefficients, and batch normalization layers are used to stabilize the data distribution of inter-layer feature maps in real time, thereby collaboratively suppressing overfitting during multi-level convolutional pooling iterations. This approach is used for parameter iteration and weight optimization of the network.
[0112] During the inference phase, the network receives the real-time fluorescence matrix and sequentially passes it through five cascaded residual convolutional modules to extract deep global spatial features. Each residual convolutional module includes two-dimensional convolution operations, batch normalization, linear rectified activation, and skip connections, progressively expanding the number of channels. Specifically, the first-level module has a kernel size of 5, expanding the feature channels to 32, followed by a max-pooling layer for dimensionality reduction. The second and third-level modules have kernel sizes of 3, progressively expanding the number of channels to 64 and 128 respectively, with each level followed by a max-pooling layer for dimensionality reduction. The fourth and fifth-level modules have kernel sizes of 3, further expanding the number of channels to 128, without pooling for dimensionality reduction, to preserve high-dimensional spatial features for subsequent analysis. Through these operations, the network can autonomously extract deep global spatial features from the overlapping and fusion of various pollutants in a multi-source mixed state. Finally, the acquired high-dimensional feature maps are flattened into one-dimensional vectors and input into a multi-layer fully connected network containing a random deactivation mechanism. The vectors are then normalized using the Softmax function, and the initial contribution weights for various pollution sources are output end-to-end. The accuracy of the convolutional neural network is shown in the appendix. Figure 4 After 50 iterations, the overall classification accuracy on the validation set reached 0.90, and after 100 iterations, the overall classification accuracy on the independent test set reached 0.91.
[0113] A Bayesian statistical inference model based on an endmember prior library is configured as a fault-tolerant backup algorithm. In practical applications, when the system evaluation module detects that the input matrix is severely affected by complex background noise in the pipeline network and the confidence of the convolutional neural network output is limited, the parsing task automatically or according to manual instructions switches to this backup engine.
[0114] In this mode, the humification index (HIX), biological index (BIX), and the ratio of the absolute intensity of the humic acid-like peak to the absolute intensity of the fulvic acid-like peak (C / A) are extracted from the core characteristic parameters derived in step three. These three independent parameters are then combined as the observation vector for the mixed water sample. Historical water samples from eight pollution sources within the catchment area were collected, each consisting of independent, unmixed, pure endmembers. Three-dimensional fluorescence spectra of 50 independent samples from each pollution source were measured, and the humification index, bioindices, and the intensity ratio of humic acid-like peaks to fulvic acid-like peaks were extracted and calculated for each sample. The mathematical statistical mean and standard deviation of each fluorescence parameter for each pollution source were calculated. Assuming that the characteristic parameters of each pollution source follow an independent normal distribution, a prior endmember distribution matrix was constructed, consisting of the expected values and variances of the three core fluorescence indicators for the eight pollution sources. A Bayesian mixture model based on fluorescence features was constructed. Multi-chain iterative sampling was performed using the Markov chain Monte Carlo algorithm. The specific sampling parameters of the Markov chain Monte Carlo algorithm were configured as follows: four parallel independent Markov chains, and a total of 20,000 iterations per Markov chain. To eliminate the nonlinear interference of the initial selection point on the posterior probability distribution, the first 5,000 iterations of each chain were directly discarded as the burning period length, retaining only the samples from the subsequent 15,000 valid iterations. The sampling step interval was set to 2, and the temporal autocorrelation of the sequence was reduced through interval sampling. Finally, 30,000 valid posterior sampling points from the four chains were merged to construct the posterior probability distribution. :
[0115] In the formula, Each pollution source contributes a weight vector to the solution; τ is a constant for the system's observation error tolerance; in this embodiment, τ is set to 0.1 (determined through cross-validation). is the prior distribution function, used to constrain the sum of the contribution weights of each pollution source to be 1 and all of them to be non-negative.
[0116] In this embodiment, the statistical convergence criterion is set as follows: the Gelman-Rubin test index is less than 1.05 and the result passes the Geweke test. Once the calculation results meet the above convergence criteria, the contribution weight components of each pollution source are calculated and output. The expected posterior probability is used as the backup source weight. :
[0117] In the formula, For mathematical expectation operators.
[0118] In one abnormal operating condition handling method of this embodiment, when both the deep learning model and the Bayesian statistical inference model fail analytically—that is, after the Markov chain Monte Carlo sampling reaches the preset maximum number of iterations, its Gelman-Rubin test index is still greater than or equal to 1.05, or it fails the Geweke test, and the Bayesian mixture model fails to meet the set statistical convergence benchmark—it indicates that the mixed water currently discharged into the drainage network contains unknown new pollution sources not covered by the eight known pollution sources, resulting in an irreconcilable nonlinear structural deviation between the actual observation vector Λ and the prior endmember distribution matrix Ψ constructed based on known endmembers.
[0119] Under this condition, the system automatically performs the following processing steps: (1) Suspend the current quantitative analysis output stream to prevent the output from having misleading erroneous contribution weights; (2) Collect mixed water samples during abnormal periods on-site for subsequent laboratory analysis with higher precision; (3) Based on the location of the abnormal pipe section indicated by the alarm and the subsequent analysis results, guide relevant management personnel to carry out targeted pollution source investigation along the pipeline network, and incorporate the characteristic information of the new pollution sources identified in the investigation into the source tracing database to enrich and iteratively update the pollution source database.
Claims
1. A method for sewer network pollution source analysis based on deep learning and water quality fluorescence fingerprinting, characterized in that, Includes the following steps: S1. Real-time acquisition of basic physicochemical index data of water bodies in drainage pipe networks, and simultaneous acquisition of meteorological characteristic data; The basic physicochemical index data and the meteorological characteristic data are aligned on the time scale to construct a standardized multimodal time series dataset; S2. Input the constructed multimodal time-series dataset into a dual-track long short-term memory network that integrates meteorological disturbance mechanisms to construct a dynamic evolution baseline for pipeline water quality, and deduce and output the dynamic baseline prediction values of each basic physicochemical index data; calculate the deviation between the actual monitored values of each basic physicochemical index data and the corresponding dynamic baseline prediction values at the same time to obtain a dynamic residual vector; input the dynamic residual vector into a pre-trained unsupervised anomaly detection and evaluation model to identify abnormal periods; S3. Based on the identified abnormal time period, extract the basic physicochemical index data and the synchronous meteorological characteristic data within the abnormal time period and concatenate them into a cross-modal input vector; input the cross-modal input vector into a pre-trained cross-modal mapping model to deduce the predicted values of each characteristic parameter of fluorescence in the pipeline water body, and reconstruct the corresponding fluorescence fingerprint spectrum matrix based on the predicted values. Specifically, the process involves inputting the constructed multimodal time-series dataset into a dual-track long short-term memory network that integrates meteorological disturbance mechanisms to build a baseline for the dynamic evolution of water quality in the pipeline network. This baseline is then used to deduce and output the predicted dynamic baseline values for each basic physicochemical indicator, including: The basic physicochemical index data sequences in the multimodal time series dataset are positionally encoded to obtain encoded time series feature vectors; The temporal feature vector is input into a dual-track long short-term memory network for parallel computation, wherein: The first calculation trajectory is the meteorological sensitivity feature derivation trajectory: the time-series feature vector is input into the first long short-term memory network to extract the time-series state and calculate the basic meteorological sensitivity prediction value; at the same time, the meteorological feature data in the multimodal time-series dataset is mapped to the meteorological hidden state; the hidden state corresponding to the basic meteorological sensitivity prediction value is concatenated with the meteorological hidden state, and the nonlinear influence residual is calculated through a response gating network composed of a linear layer and a sigmoid activation function, and the first baseline prediction value after superimposing meteorological disturbances is output; The second calculation track is an independent derivation track for non-meteorologically sensitive features: the time-series feature vector is input into the second long short-term memory network, and after random deactivation, the second baseline prediction value, which is not affected by meteorological factors, is directly output. A preset expert mask matrix is introduced, and the first baseline prediction value and the second baseline prediction value are weighted and fused according to the attributes of the basic physicochemical index data affected by meteorology, and the final dynamic baseline prediction vector is output, which is the dynamic baseline prediction value of each basic physicochemical index data. S4. The reconstructed fluorescent fingerprint matrix and the predicted values of each feature parameter are used as input feature data and input to a pre-constructed redundant analysis architecture. The redundant analysis architecture adopts a primary and backup dual-mode computing mechanism, wherein the primary computing unit performs quantitative analysis based on a deep learning model, and the backup computing unit intervenes based on a statistical inference model under preset conditions. Finally, the contribution weight distribution of each pollution source to the current abnormal water body in the pipeline network is calculated and output.
2. The sewer network pollution source analysis method based on deep learning and water quality fluorescence fingerprint according to claim 1, wherein, The basic physicochemical index data include water temperature, pH value, conductivity, total dissolved solids, chemical oxygen demand, total nitrogen, ammonia nitrogen, and total phosphorus; the meteorological characteristic data include air temperature, rainfall, and humidity.
3. The method for source apportionment of drainage pipe network pollution based on deep learning and water quality fluorescence fingerprinting as described in claim 1, characterized in that, The unsupervised anomaly detection and evaluation model is an isolated forest feature evaluation model; The step of inputting the dynamic residual vector into a pre-trained unsupervised anomaly detection and evaluation model to identify abnormal time periods includes: The dynamic residual vector is input into a pre-trained isolated forest feature evaluation model to calculate the multidimensional joint anomaly score, and the residuals of each basic physicochemical index data in the dynamic residual vector are subjected to Z-score standardization. When the multidimensional joint anomaly score is greater than the preset anomaly score threshold, and the residual standardized value of at least one basic physicochemical indicator data is greater than the preset positive deviation standard deviation threshold, the current time period is determined to be an abnormal time period.
4. The method for source apportionment of drainage pipe network pollution based on deep learning and water quality fluorescence fingerprinting as described in claim 1, characterized in that, The fluorescence characteristic parameters include the absolute intensity of specific fluorescent component peaks and the comprehensive fluorescence index characterizing the biochemical properties of the water body. Specifically, the specific fluorescent component peaks include fulvic acid-like peaks, humic acid-like peaks, tyrosine-like peaks, and tryptophan-like peaks. The comprehensive fluorescence index specifically includes the fluorescence index, the biological index, and the humification index.
5. The method for source apportionment of drainage pipe network pollution based on deep learning and water quality fluorescence fingerprinting as described in claim 1, characterized in that, The cross-modal mapping model is a multi-output regression model array based on the LightGBM algorithm; the multi-output regression model array contains multiple regressors corresponding to different fluorescence feature parameters, each regressor is trained independently and all use the LightGBM algorithm; The step of inputting the cross-modal input vector into a pre-trained cross-modal mapping model to deduce the predicted values of each feature parameter of water fluorescence, and reconstructing the corresponding fluorescence fingerprint matrix based on the predicted values, includes: The cross-modal input vector is standardized to obtain a standardized input vector; The standardized input vector is input in parallel into each regressor in the multi-output regression model array to deduce the predicted values of each fluorescence feature parameter; A two-dimensional spatial grid is set for the excitation wavelength and the emission wavelength. The theoretical excitation center and emission center coordinates of each specific fluorescence component peak are used as reference points. The corresponding absolute intensity prediction values are used as peak weights. The Gaussian kernel density distribution of each component peak is calculated in the two-dimensional spatial grid. The Gaussian kernel density distributions of all component peaks are superimposed to obtain the basic reconstruction matrix; Gaussian white noise is introduced into the basic reconstruction matrix to simulate background physical disturbances. The basic reconstruction matrix is then superimposed with the Gaussian white noise and subjected to non-negative truncation to generate the final fluorescence fingerprint matrix.
6. The method for source apportionment of drainage pipe network pollution based on deep learning and water quality fluorescence fingerprinting as described in claim 4, characterized in that, The deep learning model is a pre-trained convolutional neural network.
7. The method for source apportionment of drainage pipe network pollution based on deep learning and water quality fluorescence fingerprinting as described in claim 6, characterized in that, The pre-trained convolutional neural network is obtained through the following method: End-member water samples from each pollution source were collected, and mixed water samples were prepared according to a set volume ratio gradient. The true three-dimensional fluorescence fingerprint matrix was then measured. Using the real three-dimensional fluorescence fingerprint matrix as input samples and the known actual volume ratio of each pollution source as supervised learning labels, the convolutional neural network is subjected to parameter iteration and weight optimization to obtain the pre-trained convolutional neural network.
8. The method for source analysis of drainage pipe networks based on deep learning and water quality fluorescence fingerprinting as described in claim 6, characterized in that, When the main computing unit performs quantitative analysis based on the pre-trained convolutional neural network, it adopts the following strategy: The reconstructed fluorescence fingerprint matrix is used as the spatial input feature and passed through five levels of cascaded residual convolution modules to autonomously peel off and extract the deep global spatial features of various pollutants overlapping and fused under multi-source mixed conditions. Each level of residual convolution module includes convolution operation, batch normalization, linear rectified activation and skip connections, progressively expanding the number of channels. The first three modules are followed by a max pooling layer for dimensionality reduction. The acquired high-dimensional feature map is flattened into a one-dimensional vector and input into a multi-layer fully connected network containing a random deactivation mechanism. The probability is normalized by the Softmax function, and the initial contribution weights of various pollution sources are output end-to-end.
9. The method for source apportionment of drainage pipe network pollution based on deep learning and water quality fluorescence fingerprinting as described in claim 7 or 8, characterized in that, The statistical inference model is a Bayesian statistical inference model; The operating mechanism of the backup computing unit includes: When the noise level of the input feature data exceeds a preset noise threshold, or the confidence level of the output of the pre-trained convolutional neural network is lower than a preset confidence threshold, the parsing task switches to the Bayesian statistical inference model. The humification index, biological index, and intensity ratio of humic acid-like peak to fulvic acid-like peak are extracted from the predicted values of each characteristic parameter and used as the observation vector of the mixed water sample. By combining the prior endmember distribution matrix of the characteristic parameters corresponding to various pollution sources, a Bayesian mixture model based on fluorescence features is constructed. Multi-chain iterative sampling is performed using the Markov chain Monte Carlo algorithm to construct the posterior probability distribution; After the calculation results meet the set statistical convergence criteria, the posterior probability expectation of each pollution source contribution rate component is solved and output as a backup source tracing weight.
10. A drainage network pollution source apportionment system based on deep learning and water quality fluorescence fingerprinting, used to implement the drainage network pollution source apportionment method based on deep learning and water quality fluorescence fingerprinting as described in any one of claims 1 to 9, characterized in that, include: The data acquisition and processing module is configured to perform step S1; The baseline construction module is configured to perform step S2; The cross-modal mapping and graph reconstruction module is configured to perform step S3; and A redundant parsing module is configured to perform step S4.
11. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the drainage network pollution source analysis method based on deep learning and water quality fluorescence fingerprinting as described in any one of claims 1 to 9.
Citation Information
Patent Citations
Urban drainage pipe network pollution tracing method and device, electronic equipment and storage medium
CN119006250A
Drainage pipe network pollutant tracing method and system based on multi-level monitoring and medium
CN119475005A
Intelligent drainage pipe network pollution traceability alarm system based on invariant analysis
CN119494043A
Drainage pipe network pollutant tracing method, device, equipment, medium and program product
CN121207953A
Drainage pipe network pollution monitoring and advanced early warning system based on Internet of Things
CN121808418A