Water sample pollution tracing analysis method fusing absorption and three-dimensional fluorescence spectral characteristics
By integrating absorption and three-dimensional fluorescence spectral characteristics into a water sample pollution source tracing analysis method, the problem that existing spectral analysis methods cannot fully utilize multispectral modal information is solved, achieving high-precision prediction of pollutant indicators and source tracing of emissions, thus improving the practicality and intelligence of water quality monitoring.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- CHENGDU YIQINGYUAN TECH CO LTD
- Filing Date
- 2025-10-14
- Publication Date
- 2026-05-15
AI Technical Summary
Existing spectroscopic analysis methods rely on only a single spectral mode or two-dimensional fluorescence information, which cannot fully characterize the complex superposition effect of multiple pollutants in water bodies. Furthermore, the intrinsic correlation between different spectral modes is not fully utilized, resulting in limited accuracy in predicting pollution indicators and making it difficult to identify pollutant components and trace emission sources.
A water pollution source tracing analysis method integrating absorption and three-dimensional fluorescence spectral features is proposed. By acquiring absorption and three-dimensional fluorescence spectra, an initial spectral matrix is constructed and subjected to noise reduction and normalization. A spectral feature map is then constructed, trained using a graph matching network, and modeled using laboratory standard index values. Ultimately, this method enables the source tracing analysis of pollutant components and major emission sources.
It has achieved high-precision prediction of pollutant indicators and qualitative and quantitative tracing of pollutant components and potential emission sources, improving the practicality and intelligence of water quality monitoring systems and meeting the needs of systematic and real-time monitoring of complex water bodies.
Smart Images

Figure CN120992532B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of water sample spectral analysis for source tracing, specifically a method for water sample pollution source tracing analysis that integrates absorption-three-dimensional fluorescence spectral characteristics. Background Technology
[0002] Current water pollution monitoring technologies mainly rely on two types of methods: laboratory chemical analysis and spectroscopic measurement. Laboratory methods typically involve determining indicators such as chemical oxygen demand (COD), ammonia nitrogen, total phosphorus, and total nitrogen. These methods require collecting water samples, processing them with chemical reagents, or using specific instruments such as titration, spectrophotometry, and colorimetry to obtain water quality values. While these methods offer high accuracy, they are complex, time-consuming, and costly, and are difficult to implement for high-frequency or real-time monitoring, failing to meet the needs of continuous dynamic monitoring of urban water bodies or complex water areas. Spectroscopic measurement methods offer a rapid, non-destructive alternative. For example, ultraviolet-visible absorption spectroscopy can reflect the absorption characteristics of dissolved organic matter and inorganic ions in a water sample to different wavelengths of light, and single-wavelength fluorescence spectroscopy can be used to detect the fluorescence response of specific organic pollutants. These methods can acquire optical information about water bodies in a short time and be used to quickly estimate pollution indicators. However, single-wavelength or two-dimensional fluorescence spectroscopy can only capture the optical characteristics of some pollutants, making it difficult to simultaneously characterize the complex superposition effects of multiple pollutants in a water sample, and the completeness and multidimensional correlation of spectral information are not fully utilized. Especially when dealing with multiple organic pollutants, nitrogen and phosphorus pollutants, and their interactions in water bodies, the resolution and prediction accuracy of two-dimensional spectroscopy methods are limited.
[0003] Existing single-dimensional spectroscopic methods have significant limitations. First, single-wavelength or two-dimensional fluorescence spectroscopy cannot fully reflect the excitation-emission characteristics of different pollutants in water samples, resulting in insufficient pollutant feature extraction and affecting the accuracy of indicator prediction. Second, the inherent correlation between different spectral modes is not fully utilized, limiting the ability to identify pollutants and analyze emission sources. In addition, although laboratory chemical analysis is accurate, it is time-consuming and costly, making it difficult to meet the needs of rapid and continuous monitoring. Summary of the Invention
[0004] This invention proposes a water sample pollution source tracing analysis method that integrates absorption and three-dimensional fluorescence spectral features. It aims to solve the technical problems of existing spectral analysis methods that usually only use a single spectral mode or two-dimensional fluorescence information, which cannot fully characterize the complex superposition effect of multiple pollutants in water bodies, and the inherent correlation between different spectral modes is not fully utilized, resulting in limited accuracy of pollution index prediction and difficulty in identifying pollution components and tracing emission sources.
[0005] The water sample pollution source tracing analysis method that integrates absorption-three-dimensional fluorescence spectral characteristics includes the following steps:
[0006] S1. Obtain the absorption spectrum and three-dimensional fluorescence spectrum of the water sample to be tested. Construct an initial spectral matrix based on the absorption spectrum and three-dimensional fluorescence spectrum, and perform spectral feature preprocessing on the initial spectral matrix to obtain a denoised and normalized spectral feature matrix; wherein the three-dimensional fluorescence spectrum includes excitation spectrum and emission spectrum;
[0007] Specifically, for step S1, existing technologies typically utilize only absorption spectroscopy or single-dimensional fluorescence spectroscopy for water sample analysis. Since single-spectral methods can only capture limited chemical and optical information, overlapping signals between different pollutants are difficult to distinguish, leading to significant errors in indicator prediction. Step S1, however, simultaneously acquires absorption and three-dimensional fluorescence spectra (including excitation and emission spectra), fuses the two types of spectral data to construct an initial spectral matrix, and then performs denoising and normalization processing to obtain a unified spectral feature matrix. This process not only preserves the complementary information of different spectral channels regarding water quality components but also suppresses the interference of background noise and scale differences on subsequent analyses.
[0008] S2. Based on the spectral feature matrix, construct a spectral feature map, and map the feature points of the absorption spectrum and the three-dimensional fluorescence spectrum to nodes in the graph structure, with the spectral similarity between nodes as the edge weight;
[0009] Specifically, for step S2, existing technologies directly process the spectral matrix using linear feature extraction or principal component analysis (PCA). While this reduces dimensionality, it ignores the potential structural relationships and nonlinear correlations between different bands, making it difficult to reflect the coupling characteristics of complex pollutants in multispectral space. Step S2, however, maps feature points in the spectral feature matrix to nodes in a graph structure and uses the spectral similarity between nodes to define edge weights, constructing a spectral feature graph. This graph structure can simultaneously express the coupling relationship between absorption spectra and three-dimensional fluorescence spectra in the topological space, allowing both inter-spectral correlations and intra-spectral local structures to be explicitly represented, thus providing mineable association features for graph matching networks.
[0010] S3. The spectral feature map is trained using a machine learning model based on graph matching network, and the laboratory standard index values are used as supervision signals for modeling;
[0011] Specifically, existing spectral analysis modeling typically employs methods such as multiple linear regression, partial least squares regression, or deep neural networks. While these methods can fit the relationship between spectral features and laboratory indicators, they lack the ability to model the correspondence between different spectral modes and fail to fully utilize the complementarity between absorption spectra and three-dimensional fluorescence spectra. Step S3, however, trains the spectral feature map using a machine learning model based on a graph matching network, employing standard laboratory indicator values (including COD, ammonia nitrogen, total phosphorus, and total nitrogen) as supervisory signals to guide the model in learning node correspondences and feature space alignment across spectral modes during training. Through joint optimization of node embedding and the matching matrix, a nonlinear mapping between spectral features and indicator values can be accurately established.
[0012] S4. Input the spectral feature map of the water sample to be tested into the machine learning model based on graph matching network that has been trained, output the predicted value of the water sample index, and conduct source tracing analysis of pollutant components and main emission sources based on the index prediction results.
[0013] Specifically, for step S4, traditional water quality prediction methods typically only estimate the numerical values of indicators, failing to further explain the sources or tracing mechanisms of pollutants, thus lacking application value in environmental governance and emission liability determination. Step S4, however, inputs the spectral characteristic map of the water sample to a trained graph matching network, outputting predicted values for indicators such as COD, ammonia nitrogen, total phosphorus, and total nitrogen. It further combines these values with a database of pollutants and emission sources for comparative analysis, achieving automatic tracing of pollutants and major emission sources. Step S2, by matching the predicted indicator results with typical pollution source characteristic patterns, not only quantitatively provides the pollution level of the water sample but also qualitatively determines the types of major pollutants and their possible sources, significantly improving the practicality and intelligence of the water quality monitoring system.
[0014] The beneficial effects of the invention are:
[0015] This invention utilizes the multidimensional information of three-dimensional fluorescence and absorption spectra of water samples to achieve high-precision prediction of pollutant indicators. Based on the correlation analysis between the predicted indicators and the pollutant feature library and emission source database, it performs qualitative and quantitative source tracing of the main pollutants and potential emission sources in the water body, realizing the integration of water quality monitoring and pollution source management, and providing a systematic and real-time technical means for pollution analysis of complex water bodies. Attached Figure Description
[0016] Figure 1 This is a schematic diagram of the UAV target perception method based on beam reconstruction and environmental perception adaptation according to an embodiment of the present invention. Detailed Implementation
[0017] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings, but the scope of protection of the present invention is not limited to the following description.
[0018] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention; that is, the described embodiments are only a part of the embodiments of the invention, and not all of them. The components of the embodiments of the invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.
[0019] Therefore, the following detailed description of the embodiments of the invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the invention without inventive effort are within the scope of protection of the invention. It should be noted that relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations.
[0020] Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or machine that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or machine. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or machine that includes said element.
[0021] The features and performance of the present invention will be further described in detail below with reference to embodiments.
[0022] Example 1
[0023] Among them, such as Figure 1 A water pollution source tracing analysis method integrating absorption-three-dimensional fluorescence spectroscopy features includes the following steps:
[0024] S1. Obtain the absorption spectrum and three-dimensional fluorescence spectrum of the water sample to be tested. Construct an initial spectral matrix based on the absorption spectrum and three-dimensional fluorescence spectrum, and perform spectral feature preprocessing on the initial spectral matrix to obtain a denoised and normalized spectral feature matrix; wherein the three-dimensional fluorescence spectrum includes excitation spectrum and emission spectrum;
[0025] Specifically, existing technologies typically utilize only absorption spectra or single-dimensional fluorescence spectra for water sample analysis, making it difficult to capture the cumulative effects of multiple pollutants in complex water bodies, thus leading to significant errors in indicator prediction. Step S1 involves acquiring the absorption spectrum (the absorption intensity curve of the water sample to different wavelengths of incident light) and the three-dimensional fluorescence spectrum (a two-dimensional matrix composed of excitation and emission spectra, reflecting the fluorescence emission intensity distribution of the sample at different excitation wavelengths), fusing them to construct an initial spectral matrix (a two-dimensional matrix with wavelength as the coordinate axis and spectral intensity as the numerical value), and then performing denoising and normalization to obtain the spectral feature matrix. Denoising is used to eliminate environmental interference and instrument noise, while normalization is used to unify the intensity dimensions of different spectral channels.
[0026] S2. Based on the spectral feature matrix, construct a spectral feature map, and map the feature points of the absorption spectrum and the three-dimensional fluorescence spectrum to nodes in the graph structure, with the spectral similarity between nodes as the edge weight;
[0027] Specifically, traditional methods such as Principal Component Analysis (PCA) typically only perform linear dimensionality reduction, neglecting the nonlinear dependencies between different bands. Step S2 maps feature points in the spectral feature matrix to nodes in a graph structure (each node representing a specific band or spectral feature vector), and defines edge weights using spectral similarity (e.g., a function of cosine similarity or Euclidean distance) to construct a spectral feature map. This graph can simultaneously express the coupling relationship between absorption spectra and three-dimensional fluorescence spectra, preserving the detailed features of local bands while also characterizing the global correlations between different spectral channels, thus forming a topological input suitable for graph neural network processing.
[0028] S3. The spectral feature map is trained using a machine learning model based on graph matching network, and the laboratory standard index values are used as supervision signals for modeling;
[0029] Specifically, traditional spectral modeling methods such as multiple linear regression (MLR), partial least squares regression (PLSR), or simple deep neural networks often fail to model the correspondence between different spectral modes, making it difficult to fully utilize the complementarity between absorption spectra and three-dimensional fluorescence spectra. Step S3 uses a graph matching network (GMN) as a machine learning model to train the spectral feature map. This model can learn node correspondences (matching between absorption spectrum nodes and fluorescence spectrum nodes) and achieve alignment across modal feature spaces through joint optimization of node embedding and matching matrix.
[0030] S4. Input the spectral feature map of the water sample to be tested into the trained machine learning model based on graph matching network, output the predicted value of the water sample index, and conduct source tracing analysis of pollutant components and main emission sources based on the index prediction results.
[0031] Specifically, traditional water quality analysis is often limited to estimating the numerical values of indicators, and cannot further explain the pollutants and their sources. Step S4 inputs the spectral feature map of the water sample to be tested into the trained graph matching network to obtain the predicted values of the water sample indicators (COD, ammonia nitrogen, total phosphorus, and total nitrogen). Then, it is matched with the established pollutant database and the emission source database.
[0032] Furthermore, in step S1, the spectral feature preprocessing of the initial spectral matrix to obtain a denoised and normalized spectral matrix specifically includes the following sub-steps:
[0033] S101. Extract the main feature components by performing noise separation based on singular value decomposition on the initial spectral matrix;
[0034] S102. Normalize the spectral intensities of different dimensions to the same numerical range through normalization to obtain the normalized spectral matrix;
[0035] S103. The key band information in the normalized spectral matrix is weighted by a multi-scale feature enhancement method to obtain the spectral feature matrix.
[0036] Furthermore, step S101 specifically includes the following sub-steps:
[0037] S1011. For the acquired raw spectral matrix Perform singular value decomposition, where m represents the number of spectral sampling points and n represents the number of samples, to obtain:
[0038] ;
[0039] Among them, the Represents the original spectral matrix, the Let m represent the left singular vector matrix of SVD decomposition, and let m be the size of the matrix. Describes a diagonal singular value matrix of size m×n. Let n represent the transpose of the right singular vector matrix of the SVD decomposition, with size n×n.
[0040] S2012. Before retention The principal eigencomponents corresponding to the principal singular values, including baseline drift components, main absorption peak components, and main fluorescence emission peak components, yield the denoised spectral matrix:
[0041] ;
[0042] Among them, the Represents the denoised spectral matrix, the This represents a truncated version of the left singular vector matrix, retaining only the first few lines. The singular vectors, the This represents a truncated version of the diagonal singular value matrix, retaining only the first part. A singular value, the This represents the transpose and truncated version of the right singular vector matrix, retaining only the first part. The singular vectors, the This indicates the number of truncated singular values. Specifically, in singular value decomposition... In the singular value decomposition, singular values are arranged in order of magnitude. The first r singular values and their corresponding singular vectors are retained, while the rest are discarded. This scheme is used to reduce feature dimensionality, retaining only the main spectral modes and removing noise components corresponding to small singular values, making the spectral data smoother and more stable. In addition, the truncated rank r determines how many main spectral modes are retained in the singular value decomposition, serving as both a threshold for dimensionality reduction and denoising and a control parameter for constructing the feature vector dimension of nodes.
[0043] Specifically, the above implementation process is as follows: [The text abruptly ends here, likely due to an incomplete sentence or a formatting error.] Perform singular value decomposition And by selecting the truncated rank r according to the energy retention threshold, a low-rank approximation is obtained. .Will Each row serves as an r-dimensional feature vector for a wavelength / pixel node, which is then compared with another modality (three-dimensional fluorescence). The row vectors are used to calculate the cosine similarity between nodes, and this is used to construct the edge weight matrix of the graph, which serves as the input to the graph matching network.
[0044] Furthermore, the spectral feature map is obtained by wavelet basis expansion and singular value decomposition of the absorption spectrum and three-dimensional fluorescence spectrum, which is used to compress the spectral dimension, enhance the characterization ability of key bands, and serve as input to the graph matching network.
[0045] Further, the specific process of step S103 is as follows: On the normalized spectral matrix, spectral change features are extracted using the wavelet basis expansion method to enhance the characterization ability of key bands, and the multi-scale feature vectors are concatenated to form an enhanced spectral feature matrix. The key bands include the 254 nm absorption peak corresponding to COD, the characteristic absorption peak of ammonia nitrogen in the ultraviolet region, and the fluorescence excitation / emission peak pair corresponding to total phosphorus; wherein, the wavelet basis expansion method is specifically expressed as:
[0046] ;
[0047] Among them, the The continuous wavelet transform coefficients representing the i-th spectral feature, wherein Represents the wavelet scaling parameter, the Describing the wavelet translation parameters, the Represents the mother wavelet function, the Represents the elements of the normalized spectral matrix, the It means that the The index represents the spectral sampling point index, where i represents the spectral node index.
[0048] Specifically, the wavelet basis expansion method can simultaneously perform multi-scale analysis of the local and global features of spectral signals. In this scheme, the spectral matrix of the water sample after SVD denoising includes absorption spectra and three-dimensional fluorescence spectra. Different wavelength ranges in each spectrum have different sensitivities to pollutants, meaning there are key bands corresponding to characteristic peaks or shoulders of specific pollutants. The wavelet basis expansion formula is used to extract the spectral variation features. Furthermore, the wavelet expansion extracts multi-scale, local feature patterns from the SVD-denoised water sample spectral matrix. The i-th spectrum can be either an absorption spectrum or an excitation-emission slice of the three-dimensional fluorescence spectrum. The formula captures the broad-band absorption or emission trends of pollutants in the spectrum at different scales s, and the translation parameters are used to... Scanning the spectrum at different wavelengths identifies local peaks or subtle fluctuations; peaks correspond to characteristic absorption or fluorescence signals of specific pollutants. Normalized spectral intensity. As weights, they are used to ensure comparable contributions at different wavelengths while suppressing the impact of spectral amplitude differences on feature extraction. Wavelet basis function Used to match the local spectral morphology, so that the coefficient matrix generated by the formula... This is used to characterize local patterns and global trends in spectra. By constructing node features of the spectral feature map through a multi-scale local feature matrix, the graph matching network can distinguish the spectral features of different pollutants, enabling prediction and source tracing analysis of water sample pollution indicators based on absorption and fluorescence spectra.
[0049] Furthermore, as a preferred implementation, in addition to wavelet transform, the data can be reduced in dimensionality using methods such as PCA, LDA, KPCA, LLE, and Isomap, and the signal can be analyzed in the time and frequency domain using techniques such as Fourier transform, Hilbert transform, Gram angle field (GASF / GADF), and Markov migration field (MTF).
[0050] Furthermore, as a preferred implementation, an adaptive wavelet basis expansion method is proposed, the specific process of which is as follows:
[0051] Based on the energy distribution function of the normalized spectral matrix : At different wavelengths The energy peak distribution under the given conditions is dynamically selected to determine the optimal wavelet basis. For example, a wavelet basis that captures local variations in the high-energy band of the energy distribution function can be used as the optimal wavelet basis. Instead of fixed selection of Haar or Daubechies wavelets, this ensures that key pollution bands (254 nm for COD, 200–230 nm for ammonia nitrogen, and excitation / emission peak pairs for phosphorus) are enhanced and expanded.
[0052] Multi-scale wavelet expansion of the spectral matrix: The scale *s* is dynamically determined by the typical absorption bandwidth of the pollutant, and the translation parameter is... The corresponding spectral characteristic peaks are located at their center positions. Through multi-scale expansion, both rapid local changes (weak pollutant absorption peaks) are captured, while global trends are preserved.
[0053] The eigenvectors obtained after expansion Weighting coefficients are defined based on the differences in the spectral sensitivity of pollutants. : ; wherein, the Indicates wavelength at scale s The energy contribution is used to amplify the unfolding features of pollutant-sensitive bands and suppress the unfolding features of background noise bands through this weighting. The weighted feature vectors at different scales are then concatenated. ; Obtain the enhanced spectral feature matrix , where the symbol This indicates a splicing operation, where S represents a multi-scale set.
[0054] Furthermore, in step S2, mapping the feature points of the absorption spectrum and the three-dimensional fluorescence spectrum to nodes in a graph structure, with the spectral similarity between nodes serving as edge weights, specifically includes the following sub-steps:
[0055] S201. Take the absorption spectral feature points as node set A and the three-dimensional fluorescence spectral feature points as node set B; S202. Calculate the spectral similarity between any two nodes and use it as the edge weight to construct a weighted bipartite graph; S203. Represent the nodes as low-dimensional vectors using a graph embedding method.
[0056] Furthermore, in step S202, the specific calculation process for the edge weights is as follows: extract the feature vectors of node set A and node set B respectively, and calculate the edge weights using the cosine similarity function.
[0057] Furthermore, the laboratory standard index values include chemical oxygen demand (COD), ammonia nitrogen, total phosphorus, and total nitrogen. Specifically, these laboratory standard values are obtained using nationally or industry-standard water quality testing methods (such as the potassium dichromate method for COD determination, Nessler's reagent colorimetric method for ammonia nitrogen determination, molybdenum-antimony spectrophotometry for total phosphorus determination, and ultraviolet spectrophotometry for total nitrogen determination). These values serve as supervisory signals for the machine learning model during the training of the spectral feature map, providing a true benchmark for pollutant concentrations.
[0058] Furthermore, as a preferred embodiment, the laboratory standard index values may also include permanganate index (CODMn), BOD, TOC, NO3, NO2, PO4, turbidity, color, suspended solids, and petroleum hydrocarbons (solubility).
[0059] Furthermore, step S3 specifically includes the following sub-steps:
[0060] S301. Input the spectral feature map into the graph matching network, calculate the correspondence between absorption spectral nodes and fluorescence spectral nodes through the graph matching network, and output the node matching matrix;
[0061] S302. Based on the node matching matrix and the node information of the spectral feature map, the absorption spectrum and three-dimensional fluorescence spectrum nodes are uniformly mapped through the joint embedding function, and the joint embedding vector matrix is output.
[0062] S303. Match the joint embedding vector matrix with the laboratory standard index values, and establish the mapping relationship between the embedding vector and the water quality index through nonlinear regression;
[0063] S304. After training is completed, a combined prediction model of graph matching network and regression model is obtained.
[0064] Specifically, the obtained node matching matrix, combined with the node vector representation in the spectral feature map, maps feature points originally belonging to different spectral domains (absorption spectral domain and three-dimensional fluorescence spectral domain) into the same embedding space. Through a joint embedding function, the similarity relationship between cross-domain nodes is maintained in the low-dimensional space, allowing absorption spectral features and fluorescence spectral features to be complementary, thus obtaining a joint embedding vector matrix. Each row of this matrix corresponds to a cross-domain fused feature point. Known laboratory test index values (Chemical Oxygen Demand (COD), Ammonia Nitrogen, Total Phosphorus, Total Nitrogen, etc.) are used as supervision signals to model the joint embedding vector. By introducing a nonlinear regression model (such as a deep regression network, multilayer perceptron, or kernel regression), the complex nonlinear relationship between spectral features and water quality indicators is captured, achieving accurate mapping between the joint embedding vector and water quality indicator values. Furthermore, after training, an end-to-end combined prediction model is obtained, where the graph matching network is responsible for extracting and fusing cross-spectral features, while the nonlinear regression module is responsible for mapping the fused features to water quality indicators. This combined model can not only automatically adapt to the complex spectral features of different water samples but also stably output predicted values of key water quality indicators.
[0065] For example, the graph matching network is constructed through shallow graph convolutional layers and an attention weighting mechanism, and is used to calculate the correspondence between absorption spectral nodes and three-dimensional fluorescence spectral nodes and output a node matching matrix.
[0066] Furthermore, step S301 specifically includes the following sub-steps:
[0067] S3011. Encode the feature vectors of the absorption spectral nodes and the three-dimensional fluorescence spectral nodes respectively:
[0068] ;
[0069] ;
[0070] Among them, the The encoding vector representing absorption spectrum node i, the The absorption spectrum node feature encoding function, the The input vector represents the i-th feature point of the absorption spectrum, wherein This represents the total number of absorption spectral nodes. The encoding vector representing the three-dimensional fluorescence spectral node j, the The three-dimensional fluorescence spectral node feature encoding function, the The input vector representing the j-th feature point of the three-dimensional fluorescence spectrum, wherein This represents the total number of nodes in the three-dimensional fluorescence spectrum.
[0071] Specifically, in constructing spectral feature maps for absorption and 3D fluorescence spectra, each spectral feature point is mapped to a graph node. To enable the graph matching network to effectively process spectral information from different modalities, the original features of the nodes need to be encoded. Through feature encoding, the original features of absorption and 3D fluorescence nodes are mapped to a unified embedding space, achieving feature abstraction and alignment for different modal spectra. Furthermore, a nonlinear transformation is applied to each node using an encoding function, allowing the node embedding vector to simultaneously reflect the global trend and local peak patterns of the spectrum. This enhances the characterization of pollutant characteristic bands, providing comparability and distinguishability for subsequent node similarity calculations and graph matching. This enables the graph matching network to establish accurate node correspondences between multimodal spectra and capture the complex mixing characteristics of pollutants in water samples. The encoding function... and The nonlinear transformation is performed, which converts the original feature space of each node (including the spectral values and wavelet coefficients after SVD denoising) into a unified vector space, so that nodes of different modes can be directly similar in this space.
[0072] S3012. Calculate the similarity between absorption spectral nodes and fluorescence spectral nodes using the encoded feature vectors of the absorption spectral nodes and the three-dimensional fluorescence spectral nodes as the initial matching matrix:
[0073] ;
[0074] Among them, the This represents the node similarity after cosine similarity calculation;
[0075] S3013. Aggregate the neighbor node information in the spectral feature map and update the node embedding:
[0076] ;
[0077] ;
[0078] Among them, the This indicates the aggregation and embedding of absorption spectral nodes, the Represents the activation function, the The neighbor index of absorption spectrum node i is represented by the following. Denotes the set of neighbors of node i, the This represents the aggregation layer weight matrix, the This represents the node similarity between absorbing node i and its neighbor k. The encoding vector representing the absorption spectrum node k, the The aggregation layer bias vector is represented by the following. This indicates the aggregation and embedding of three-dimensional fluorescence spectral nodes, the The neighbor index of the three-dimensional fluorescence spectroscopy node j is represented by the following: Let j represent the set of neighbors of node j. Indicates neighbors The node similarity with node j in the three-dimensional fluorescence spectrum, the Indicates fluorescence spectral nodes The encoded vector;
[0079] Specifically, the above scheme is used to update the embedding vector of each spectral feature map node, fusing the node's own features with the feature information of its neighboring nodes to generate an embedding vector that contains both the node's own information and local spectral structure information. Specifically, by using a weighted summation of neighboring nodes, the feature information of locally similar nodes is incorporated into the current node, ensuring that each node's embedding reflects not only its own spectral features but also local spectral structure information. Furthermore, during the training of the graph matching network, and The updated embedding ensures that the features of absorption and fluorescence spectral nodes are comparable in the same embedding space, thereby supporting node matching and index prediction.
[0080] Furthermore, linear mapping and nonlinear functions It allows node embedding to capture nonlinear relationships between spectral features, enhances the ability to distinguish key bands, and makes the matching matrix reflect the relationships between corresponding nodes;
[0081] S3014. Calculate the final node matching matrix using the updated node embeddings: ;
[0082] Among them, the The elements in the node matching matrix are represented by the following: This represents the traversal variable in the index set of all three-dimensional fluorescence spectral nodes.
[0083] Specifically, in the above scheme, the calculation results are the matching probabilities of each element in the node matching, namely the matching probability of absorption spectral node i and three-dimensional fluorescence spectral node j. These need to be organized according to model row i and column j to obtain the complete node matching matrix. :
[0084] .
[0085] Furthermore, step S4 specifically includes the following sub-steps:
[0086] S401. Input the spectral feature map of the spectral matrix of the water sample to be tested into the trained machine learning model based on graph matching network, and output the corresponding node matching matrix of the water sample to be tested.
[0087] S402. By performing a unified mapping of the absorption spectrum and three-dimensional fluorescence spectrum nodes of the water sample to be tested through cross-modal joint embedding, the joint embedding vector matrix of the water sample to be tested is output.
[0088] S403. Input the joint embedding vector matrix of the water sample to be tested into the trained nonlinear regression model, and output the predicted water quality index value;
[0089] S404. Based on the predicted water quality index values, combined with the established pollutant component database and emission source database, analyze and output information on the distribution of pollutant components and major emission sources.
[0090] Specifically, the pollution component and emission source tracing analysis is obtained by performing similarity matching and cluster analysis between the predicted index results and the emission source database, which is used to identify the pollutant component types and infer the main emission sources; the trained machine learning model based on graph matching network, which is the combined prediction model of graph matching network and regression model mentioned in step S304 above, the graph matching network part is implemented based on shallow graph convolutional network or graph attention network, the nonlinear regression model, which is the combined prediction model of graph matching network and regression model, the regression model part is implemented based on deep regression network, multilayer perceptron or kernel regression, and the cross-modal joint embedding is the node embedding in step S3013.
[0091] Example 2
[0092] Furthermore, as a preferred embodiment of the above embodiments, an example of constructing the encoding function in step S3011 of the above embodiments is provided, wherein the encoding function in this embodiment... and Based on the normalization and weighted transformation of spectral features, a method is constructed for analyzing absorption spectral nodes. and fluorescence spectral nodes The original feature vectors are subjected to a nonlinear mapping. Specifically, it is defined as follows:
[0093] ;
[0094] ;
[0095] Among them, the and Indicates adjustment and The weight parameters of the feature contribution, the and Indicates adjustment and Weight parameters for the contribution of each feature.
[0096] The specific process is as follows:
[0097] The input feature vector of each node or The node input matrix of the spectral feature map is constructed by constructing a complete node vector composed of spectral values after SVD denoising and multi-scale local feature coefficients extracted by wavelet expansion.
[0098] Calculate the adjacency matrix for each node. and Neighboring nodes are weighted by spectral similarity, and adjacency information and node features are input into the graph convolutional layer;
[0099] In shallow graph convolution, the embedding of each node follows the formula Update; among them, , Let i represent the set of neighboring nodes. and These are trainable convolutional weights and biases. It uses a non-linear activation function, such as ReLU. One to two shallowly stacked layers complete feature encoding, outputting the final embedding vector. This vector preserves the local spectral features of nodes while integrating information from neighboring nodes, achieving cross-modal feature alignment. By training the weights of the graph convolutional layers through gradient descent, the embedding vector gradually reflects the similarity and difference of pollutant features in the water sample spectrum, thereby ensuring that spectral feature map nodes can effectively participate in indicator prediction and pollution source tracing analysis in the graph matching network.
[0100] Example 3
[0101] As a preferred embodiment of the above embodiments, step S404 involves an analysis process for tracing pollutant components and emission sources, specifically including the following steps:
[0102] S4041. Map the laboratory standard values of chemical oxygen demand (COD), ammonia nitrogen, total phosphorus, and total nitrogen to typical pollutant categories respectively:
[0103] COD: Organic pollutants (such as organic matter in industrial wastewater and kitchen wastewater);
[0104] Ammonia nitrogen (NH3-N): Nitrogen-containing pollutants (such as domestic sewage and aquaculture wastewater);
[0105] Total phosphorus (TP): eutrophic substances (such as agricultural non-point sources, detergents);
[0106] Total nitrogen (TN): nitrogen oxides and nitrogen-containing wastewater (such as fertilizer runoff and industrial waste liquid);
[0107] S4042. Based on the concentration levels of each indicator predicted from the spectral feature map, and combined with their variation patterns and the degree of exceedance, extract the characteristics of the pollutant components:
[0108] High COD indicates strong organic matter emissions, accompanied by high UV absorption characteristic peaks;
[0109] High ammonia nitrogen: This indicates the decomposition of proteins or nitrogen-containing compounds, with the fluorescence spectrum showing peaks resembling protein substances.
[0110] High total phosphorus: associated with algae and inorganic phosphates, with phosphate characteristic peaks appearing in the excitation-emission matrix;
[0111] High total nitrogen: associated with nitrates and ammonium salts, with enhanced absorption spectra in specific wavelength bands;
[0112] S4043. Using the established pollutant component-emission source database, match the extracted pollutant component characteristics with typical emission source types:
[0113] Industrial wastewater sources: elevated COD and TN levels; accompanied by complex organic spectral characteristics;
[0114] Domestic sewage source: COD and NH3-N increased; accompanied by protein fluorescence peaks;
[0115] Agricultural non-point source pollution: Increased TN and TP levels; manifested as nitrogen and phosphorus fertilizer runoff characteristics;
[0116] Aquaculture source: Increased NH3-N and TP; fluorescence peaks correspond to amino acids and phosphates;
[0117] S4044. By matching the pollutant components with the emission sources and combining the spatiotemporal distribution characteristics, output the main pollutant component types and corresponding emission source categories of the water sample to be tested.
[0118] The above description is merely a preferred embodiment of the present invention. It should be understood that the present invention is not limited to the forms disclosed herein and should not be construed as excluding other embodiments. It can be used in various other combinations, modifications, and environments, and can be altered within the scope of the concept described herein through the above teachings or related technologies or knowledge. Modifications and variations made by those skilled in the art that do not depart from the spirit and scope of the present invention should be within the protection scope of the appended claims.
Claims
1. A method for tracing the source of pollution in water samples using a fusion absorption-three-dimensional fluorescence spectral characteristic method, characterized in that, Includes the following steps: S1. Obtain the absorption spectrum and three-dimensional fluorescence spectrum of the water sample to be tested, construct an initial spectral matrix based on the absorption spectrum and three-dimensional fluorescence spectrum, and perform spectral feature preprocessing on the initial spectral matrix to obtain a denoised and normalized spectral feature matrix; The three-dimensional fluorescence spectrum includes excitation and emission spectra; S2. Based on the spectral feature matrix, construct a spectral feature map, and map the feature points of the absorption spectrum and the three-dimensional fluorescence spectrum to nodes in the graph structure, with the spectral similarity between nodes as the edge weight; S3. The spectral feature map is trained using a machine learning model based on graph matching network, and the laboratory standard index values are used as supervision signals for modeling; S4. Input the spectral feature map of the water sample to be tested into the trained machine learning model based on graph matching network, output the predicted value of the water sample index, and conduct source tracing analysis of pollutant components and main emission sources based on the index prediction results. Specifically, step S2 includes the following sub-steps: S201. Take the absorption spectral feature points as node set A, and the three-dimensional fluorescence spectral feature points as node set B; S202. Calculate the spectral similarity between any two nodes as the edge weight, and construct a weighted bipartite graph. The specific calculation process for the edge weight is as follows: extract the feature vectors of node set A and node set B respectively, and calculate the edge weight using the cosine similarity function. S203. Represent nodes as low-dimensional vectors using graph embedding methods; Step S3 specifically includes the following sub-steps: S301. Input the spectral feature map into the graph matching network, calculate the correspondence between absorption spectral nodes and fluorescence spectral nodes through the graph matching network, and output the node matching matrix; S302. Based on the node matching matrix and the node information of the spectral feature map, the absorption spectrum and three-dimensional fluorescence spectrum nodes are uniformly mapped through the joint embedding function, and the joint embedding vector matrix is output. S303. Match the joint embedding vector matrix with the laboratory standard index values, and establish the mapping relationship between the embedding vector and the water quality index through nonlinear regression; S304. After training is completed, a combined prediction model of graph matching network and regression model is obtained.
2. The water sample pollution source tracing analysis method based on fused absorption-three-dimensional fluorescence spectral characteristics as described in claim 1, characterized in that, In step S1, the spectral feature preprocessing of the initial spectral matrix to obtain a denoised and normalized spectral matrix specifically includes the following sub-steps: S101. Extract the main feature components by performing noise separation based on singular value decomposition on the initial spectral matrix; S102. Normalize the spectral intensities of different dimensions to the same numerical range through normalization to obtain the normalized spectral matrix; S103. The key band information in the normalized spectral matrix is weighted by a multi-scale feature enhancement method to obtain the spectral feature matrix.
3. The water sample pollution source tracing analysis method based on fused absorption-three-dimensional fluorescence spectral characteristics as described in claim 2, characterized in that, Step S101 specifically includes the following sub-steps: S1011. For the acquired raw spectral matrix Perform singular value decomposition, where m represents the number of spectral sampling points and n represents the number of samples, to obtain: ; Among them, the Represents the original spectral matrix, the Let m represent the left singular vector matrix of SVD decomposition, and let m be the size of the matrix. Describes a diagonal singular value matrix of size m×n. This represents the transpose of the right singular vector matrix of SVD decomposition, with size n×n; S1012. Before retention The principal eigencomponents corresponding to the principal singular values, including baseline drift components, main absorption peak components, and main fluorescence emission peak components, yield the denoised spectral matrix: ; Among them, the Represents the denoised spectral matrix, the This represents a truncated version of the left singular vector matrix, retaining only the first few lines. The singular vectors, the This represents a truncated version of the diagonal singular value matrix, retaining only the first part. A singular value, the This represents the transpose and truncated version of the right singular vector matrix, retaining only the first part. The singular vectors, the This indicates the number of singular values truncated.
4. The water sample pollution source tracing analysis method based on fused absorption-three-dimensional fluorescence spectral characteristics as described in claim 2, characterized in that, The specific process of step S103 is as follows: On the normalized spectral matrix, spectral change features are extracted using the wavelet basis expansion method to enhance the characterization ability of key bands, and the multi-scale feature vectors are concatenated to form an enhanced spectral feature matrix. The key bands include the 254 nm absorption peak corresponding to COD, the characteristic absorption peak of ammonia nitrogen in the ultraviolet region, and the fluorescence excitation / emission peak pair corresponding to total phosphorus; wherein, the wavelet basis expansion method is specifically expressed as follows: ; Among them, the The continuous wavelet transform coefficients representing the i-th spectral feature, wherein Represents the wavelet scaling parameter, the Describing the wavelet translation parameters, the Represents the mother wavelet function, the Represents the elements of the normalized spectral matrix, the The index represents the spectral sampling point index, where i represents the spectral node index.
5. The water sample pollution source tracing analysis method based on fused absorption-three-dimensional fluorescence spectral characteristics as described in claim 1, characterized in that, The laboratory standard index values include chemical oxygen demand (COD), ammonia nitrogen, total phosphorus, and total nitrogen.
6. The water sample pollution source tracing analysis method based on fused absorption-three-dimensional fluorescence spectral characteristics as described in claim 1, characterized in that, Step S301 specifically includes the following sub-steps: S3011. Encode the feature vectors of the absorption spectral nodes and the three-dimensional fluorescence spectral nodes respectively: ; ; Among them, the The encoding vector representing absorption spectrum node i, the The absorption spectrum node feature encoding function, the The input vector represents the i-th feature point of the absorption spectrum, wherein This represents the total number of absorption spectral nodes. The encoding vector representing the three-dimensional fluorescence spectral node j, the The three-dimensional fluorescence spectral node feature encoding function, the The input vector representing the j-th feature point of the three-dimensional fluorescence spectrum, wherein This represents the total number of nodes in the three-dimensional fluorescence spectrum. S3012. Calculate the similarity between absorption spectral nodes and fluorescence spectral nodes using the encoded feature vectors of the absorption spectral nodes and the three-dimensional fluorescence spectral nodes as the initial matching matrix: ; Among them, the This represents the node similarity after cosine similarity calculation; S3013. Aggregate the neighbor node information in the spectral feature map and update the node embedding: ; ; Among them, the This indicates the aggregation and embedding of absorption spectral nodes, the Represents the activation function, the The neighbor index of absorption spectrum node i is represented by the following. Denotes the set of neighbors of node i, the The aggregate layer weight matrix is represented by the following: This represents the node similarity between absorbing node i and its neighbor k. The encoding vector representing the absorption spectrum node k, the The aggregation layer bias vector is represented by the following. This indicates the aggregation and embedding of three-dimensional fluorescence spectral nodes, the The neighbor index of the three-dimensional fluorescence spectroscopy node j is represented by the following: Let j represent the set of neighbors of node j. Indicates neighbors The node similarity with node j in the three-dimensional fluorescence spectrum, the Indicates fluorescence spectral nodes The encoded vector; S3014. Calculate the final node matching matrix using the updated node embeddings: ; Among them, the The elements in the node matching matrix are represented by the following: This represents the traversal variable in the index set of all three-dimensional fluorescence spectral nodes.
7. The water sample pollution source tracing analysis method based on fused absorption-three-dimensional fluorescence spectral characteristics as described in claim 1, characterized in that, Step S4 specifically includes the following sub-steps: S401. Input the spectral feature map of the spectral matrix of the water sample to be tested into the trained machine learning model based on graph matching network, and output the corresponding node matching matrix of the water sample to be tested. S402. By using a cross-modal joint embedding function, the nodes of the absorption spectrum and three-dimensional fluorescence spectrum of the water sample to be tested are uniformly mapped, and the joint embedding vector matrix of the water sample to be tested is output. S403. Input the joint embedding vector matrix of the water sample to be tested into the trained nonlinear regression model, and output the predicted water quality index value; S404. Based on the predicted water quality index values, combined with the established spectral-component correlation model and emission source database, analyze and output the distribution of pollutant components and information on major emission sources.