A method for quantitatively identifying pollution sources of mixed water bodies by three-dimensional fluorescence spectrum combined with MixSIAR model

By combining three-dimensional fluorescence spectroscopy and the MixSIAR model, characteristic wavelength points are screened and combined, solving the problems of large data volume and high complexity in the identification of mixed water pollution sources, and achieving the effect of accurately tracing the source of mixed water pollution sources with a small amount of data.

CN118824390BActive Publication Date: 2026-02-03YANGTZE DELTA REGION INST OF TSINGHUA UNIV ZHEJIANG
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202410824773.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-25
Publication Date
2026-02-03
Estimated Expiration
2044-06-25

AI Technical Summary

Technical Problem

Existing technologies are insufficient to effectively identify multiple pollution sources in mixed water bodies, and existing methods involve large amounts of test data and high complexity.

Method used

Three-dimensional fluorescence spectroscopy combined with the MixSIAR model was used. Feature wavelength points were screened through a random forest model, and the feature wavelength points were combined and the pollution contribution rate was output using the MixSIAR model. The best feature wavelength points were selected for tracing the pollution sources of mixed water bodies.

Benefits of technology

With a small amount of test data, it can accurately identify pollution sources in mixed water bodies, and has high application and promotion prospects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118824390B_ABST
    Figure CN118824390B_ABST
Patent Text Reader

Abstract

The application provides a method for quantitatively identifying a pollution source of a mixed water body by combining three-dimensional fluorescence spectroscopy with a MixSIAR model, and the method comprises the following steps: screening characteristic wavelength points by using a random forest model; combining the characteristic wavelength points two by two, and further screening the characteristic wavelength points by using a pollution contribution rate of a single pollution source output by the MixSIAR model to obtain optimal characteristic wavelength points; and inputting data under the optimal characteristic wavelength points of the mixed water body into the MixSIAR model to determine the pollution source of the mixed water body by the contribution rate. According to the application, the data of the optimal characteristic wavelength points can be used to effectively screen the pollution source of the mixed water body, the pollution source of the mixed water body can be traced back more accurately under the condition that the amount of test data is small, and the application has a very high application and promotion prospect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of water pollution source tracing technology, and in particular to a method for quantitatively identifying mixed water pollution sources by combining three-dimensional fluorescence spectroscopy with a MixSIAR model. Background Technology

[0002] Water pollution source tracing is a crucial part of the "investigate, measure, trace, and treat" work model. Efficient and accurate identification of pollution sources is of significant practical importance for the precise control of water pollution, effectively managing pollutant discharge into rivers and ensuring surface water quality meets standards. Water pollution sources are complex, especially within industrial parks. In practical problem tracing, water pollution sources are not often single-source but rather a mixture of multiple pollution sources.

[0003] Patent application CN116595461A discloses a method for tracing the source of sewage discharge from rainwater outlets on sunny days based on random forest identification. This method acquires three-dimensional fluorescence spectral data of sewage samples within the area to be traced, inputs this data into a random forest model for training, and constructs a three-dimensional fluorescence identification model for the pollution source. The three-dimensional fluorescence spectral data of sewage samples from the sunny discharge outlets to be traced are then input into the identification model to obtain the tracing results. Patent application CN113033623A discloses a method and system for identifying pollution sources based on ultraviolet-visible absorption spectroscopy. This method acquires ultraviolet-visible absorption spectral data of pollution source samples, establishes and trains a pollution source identification model, and uses the trained model to identify the pollution source. Patent application CN113311081A discloses a method and device for identifying pollution sources based on three-dimensional liquid chromatography fingerprints. This method collects three-dimensional liquid chromatography fingerprints of pollution source samples and samples to be identified, establishes a pollution source identification model, and identifies the pollution source to which the sample to be identified belongs. The above methods can effectively identify single pollution sources, but they cannot identify pollution sources in mixed water bodies.

[0004] Currently, for the identification of pollution in mixed water bodies, only the invention patent application with announcement number CN115219472B discloses a method and system for quantitatively identifying multiple pollution sources in mixed water bodies. This method constructs a quantitative identification model by mixing the three-dimensional fluorescence spectral data of a single pollution source, thereby quantitatively identifying the sample to be tested. However, it can be seen from its application text that it is necessary to test the fluorescence data of 14,400 mixed pollution sources for just four types of pollution sources. The process is extremely complex and the amount of testing is large.

[0005] Therefore, it is necessary to provide a method for tracing the sources of pollution in mixed water bodies that requires less test data but still yields relatively accurate results, in order to solve the above-mentioned technical problems. Summary of the Invention

[0006] This invention provides a method for quantitatively identifying mixed water pollution sources using three-dimensional fluorescence spectroscopy combined with the MixSIAR model. This method can accurately trace the sources of mixed water pollution even with a small amount of test data, which is beneficial for its application and promotion.

[0007] The specific technical solution is as follows:

[0008] This invention provides a method for quantitatively identifying mixed water pollution sources using three-dimensional fluorescence spectroscopy combined with the MixSIAR model, comprising the following steps:

[0009] S1: Collect single-source wastewater samples from enterprises with different pollution sources around the mixed water body and obtain three-dimensional fluorescence spectral data of each water sample;

[0010] S2: The three-dimensional fluorescence spectral data of S1 are preprocessed. The preprocessed fluorescence spectral data and their corresponding pollution source types are used as inputs to construct a random forest model with wavelength points as variables. Several feature wavelength points with the highest importance in distinguishing the pollution source types of each sample are selected.

[0011] S3: Combine the fluorescence intensity data and pollution source type data corresponding to each pair of characteristic wavelength points to obtain the characteristic data group under each combination of characteristic wavelength points;

[0012] S4: Collect single-source drainage water samples from known pollution sources around the mixed water body, obtain three-dimensional fluorescence spectral data of several water samples, and screen the characteristic fluorescence data of each water sample under each characteristic wavelength point combination as several validation groups for MixSIAR model verification.

[0013] S5: Input the feature data set and the validation set under each combination of the same characteristic wavelength points into the MixSIAR model, output the contribution rate of each pollution source, and then verify the accuracy of the model output results according to the actual pollution source type of the validation set, so as to obtain the best combination of characteristic wavelength points with the highest accuracy.

[0014] S6: Collect samples of the mixed water body to be tested, obtain the three-dimensional fluorescence spectrum data of the mixed water body samples, and after the data is preprocessed, select the characteristic fluorescence data of the mixed water body samples under the optimal characteristic wavelength point combination, and input them together with the characteristic data group under the optimal characteristic wavelength point combination into the MixSIAR model to output the different contribution rates of various pollution sources in the mixed water body samples to be tested, and determine the pollution sources of the mixed water body by the contribution rate.

[0015] The "single-source wastewater sample" mentioned above refers to wastewater samples from a single pollution source enterprise; "wavelength point" refers to the data point corresponding to a certain combination of excitation and emission wavelengths, for example: excitation wavelength (Ex) = 250nm, emission wavelength (Em) = 455nm, the wavelength point is Ex / Em = 250 / 455nm; "characteristic data set" refers to characteristic wavelength points and their corresponding fluorescence intensity data, pollution source type data; "characteristic fluorescence data" refers to the fluorescence intensity data corresponding to characteristic wavelength points; "validation set" refers to the characteristic fluorescence data of wastewater samples from a single pollution source enterprise that have been re-collected under various combinations of characteristic wavelength points; "optimal combination of characteristic wavelength points" refers to the combination of characteristic wavelength points obtained after validation by the validation set. "Fluorescence spectral data": matrix data composed of each wavelength point and its corresponding fluorescence intensity data.

[0016] Furthermore, in S1, the mixed water body includes, but is not limited to, rainwater inlet drainage samples, sewage inlet drainage samples, surface water, and groundwater.

[0017] Furthermore, in S1, the number of pollution source types is 5 or less (including 5); the method for collecting single-source drainage water samples is as follows: during the enterprise's production process, water samples are collected every 1 to 2 hours before the enterprise discharges into the pipeline network, for a total of 15 times; mixed water samples are collected every 5 minutes, for a total of 3 times.

[0018] Furthermore, in S1, the instrument parameters for acquiring the three-dimensional fluorescence spectrum data are as follows: 150W xenon lamp, scanning speed of 12000nm / min, scanning range of excitation wavelength of 220-450nm, scanning interval of 1-5nm, scanning range of emission wavelength of 260-600nm, scanning interval of 1-10nm; and slit width of 5nm for both excitation and emission wavelengths.

[0019] Further, in S2, the preprocessing method includes:

[0020] (S2-1) Delaunay interpolation combined with Gaussian function is used for smooth interpolation to remove Rayleigh scattering and Raman scattering;

[0021] (S2-2) Expand the three-dimensional fluorescence spectral data along the direction of the excitation wavelength to form a one-dimensional vector form in which the data points of adjacent rows are connected end to end;

[0022] (S2-3) Normalize the three-dimensional fluorescence intensity data to obtain normalized fluorescence intensity data with values ​​of 0 to 1.

[0023] Furthermore, in S2, the random forest model uses decision trees as the base classifier, and optimizes the model parameters using fluorescence spectral data and their corresponding pollution source types. The optimized parameters include n_estimators, max_features, min_sample_leaf, and min_samples_split.

[0024] Furthermore, in S2, the top ten characteristic wavelength points of importance in distinguishing the pollution source types of each sample were selected.

[0025] Furthermore, before combining S3, the characteristic wavelength points are first screened for inconsistencies. The screening principle is as follows:

[0026] (S3-1) In the data obtained in S6, if the fluorescence characteristic value corresponding to a certain characteristic wavelength point of a certain water sample exceeds the range of fluorescence characteristic values ​​corresponding to the characteristic wavelength point formed by the data obtained in S1, then delete the characteristic wavelength point; otherwise, retain the characteristic wavelength point.

[0027] (S3-2) If there is no significant difference in the fluorescence characteristic values ​​of different pollution source enterprises corresponding to a certain characteristic wavelength point in all the data obtained in S1, then delete the characteristic wavelength point; otherwise, retain the characteristic wavelength point.

[0028] Furthermore, in S3, the MixSIAR model is based on the Bayesian probability method, using j characteristic fluorescence data to quantify the contribution of k pollution sources to the receptor. The model sets the contribution rate of each source to 1, uses the Dirichlet distribution as the prior distribution of pollution source contributions, and performs iterative calculations using a Monte Carlo Markov chain in the JAGS function. Based on the final output, the pollution contribution rate of each pollution source in the validation group sample is calculated.

[0029]

[0030]

[0031] In the formula, i represents the sample number of the verification group; j represents the characteristic wavelength point number; k represents the enterprise type number; X ij S is the fluorescence intensity value at the j-th characteristic wavelength point of the i-th validation group sample. jk p is the fluorescence intensity value at the j-th characteristic wavelength point of the pollution source enterprise k. k This represents the contribution rate of polluting enterprise k to the validation group sample, with a mean value of μ. jk The standard deviation is ω jk ε ij It is the residual, with a mean of zero and a standard deviation of σ. j N represents the Nth sample.

[0032] Furthermore, the data files input into the MixSIAR model include: the Mixture file, the Sources file, and the Discrimination file; among them, the Mixture file contains the validation group data, the Sources file contains the feature data group, and all data in the Discrimination file are set to 0. The format of all three data files is CSV; the Markov chain Monte Carlo run length is set to "normal"; and the sum of the pollution contribution rates of each pollution source in the validation group is 1.

[0033] Furthermore, in S5, the method with the highest accuracy is the one where the output pollutant type result is the same as the actual pollutant type result of the validation group, and the pollution contribution rate is the highest.

[0034] Compared with the prior art, the present invention has the following beneficial effects:

[0035] This invention uses a random forest model to screen characteristic wavelength points, then combines the characteristic wavelength points in pairs, and uses the pollution contribution rate output by the MixSIAR model to further screen the characteristic wavelength points to obtain the optimal characteristic wavelength points. This invention can effectively screen out the pollution sources of mixed water bodies using only the data of the optimal characteristic wavelength points. It can accurately trace the pollution sources of mixed water bodies with a relatively small amount of test data, and has a very high application and promotion prospect. Attached Figure Description

[0036] Figure 1 This is a flowchart illustrating the method for quantitatively identifying pollution sources in mixed water bodies in Example 1.

[0037] Figure 2 This is a distribution map of the importance of fluorescence features output by the random forest model in Example 1;

[0038] The horizontal axis represents the importance of the wavelength points; the vertical axis represents the wavelength point numbers, which are assigned as 1, 2, 3, ... Detailed Implementation

[0039] The present invention will be further described below with reference to specific embodiments. The following are only specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto.

[0040] Example 1

[0041] This invention provides a method for quantitatively identifying mixed water pollution sources using three-dimensional fluorescence spectroscopy combined with the MixSIAR model, specifically including the following steps:

[0042] S1: Collect single-source wastewater samples from enterprises with different pollution sources around the mixed water body and obtain three-dimensional fluorescence spectral data of each water sample;

[0043] The number of pollution source types is 5 or less (including 5); the single-source drainage water sample collection method is as follows: during the enterprise's production process, water samples are collected every 1 to 2 hours before the enterprise discharges into the pipeline network, for a total of 15 times; mixed water samples are collected every 5 minutes, for a total of 3 times.

[0044] The instrument parameters for acquiring three-dimensional fluorescence spectral data are as follows: 150W xenon lamp, scanning speed of 12000nm / min, scanning range of excitation wavelength of 220-450nm, scanning interval of 1-5nm, scanning range of emission wavelength of 260-600nm, scanning interval of 1-10nm; the slit width of both excitation and emission wavelengths is 5nm.

[0045] S2: The three-dimensional fluorescence spectral data of S1 are preprocessed. The preprocessed fluorescence spectral data and their corresponding pollution source types are used as inputs to construct a random forest model with wavelength points as variables. Several feature wavelength points with the highest importance in distinguishing the pollution source types of each sample are selected.

[0046] Preprocessing methods include:

[0047] (S2-1) Delaunay interpolation combined with Gaussian function is used for smooth interpolation to remove Rayleigh scattering and Raman scattering;

[0048] (S2-2) Expand the three-dimensional fluorescence spectral data along the direction of the excitation wavelength to form a one-dimensional vector form in which the data points of adjacent rows are connected end to end;

[0049] (S2-3) Normalize the three-dimensional fluorescence spectral data to obtain normalized fluorescence intensity data with values ​​of 0 to 1.

[0050] The random forest model uses decision trees as the base classifier. The model parameters are optimized using fluorescence spectral data and their corresponding pollution source types. The optimized parameters include n_estimators, max_features, min_sample_leaf, and min_samples_split.

[0051] S3: Combine the fluorescence intensity data and pollution source type data corresponding to each pair of characteristic wavelength points to obtain the characteristic data group under each combination of characteristic wavelength points;

[0052] Before combining S3, the characteristic wavelength points are first screened for inconsistencies. The screening principle is as follows:

[0053] (S3-1) In the data obtained in S6, if the fluorescence characteristic value corresponding to a certain characteristic wavelength point of a certain water sample exceeds the range of fluorescence characteristic values ​​corresponding to the characteristic wavelength point formed by the data obtained in S1, then delete the characteristic wavelength point; otherwise, retain the characteristic wavelength point.

[0054] (S3-2) If there is no significant difference in the fluorescence characteristic values ​​of different pollution source enterprises corresponding to a certain characteristic wavelength point in all the data obtained in S1, then delete the characteristic wavelength point; otherwise, retain the characteristic wavelength point.

[0055] S4: Collect single-source drainage water samples from known pollution sources around the mixed water body, obtain three-dimensional fluorescence spectral data of several water samples, and screen the characteristic fluorescence data of each water sample under each characteristic wavelength point combination as several validation groups for MixSIAR model verification.

[0056] The MixSIAR model, based on Bayesian probabilistic methods, uses j feature fluorescence data to quantify the contribution of k pollution sources to the receptor. The model sets the contribution rate of each source to 1, uses a Dirichlet distribution as the prior distribution of pollution source contributions, and performs iterative calculations using a Monte Carlo Markov chain in the JAGS function. Based on the final output, the pollution contribution rate of each pollution source in the validation group sample is calculated.

[0057]

[0058] In the formula, i represents the sample number of the verification group; j represents the characteristic wavelength point number; k represents the enterprise type number; X ij S is the fluorescence intensity value at the j-th characteristic wavelength point of the i-th validation group sample. jk p is the fluorescence intensity value at the j-th characteristic wavelength point of the pollution source enterprise k. k This represents the contribution rate of polluting enterprise k to the validation group sample, with a mean value of μ. jk The standard deviation is ω jk ε ij It is the residual, with a mean of zero and a standard deviation of σ. j N represents the Nth sample.

[0059] The data files input into the MixSIAR model include: Mixture file, Sources file, and Discrimination file; among them, the Mixture file contains the validation group data, the Sources file contains the feature data group, and all data in the Discrimination file are set to 0. The format of all three data files is CSV; the Markov chain Monte Carlo run length is set to "normal"; the sum of the pollution contribution rates of each pollution source in the validation group is 1.

[0060] S5: Input the feature data set and the validation set under each combination of the same characteristic wavelength points into the MixSIAR model, output the contribution rate of each pollution source, and then verify the accuracy of the model output results according to the actual pollution source type of the validation set, so as to obtain the best combination of characteristic wavelength points with the highest accuracy.

[0061] The method with the highest accuracy is the one where the output pollutant type result is the same as the actual pollutant type result of the validation group, and the pollution contribution rate is the highest.

[0062] S6: Collect samples of the mixed water body to be tested, obtain three-dimensional fluorescence spectral data of the mixed water body samples, select the characteristic fluorescence data of the mixed water body samples under the optimal characteristic wavelength point combination after data preprocessing, and input them together with the characteristic data group under the optimal characteristic wavelength point combination into the MixSIAR model, output the different contribution rates of various pollution sources in the mixed water body sample to be tested, and determine the pollution sources of the mixed water body by the contribution rate.

[0063] The "single-source wastewater sample" mentioned above refers to wastewater samples from a single pollution source enterprise; "wavelength point" refers to the data point corresponding to a certain combination of excitation and emission wavelengths, for example: excitation wavelength (Ex) = 250nm, emission wavelength (Em) = 455nm, the wavelength point is Ex / Em = 250 / 455nm; "characteristic data set" refers to characteristic wavelength points and their corresponding fluorescence intensity data, pollution source type data; "characteristic fluorescence data" refers to the fluorescence intensity data corresponding to characteristic wavelength points; "validation set" refers to the characteristic fluorescence data of re-collected wastewater samples from a single pollution source enterprise under various combinations of characteristic wavelength points; "optimal combination of characteristic wavelength points" refers to the combination of characteristic wavelength points obtained after validation by the validation set. "Fluorescence spectral data": matrix data composed of each wavelength point and its corresponding fluorescence intensity data.

[0064] Application Example 1

[0065] This application case uses the method provided in Example 1 to trace the source of water samples from rainwater inlets that may be affected by three pollution sources.

[0066] The specific method is as follows:

[0067] S1: Collect single-source wastewater samples from three pollution source enterprises (A, B, and C) around the storm drain inlet. Enterprise A is engaged in wool textile dyeing and finishing, enterprise B in papermaking, and enterprise C in chemical fiber dyeing and finishing. Obtain three-dimensional fluorescence spectral data for each water sample. The single-source wastewater samples were collected as follows: during the enterprise's production process, water samples were collected every 1-2 hours before the wastewater was discharged into the pipe network, for a total of 15 times; storm drain samples were collected every 5 minutes, for a total of 3 times.

[0068] The instrument parameters for acquiring the three-dimensional fluorescence spectral data are as follows: 150W xenon lamp, scanning speed of 12000nm / min, scanning range of excitation wavelength of 220-450nm, scanning interval of 1-5nm, scanning range of emission wavelength of 260-600nm, scanning interval of 1-10nm; the slit width of both excitation and emission wavelengths is 5nm.

[0069] S2: The three-dimensional fluorescence spectral data of S1 are preprocessed. The preprocessed fluorescence spectral data and their corresponding pollution source types are used as inputs to construct a random forest model with wavelength points as variables. Several feature wavelength points with the highest importance in distinguishing the pollution source types of each sample are selected.

[0070] In S2, the preprocessing method includes:

[0071] (S2-1) Delaunay interpolation combined with Gaussian function is used for smooth interpolation to remove Rayleigh scattering and Raman scattering;

[0072] (S2-2) Expand the three-dimensional fluorescence spectral data along the direction of the excitation wavelength to form a one-dimensional vector form in which the data points of adjacent rows are connected end to end;

[0073] (S2-3) Normalize the three-dimensional fluorescence spectral data to obtain normalized fluorescence intensity data with values ​​of 0 to 1.

[0074] The random forest model uses decision trees as the base classifier. The model parameters are optimized using fluorescence spectral data and their corresponding pollution source types. The parameter results are n_estimators = 120, min_sample_leaf = 1.1, min_samples_split = 1, and max_features = 126.

[0075] The top ten most important characteristic wavelengths for distinguishing the pollution source types of each sample were selected, designated as characteristic wavelengths 1-10. Before combining the S3 samples, the characteristic wavelengths were screened for any unreasonableness.

[0076] The selection principle is as follows:

[0077] (S3-1) In the data obtained in S6, if the fluorescence characteristic value corresponding to a certain characteristic wavelength point of a certain water sample exceeds the range of fluorescence characteristic values ​​corresponding to the characteristic wavelength point formed by the data obtained in S1, then delete the characteristic wavelength point; otherwise, retain the characteristic wavelength point.

[0078] (S3-2) If there is no significant difference in the fluorescence characteristic values ​​of different pollution source enterprises corresponding to a certain characteristic wavelength point in all the data obtained in S1, then delete the characteristic wavelength point; otherwise, retain the characteristic wavelength point.

[0079] Based on screening principle (S3-1), characteristic wavelength points 1, 7, and 9 are removed. Based on screening principle (S3-2), characteristic wavelength points 3, 4, 8, and 10 are removed. The selected characteristic wavelength points are: characteristic wavelength point 2 (excitation wavelength / emission wavelength (Ex / Em) = 250 / 455nm), characteristic wavelength point 5 (Ex / Em = 235 / 350nm), and characteristic wavelength point 6 (Ex / Em = 360 / 440nm).

[0080] S4: Collect single-source drainage water samples from known pollution sources around the rainwater inlet, obtain three-dimensional fluorescence spectral data of several water samples, and screen the characteristic fluorescence data of each water sample under each characteristic wavelength point combination as several verification groups for MixSIAR model verification.

[0081] The data files input into the MixSIAR model include: Mixture file, Sources file, and Discrimination file; among them, the Mixture file contains the validation group data, the Sources file contains the feature data group, and all data in the Discrimination file are set to 0. The format of all three data files is CSV; the Markov chain Monte Carlo run length is set to "normal"; the sum of the pollution contribution rates of each pollution source in the validation group is 1.

[0082] S5: Input the feature data set and the validation set under each combination of the same characteristic wavelength points into the MixSIAR model, output the contribution rate of each pollution source, and then verify the accuracy of the model output results according to the actual pollution source type of the validation set, so as to obtain the best combination of characteristic wavelength points with the highest accuracy.

[0083] The method with the highest accuracy is the one where the output pollutant type result is the same as the actual pollutant type result of the validation group, and the pollution contribution rate is the highest.

[0084] The experimental results are as follows: the optimal combination of characteristic wavelengths is characteristic wavelengths 2 and 6.

[0085] Model output results Feature 2 & Feature 5 Feature 2 & Feature 6 Feature 5 & Feature 6 A:B:C (A as the single source of pollution) 88:3:9 93:3:4 87:7:6 A:B:C (B as the single source of pollution) 10:87:3 2:90:7 9:86:5 A:B:C (C as the single source of pollution) 3:6:91 5:3:92 6:5:89

[0086] S6: Collect water samples from the rainwater inlet to be tested, obtain the three-dimensional fluorescence spectrum data of the rainwater inlet water samples, and after data preprocessing, select the characteristic fluorescence data of the rainwater inlet water samples under the optimal characteristic wavelength point combination, and input them together with the characteristic data group under the optimal characteristic wavelength point combination into the MixSIAR model to output the different contribution rates of various pollution sources in the rainwater inlet water samples to be tested, and determine the pollution sources of the rainwater inlet water samples by the contribution rate.

[0087] The source tracing results show that in the rainwater inlet drainage samples, the contribution rate of enterprise A was 69.6%, with an error of 10.4% compared to the actual rate; the contribution rate of enterprise B was 12.5%, with an error of 2.5% compared to the actual rate; and the contribution rate of enterprise C was 17.9%, with an error of 7.9% compared to the actual rate.

Claims

1. A method for quantitatively identifying mixed water pollution sources using three-dimensional fluorescence spectroscopy combined with a MixSIAR model, characterized in that, Includes the following steps: S1: Collect single-source wastewater samples from enterprises with different pollution sources around the mixed water body and obtain three-dimensional fluorescence spectral data of each water sample; S2: The three-dimensional fluorescence spectral data of S1 are preprocessed. The preprocessed fluorescence spectral data and their corresponding pollution source types are used as inputs to construct a random forest model with wavelength points as variables. Several feature wavelength points with the highest importance in distinguishing the pollution source types of each sample are selected. S3: Combine the fluorescence intensity data and pollution source type data corresponding to each pair of characteristic wavelength points to obtain the characteristic data group under each combination of characteristic wavelength points; S4: Collect single-source drainage water samples from known pollution sources around the mixed water body, obtain three-dimensional fluorescence spectral data of several water samples, and screen the characteristic fluorescence data of each water sample under each characteristic wavelength point combination as several validation groups for MixSIAR model validation. S5: Input the feature data set and the validation set under each combination of the same characteristic wavelength points into the MixSIAR model, output the contribution rate of each pollution source, and then verify the accuracy of the model output results according to the actual pollution source type of the validation set, so as to obtain the best combination of characteristic wavelength points with the highest accuracy. S6: Collect samples of the mixed water body to be tested, obtain three-dimensional fluorescence spectral data of the mixed water body samples, select the characteristic fluorescence data of the mixed water body samples under the optimal characteristic wavelength point combination after data preprocessing, and input them together with the characteristic data group under the optimal characteristic wavelength point combination into the MixSIAR model, output the different contribution rates of various pollution sources in the mixed water body sample to be tested, and determine the pollution sources of the mixed water body by the contribution rate.

2. The method for quantitatively identifying mixed water pollution sources using three-dimensional fluorescence spectroscopy combined with the MixSIAR model as described in claim 1, characterized in that, In S1, the number of pollution source types is less than 5; the method for collecting single-source drainage water samples is as follows: during the enterprise's production process, water samples are collected every 1 to 2 hours before the enterprise discharges into the pipeline network, for a total of 15 times; mixed water samples are collected every 5 minutes, for a total of 3 times.

3. The method for quantitatively identifying mixed water pollution sources using three-dimensional fluorescence spectroscopy combined with the MixSIAR model as described in claim 1, characterized in that, In S1, the instrument parameters for acquiring the three-dimensional fluorescence spectrum data are as follows: 150W xenon lamp, scanning speed of 12000 nm / min, scanning range of excitation wavelength of 220–450 nm, scanning interval of 1–5 nm, scanning range of emission wavelength of 260–600 nm, scanning interval of 1–10 nm; and slit width of 5 nm for both excitation and emission wavelengths.

4. The method for quantitatively identifying mixed water pollution sources using three-dimensional fluorescence spectroscopy combined with the MixSIAR model as described in claim 1, characterized in that, In S2, the preprocessing method includes: (S2-1) Delaunay interpolation combined with Gaussian function is used for smooth interpolation to remove Rayleigh scattering and Raman scattering; (S2-2) Unfold the three-dimensional fluorescence spectral data along the direction of the excitation wavelength to form a one-dimensional vector form in which the data points of adjacent rows are connected end to end; (S2-3) Normalize the three-dimensional fluorescence spectral data to obtain normalized fluorescence intensity data with values ​​of 0 to 1.

5. The method for quantitatively identifying mixed water pollution sources using three-dimensional fluorescence spectroscopy combined with the MixSIAR model as described in claim 1, characterized in that, In S2, the random forest model uses decision trees as the base classifier. The model parameters are optimized using fluorescence spectral data and their corresponding pollution source types. The optimized parameters include n_estimators, max_features, min_sample_leaf, and min_samples_split.

6. The method for quantitatively identifying mixed water pollution sources using three-dimensional fluorescence spectroscopy combined with the MixSIAR model as described in claim 1, characterized in that, In S2, the top ten characteristic wavelength points with the highest importance in distinguishing the pollution source types of each sample were selected. Before combining S3, the characteristic wavelength points are first screened for inconsistencies. The screening principle is as follows: (S3-1) In the data obtained in S6, if the fluorescence characteristic value corresponding to a certain characteristic wavelength point of a certain water sample exceeds the range of fluorescence characteristic values ​​corresponding to the characteristic wavelength point formed by the data obtained in S1, then delete the characteristic wavelength point; otherwise, retain the characteristic wavelength point. (S3-2) If there is no significant difference in the fluorescence characteristic values ​​of different pollution source enterprises corresponding to a certain characteristic wavelength point in all the data obtained in S1, then delete the characteristic wavelength point; otherwise, retain the characteristic wavelength point.

7. The method for quantitatively identifying mixed water pollution sources using three-dimensional fluorescence spectroscopy combined with the MixSIAR model as described in claim 1, characterized in that, In S3, the MixSIAR model is based on the Bayesian probability method and uses j feature fluorescence data to quantify the contribution of k pollution sources to the receptor. The model sets the contribution rate of each source to 1, uses the Dirichlet distribution as the prior distribution of pollution source contribution, and performs iterative calculation through the Monte Carlo Markov chain in the JAGS function. Based on the final output, the pollution contribution rate of each pollution source in the validation group sample is calculated. (1); (2); (3); In the formula, i represents the sample number of the verification group; j represents the characteristic wavelength point number; and k represents the enterprise type number. It is the fluorescence intensity value at the j-th characteristic wavelength point of the i-th validation group sample. It is the fluorescence intensity value at the j-th characteristic wavelength point of pollution source enterprise k. This represents the contribution rate of polluting enterprise k to the validation group sample, with a mean of [value missing]. The standard deviation is ; These are residuals, with a mean of zero and a standard deviation of [value missing]. N represents the Nth sample.

8. The method for quantitatively identifying mixed water pollution sources using three-dimensional fluorescence spectroscopy combined with the MixSIAR model as described in claim 1, characterized in that, The data files input into the MixSIAR model include: Mixture file, Sources file, and Discrimination file; among them, the Mixture file contains the validation group data, the Sources file contains the feature data group, and all data in the Discrimination file are set to 0. The format of all three data files is CSV; the Markov chain Monte Carlo run length is set to "normal"; the sum of the pollution contribution rates of each pollution source in the validation group is 1.

9. The method for quantitatively identifying mixed water pollution sources using three-dimensional fluorescence spectroscopy combined with the MixSIAR model as described in claim 1, characterized in that, In S5, the method with the highest accuracy is the one where the output pollutant type result is the same as the actual pollutant type result of the validation group, and the pollution contribution rate is the highest.

Citation Information

Patent Citations

  • Pollution source identification method and system based on ultraviolet-visible absorption spectrum

    CN113033623A

  • Pollution source identification method and device based on three-dimensional liquid chromatography fingerprints

    CN113311081A

  • A method and system for quantitatively identifying multiple pollution sources in mixed water bodies

    CN115219472B

  • Gutter inlet sunny day pollution discharge tracing method based on random forest recognition

    CN116595461A

  • Underground water nitrogen pollution source quantitative analysis method and device based on multiple tracers

    CN111624679A