Pollution source identification model optimization method and system based on three-dimensional fluorescence data enhancement

By employing methods such as cross-temporal sampling, spectral enhancement, and hierarchical iterative training, the problem of sample acquisition difficulties in pollution source identification models was solved, improving the model's stability and generalization ability, and achieving accurate identification of diverse pollution sources.

CN121834342APending Publication Date: 2026-04-10YANGTZE DELTA REGION INST OF TSINGHUA UNIV ZHEJIANG
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-29
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

In existing technologies, pollution source identification models are biased towards certain pollution source categories with obvious characteristics due to the difficulty in obtaining samples, which affects the stability and generalization ability of the models, and they are prone to overfitting, especially when faced with diverse pollution sources.

Method used

Multiple pollution source water samples were collected using a cross-temporal sampling rule. Three-dimensional fluorescence scanning and spectral homology enhancement were performed to generate a labeled training set. The pollution source water sample set was then subjected to destructive multi-granularity cross-mixing to generate a mixed pollution source water sample set, which was then spectrally decoupled and calibrated to form a label vector training set. A hierarchical iterative training method was used to optimize the pollution source component identification model.

Benefits of technology

Ensuring sample diversity and representativeness improves the model's generalization ability and recognition accuracy in complex scenarios, enabling it to handle both single and multi-source pollution identification tasks, thus enhancing the model's practicality and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121834342A_ABST
    Figure CN121834342A_ABST
Patent Text Reader

Abstract

The invention provides a pollution source identification model optimization method and system based on three-dimensional fluorescence data enhancement, and relates to the technical field of environmental pollution source identification, and the method comprises the steps: carrying out the three-dimensional fluorescence scanning after a plurality of pollution source water sample sets are collected across time and space, and obtaining a plurality of original three-dimensional fluorescence spectrum data sets; performing spectrum homologous fidelity enhancement to obtain a plurality of labeled training sets; carrying out destructive multi-granularity cross mixing to obtain a mixed pollution source water sample set; carrying out three-dimensional fluorescence scanning to obtain a mixed pollution spectrum set, and then carrying out spectrum multi-source decoupling calibration to obtain a label vector training set; and performing hierarchical iteration training output of the pollution source component identification model. The technical problems that in the prior art, a water body pollution data sample is difficult to obtain, a small number of samples cause the model to bias certain pollution source categories with obvious characteristics, the model is prone to overfitting when facing diversified pollution sources, and then the stability and generalization ability of the model are affected are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of environmental pollution source identification technology, specifically to a method and system for optimizing pollution source identification models based on three-dimensional fluorescence data enhancement. Background Technology

[0002] In environmental pollution source identification and monitoring, the use of three-dimensional fluorescence spectroscopy (EEM) technology for water quality monitoring and pollution source tracing has become a research and application hotspot. By analyzing the fluorescence characteristics of dissolved organic matter, three-dimensional fluorescence technology can identify pollution sources and their changes in water bodies, providing an effective tool for water quality assessment and pollution source tracing. The fluorescence characteristics of pollutants in water bodies exhibit significant variability with time, location, and different industrial activities. Due to the diverse types of pollution sources, the difficulty of sampling, and the dynamic changes of pollutants in water bodies, obtaining pollution source data presents considerable challenges.

[0003] Existing data samples are often limited, especially when there are many types of pollution sources, sampling difficulties, and high water variability, making it difficult to collect enough high-quality samples. A small number of samples can cause the model to become biased towards certain pollution source categories with distinct characteristics, making the model prone to overfitting when faced with diverse pollution sources, thus affecting the model's stability and generalization ability. Furthermore, fluctuations in pollutant composition and concentration further exacerbate the instability of sample quality, resulting in poor model adaptability to new data. Summary of the Invention

[0004] This application provides a method and system for optimizing pollution source identification models based on three-dimensional fluorescence data enhancement. It aims to solve the technical problems of existing technologies, such as the difficulty in obtaining water pollution data samples, the small number of samples causing the model to be biased towards certain pollution source categories with obvious characteristics, making the model prone to overfitting when facing diverse pollution sources, and thus affecting the stability and generalization ability of the model.

[0005] The first aspect disclosed in this application provides an optimization method for a pollution source identification model based on three-dimensional fluorescence data enhancement. The method includes: collecting multiple pollution source water sample sets across time and space based on sampling rules, performing three-dimensional fluorescence scanning to obtain multiple original three-dimensional fluorescence spectral datasets; performing spectral homology enhancement on the multiple original three-dimensional fluorescence spectral datasets to obtain multiple labeled training sets; performing destructive multi-granularity cross-mixing on the multiple pollution source water sample sets to obtain a mixed pollution source water sample set, wherein each mixed pollution source water sample in the mixed pollution source water sample set carries a multi-pollution source concentration vector; performing three-dimensional fluorescence scanning on the mixed pollution source water sample set to obtain a mixed pollution spectral set, and then performing spectral multi-source decoupling calibration to obtain a label vector training set; using the multiple labeled training sets as single-pollution training samples, the multiple original three-dimensional fluorescence spectral datasets as single-pollution test data, the label vector training set as mixed pollution training samples, and the mixed pollution spectral set as mixed pollution test data, performing hierarchical iterative training output for a pollution source component identification model.

[0006] The second aspect of this application discloses a pollution source identification model optimization system based on three-dimensional fluorescence data enhancement. The system is used in the aforementioned pollution source identification model optimization method based on three-dimensional fluorescence data enhancement. The system includes: a three-dimensional fluorescence scanning module, used to collect multiple pollution source water sample sets across time and space based on sampling rules, and then perform three-dimensional fluorescence scanning to obtain multiple original three-dimensional fluorescence spectral datasets; a spectral homology enhancement module, used to perform spectral homology enhancement on the multiple original three-dimensional fluorescence spectral datasets to obtain multiple labeled training sets; and a multi-granularity cross-mixing module, used to perform destructive multi-granularity cross-mixing on the multiple pollution source water sample sets to obtain mixed pollution... The system comprises: a water sample set containing multiple pollution sources, each sample in the mixed pollution source water sample set carrying a concentration vector of multiple pollution sources; a spectral multi-source decoupling calibration module for performing three-dimensional fluorescence scanning on the mixed pollution source water sample set to obtain a mixed pollution spectral set, followed by spectral multi-source decoupling calibration to obtain a label vector training set; and a model hierarchical iterative training module for using the multiple labeled training sets as single pollution training samples, the multiple original three-dimensional fluorescence spectral datasets as single pollution test data, the label vector training set as mixed pollution training samples, and the mixed pollution spectral set as mixed pollution test data to perform hierarchical iterative training output of the pollution source component identification model.

[0007] One or more technical solutions provided in this application have at least the following beneficial effects: Employing a cross-temporal sampling rule, pollutant source water samples were collected from multiple different water bodies. This sampling method ensures sample diversity and representativeness, covering different pollutant source types, concentrations, and distribution states. Spectral data was acquired through three-dimensional fluorescence scanning, enabling high-precision capture of the fluorescence characteristics of various pollutants in the water samples, providing accurate basic data for subsequent analysis. The original three-dimensional fluorescence spectral data was perturbed and enhanced to expand the dataset. This enhancement method ensured spectral consistency between the enhanced and original data while incorporating noise and other perturbation factors to increase data diversity. Finally, the enhanced data was combined with corresponding pollutant source labels to form a labeled training set. A destructive multi-granularity cross-mixing method was used to combine multiple pollutant source water samples to simulate complex pollution environments. Each mixed pollutant water sample contained concentration vectors of multiple pollutants, more realistically reflecting the coexistence of multiple pollutants in actual water bodies. This method effectively generated mixed pollutant samples. The use of mixed pollution source water sample sets enhances the model's practicality. Three-dimensional fluorescence scanning of the mixed pollution source water sample set yields a mixed pollution spectral set. Through multi-source spectral decoupling calibration, the independent spectral features of each pollution source are decoupled from the mixed pollution spectral data. The pollution source concentration information corresponding to each spectral sample is used as a label to form the final label vector training set. This decoupling technique ensures that the generated label vector training set accurately corresponds to the concentration information of each pollution source, providing high-quality data for subsequent model training. By using different datasets, the pollution source component identification model undergoes hierarchical iterative training to continuously optimize its predictive ability. The model first learns the identification ability of a single pollution source, and then further optimizes it using mixed pollution source data, enabling the model to identify complex mixed pollution source scenarios. This staged training method enhances the model's generalization ability and identification accuracy in complex scenarios, allowing it to handle both single and multi-pollution source identification tasks simultaneously.

[0008] The above description is merely an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this application more obvious and understandable, specific embodiments of this application are given below. Attached Figure Description

[0009] Figure 1 This is a schematic diagram of the process for optimizing a pollution source identification model based on three-dimensional fluorescence data enhancement, as provided in an embodiment of this application.

[0010] Figure 2 A schematic diagram of the structure of the pollution source identification model optimization system based on three-dimensional fluorescence data enhancement provided in the embodiments of this application.

[0011] Figure labeling: 3D fluorescence scanning module 10, spectral homology enhancement module 20, multi-granularity cross-mixing module 30, spectral multi-source decoupling calibration module 40, model hierarchical iterative training module 50. Detailed Implementation

[0012] To further illustrate the technical means and effects of the present invention in achieving its intended purpose, the following detailed description of the specific implementation methods, structures, features, and effects of the present invention, in conjunction with the accompanying drawings and preferred embodiments, is provided below.

[0013] Example 1, as Figure 1 As shown in the embodiments of this application, a method for optimizing a pollution source identification model based on three-dimensional fluorescence data enhancement is provided. The method includes: A100: Based on sampling rules, multiple pollution source water samples are collected across time and space, and then three-dimensional fluorescence scanning is performed to obtain multiple original three-dimensional fluorescence spectrum datasets.

[0014] Sampling rules were established to ensure the collection of water samples from different spatiotemporal environments. These samples were representative of various pollution sources, such as agricultural drainage, industrial wastewater, and domestic sewage. Sampling spanned different time periods and geographical locations to ensure the data reflected the changes in fluorescence characteristics of different pollution sources under varying environmental conditions. Three-dimensional fluorescence scanning is a method for analyzing the fluorescence characteristics of different pollutants in water samples through light excitation and emission. During scanning, excitation light irradiates the water sample, generating emission light of different wavelengths. By recording the emission spectra, an excitation-emission matrix for the water sample can be obtained. Multiple water samples from various pollution sources were scanned using a three-dimensional fluorescence scanner to obtain multiple raw three-dimensional fluorescence spectrum datasets. These datasets reflect the fluorescence characteristics of dissolved organic matter and pollution sources in the water. Exemplary parameter settings are as follows: excitation wavelength (Ex) range 220-450 nm, emission wavelength (Em) range 260-600 nm, scan interval Ex 5 nm / Em 1 nm, scan speed 2400 nm / min, slit width 5 nm.

[0015] A200: Perform spectral homology enhancement on the multiple original three-dimensional fluorescence spectral datasets to obtain multiple labeled training sets.

[0016] To enhance the stability of the dataset, spectral homology enhancement was performed on multiple original 3D fluorescence spectral datasets to generate more representative and consistent training data. Specifically, for multiple original 3D fluorescence spectral datasets, random stacking and noise introduction were used to simulate errors and variability in the real environment. Random stacking refers to selecting other samples of the same category and linearly combining them according to a certain proportion to generate new synthetic spectral data. This process preserves the fluorescence characteristics of the original data categories while increasing variability. Noise introduction refers to adding Gaussian noise or environmental background noise to the enhanced data to simulate unavoidable noise in experiments, such as instrument noise and environmental noise. These two steps work together to expand the dataset and increase the model's generalization ability. To ensure the quality of the enhanced data and its similarity to the original data, the reconstruction error between the enhanced samples and the original samples was calculated to select effective enhanced samples. These samples were then further labeled, i.e., assigned corresponding pollution source type labels, such as agricultural, industrial, or domestic sewage, to the enhanced spectral data, resulting in multiple labeled training sets for model training.

[0017] A300: The multiple pollution source water sample sets are subjected to destructive multi-granularity cross-mixing to obtain a mixed pollution source water sample set, wherein each mixed pollution source water sample in the mixed pollution source water sample set carries a multi-pollution source concentration vector.

[0018] Different pollution sources coexist in water bodies, necessitating the mixing of multiple pollution source water samples to simulate cross-contamination. Based on the combination of different pollution sources, multiple pollution source mixing scenarios are created. Each mixing scenario contains multiple pollution source water samples, forming different pollution combinations. For example, one scenario might be a mixture of agricultural and industrial pollution, while another might be a mixture of domestic sewage and agricultural pollution. Multiple pollution source water sample sets are cross-mixed, and multiple mixed samples are generated based on different mixing ratios. These samples contain the concentration characteristics of different pollution sources. Weighted combinations of pollution source water samples can be performed using preset mixing ratios, such as 30% agricultural pollution and 70% industrial pollution.

[0019] Each mixed-source water sample in the pool contains a multi-source concentration vector, which represents the proportion of each pollution source in the water sample. For example, in a mixed agricultural and industrial water sample, the concentration vector could be: agricultural pollution 0.3, industrial pollution 0.7. This concentration vector helps to distinguish the contribution of different pollution sources in the subsequent identification process.

[0020] A400: Perform three-dimensional fluorescence scanning on the mixed pollution source water sample set to obtain the mixed pollution spectrum set, and then perform spectral multi-source decoupling calibration to obtain the label vector training set.

[0021] Similar to the three-dimensional fluorescence scanning of water samples from multiple pollution sources, this scanned sample was a mixed pollution source water sample. Using the same instrument, within the specified excitation and emission wavelength range, the mixed pollution source water sample was subjected to three-dimensional fluorescence scanning, and spectral characteristics were recorded. The obtained mixed pollution spectral set contained fluorescence signals from different pollution sources, which can reflect the different concentrations and characteristics of multiple pollution sources in the water sample.

[0022] Since each mixed pollution source water sample contains multiple pollution sources, the spectral signals of these pollution sources overlap during the scanning process. Therefore, multi-source decoupling calibration is required to separate the fluorescence characteristics of each pollution source from the mixed signal. Specifically, spectral decoupling technology is used to extract the independent fluorescence characteristics of each pollution source from the mixed spectrum. For example, principal component analysis is used to reduce the dimensionality of the multidimensional spectral data, extract the most representative components through linear combination, and obtain the independent fluorescence characteristics of each pollution source through spectral decoupling.

[0023] By combining the multi-source concentration vectors of mixed-source water samples with the decoupled source feature data, each mixed-source sample is labeled with a source concentration vector, representing the relative concentration of different sources in the water sample. This generates a label vector training set, which is used to train a source component identification model so that the concentration of different sources can be predicted based on fluorescence data in future identification processes.

[0024] A500: The multiple labeled training sets are used as single-pollution training samples, the multiple original three-dimensional fluorescence spectrum datasets are used as single-pollution test data, the label vector training set is used as mixed-pollution training samples, and the mixed-pollution spectrum set is used as mixed-pollution test data to perform hierarchical iterative training output of the pollution source component identification model.

[0025] Multiple labeled training sets contain single water sample data for each pollution source and its corresponding label. Each training sample contains only the features and label of one pollution source, used for individual pollution source identification. During the testing phase, single-pollution water samples obtained from the original 3D fluorescence spectroscopy dataset are used for single pollution source identification. Each sample in the test dataset contains only one pollution source, and prediction is performed using a single-pollution model. To ensure the independence of the test dataset and the generalization ability of the model, K original 3D fluorescence spectroscopy data points are removed from the original dataset to form an independent test set. This ensures that the test set does not overlap with the training set, avoiding overfitting.

[0026] The label vector training set contains mixed samples of multiple pollution sources and corresponding pollution source concentration vector labels. This dataset is used to train a model capable of identifying complex scenarios where multiple pollution sources coexist. The characteristic of mixed pollution samples is that each water sample contains multiple pollution sources, and their fluorescence signals are mixed together. The test data here also removes independent data that may cause perturbations to avoid overfitting. During the testing phase, the performance of the pollution source component identification model is tested using a mixed pollution spectrum set that has undergone 3D fluorescence scanning. This data is used to verify whether the model can accurately identify multiple pollution sources from complex polluted environments and predict their concentrations. Similar to the single-pollution test data, the mixed pollution test data here also removes independent data that may cause perturbations to ensure that the test set does not overlap with the training set.

[0027] By iteratively training the aforementioned training set, the pollution source component identification model is optimized. A multi-stage training approach is adopted, first training a single pollution source identification model, then training a mixed pollution source identification model. This hierarchical training helps to gradually improve the model's complexity and recognition ability. During the training process, based on the model's performance on the test set, the model's hyperparameters are gradually optimized to improve the model's accuracy in different pollution source environments. Through multiple rounds of iterative training and validation, the model's performance in the pollution source component identification task is continuously improved.

[0028] Furthermore, the method also includes: A510: Pre-build the pollution source identification model based on the gradient boosting machine; A520: Use the multiple labeled training sets as single pollution training samples and the multiple original three-dimensional fluorescence spectrum datasets as single pollution test data to iteratively train and verify the pollution source identification model until the model's average accuracy and variance meet the first set of preset thresholds, and output the pollution source single-class model; A530: Use the labeled vector training set as mixed pollution training samples and the mixed pollution spectrum set as mixed pollution test data to iteratively train and verify the pollution source single-class model until the model's average accuracy and variance meet the second set of preset thresholds, and output the pollution source component identification model.

[0029] Gradient boosting machines (GPMs) are an ensemble learning method based on decision trees. By constructing multiple weak classifiers, such as regression trees, and using weighted summation, they progressively improve the model's prediction accuracy. In pollution source identification tasks, GPMs combine multiple tree models into a powerful predictor, effectively capturing nonlinear relationships and complex features in the data. This paper selects the gradient boosting machine framework and sets initial parameters, such as tree depth, learning rate, and maximum number of iterations, to pre-build a pollution source identification model, serving as the basis for subsequent iterative training.

[0030] Multiple labeled training sets are used as single-pollution training samples and input into a pre-built pollution source identification model. During training, multiple original 3D fluorescence spectral datasets are used as test data for single pollution sources. The test sets are different from the training sets and are used to verify the model's generalization ability on new samples. In iterative training, the pollution source identification model repeatedly adjusts its prediction weights to minimize the loss function, such as mean squared error or cross-entropy. After each training iteration, the model is fine-tuned based on the prediction results of the test data. The training process continues until the model's average accuracy and variance meet a first set of preset thresholds. The first set of preset thresholds is obtained through cross-validation and is used to ensure that the model has good stability and accuracy on different datasets. When the model training is complete and the stopping condition is met, the final output is a single-class pollution source model, which can independently identify the concentration of a specific pollution source.

[0031] The label vector training set was used as the training sample for mixed pollution, containing mixed water samples from multiple pollution sources and their corresponding pollution source concentration vector labels. This data was used to train the pollution source component identification model. The test data was a mixed pollution spectral set, which contained mixed features of multiple pollution sources. The test set was used to test the mixed pollution source model's ability to identify different pollution sources simultaneously.

[0032] The model is trained iteratively using training data from mixed pollution sources. In each round, the model parameters are continuously adjusted to optimize its predictive ability. To ensure the independence of the test dataset and the model's generalization ability, K original 3D fluorescence spectra are removed from the first original 3D fluorescence spectra dataset to create an independent test set. This ensures that the test set does not overlap with the training set, avoiding overfitting. During training, the model's performance is verified in each iteration by calculating its average accuracy and variance. If the results meet a second set of preset thresholds, the model is considered to have completed training and can effectively identify mixed pollution sources. At this point, the training process stops, and the final pollution source component identification model is output. This model can simultaneously handle multiple pollution sources and predict their concentrations.

[0033] Furthermore, the method involves performing spectral homology enhancement on the multiple original three-dimensional fluorescence spectral datasets to obtain multiple labeled training sets, the method comprising: A210: Perturb and enhance the multiple original three-dimensional fluorescence spectral datasets to obtain multiple enhanced three-dimensional fluorescence spectral datasets; A220: Calculate the reconstruction error of the multiple enhanced three-dimensional fluorescence spectral datasets based on the multiple original three-dimensional fluorescence spectral datasets to select multiple effective three-dimensional fluorescence spectral datasets; A230: Identify the multiple effective three-dimensional fluorescence spectral datasets using multiple pollution source labels from the multiple pollution source water sample sets to obtain the multiple labeled training sets.

[0034] Perturbation augmentation is a data augmentation technique that generates new samples by perturbing the original data, thereby expanding the training dataset and improving the model's generalization ability. For multiple original 3D fluorescence spectroscopy datasets, perturbation augmentation can be performed in the following ways: adding Gaussian noise or other types of random noise to simulate noise conditions that may occur during actual measurements. This method can effectively prevent model overfitting and enhance the model's robustness to data perturbations. Amplitude scaling or emission / excitation spectrum shifting of the original 3D fluorescence data can simulate spectral changes under different concentrations or environmental variations. Small-scale perturbations can be applied to the data within specific wavelength regions to simulate instabilities or different manifestations of pollution sources in experiments. Through these perturbations, multiple original 3D fluorescence spectroscopy datasets are augmented, resulting in multiple augmented 3D fluorescence spectroscopy datasets with greater diversity.

[0035] The effectiveness of each augmented dataset is evaluated by calculating the reconstruction error between multiple augmented 3D fluorescence spectroscopy datasets and multiple original 3D fluorescence spectroscopy datasets. Based on the calculation results of the reconstruction error, augmented datasets with smaller errors and high similarity to the original data are selected as multiple effective 3D fluorescence spectroscopy datasets.

[0036] For example, the reconstructive error is calculated: ; Represents the original sample data. This is a reconstruction error. This represents the enhanced sample. This represents the total number of samples. This is achieved through analysis of the original data samples. Augmentation strategies are applied to generate augmented samples. The differences between the augmented samples and the original samples are compared, and the squared error of each sample is calculated. The errors of all samples are averaged to obtain the overall reconstruction error. Based on the magnitude of the reconstruction error, effective augmented samples are selected. A lower reconstruction error indicates that the augmented sample has a higher similarity to the original sample and can better preserve the basic features of the sample. By controlling and optimizing the reconstruction error, it is possible to ensure that the sample features remain consistent with the original samples during the data augmentation process, thereby improving the stability and generalization ability of the model. Augmented samples with a reconstruction error below 10% are selected.

[0037] By matching the selected valid 3D fluorescence spectral datasets with pollution source labels, corresponding labels are added to each enhanced 3D fluorescence spectral dataset. These labels identify the pollution source category and its concentration in each spectral dataset. This results in multiple labeled training sets containing rich 3D fluorescence spectral features, with each dataset accompanied by pollution source label information.

[0038] Furthermore, the method involves perturbing and enhancing the multiple original three-dimensional fluorescence spectral datasets to obtain multiple enhanced three-dimensional fluorescence spectral datasets, the method comprising: A211: Randomly select K original three-dimensional fluorescence spectral data from the first original three-dimensional fluorescence spectral dataset of the first pollution source; A212: After presetting M initial weighting coefficients, perform random perturbation to obtain M sets of incremental weighting coefficients; A213: Use the M sets of incremental weighting coefficients to perform weighted linear superposition on the K original three-dimensional fluorescence spectral data to output K×M sets of intermediate three-dimensional fluorescence spectral data; A214: Inject water environment background noise into the K×M sets of intermediate three-dimensional fluorescence spectral data to obtain K×M sets of enhanced three-dimensional fluorescence spectral data, which constitute the first enhanced three-dimensional fluorescence spectral dataset.

[0039] From the first original three-dimensional fluorescence spectrum dataset of the first pollution source, K original three-dimensional fluorescence spectrum data are randomly selected. Random sampling methods, such as simple random sampling or equal probability sampling, are used to ensure that the selected samples are representative.

[0040] M initial weighting coefficients are set. These initial weighting coefficients are key parameters used to weight and superimpose subsequent spectral data. The initial weighting coefficients can be manually set or determined based on prior knowledge to simulate different pollution source concentrations and their varying effects on the three-dimensional fluorescence spectrum. After obtaining the M initial weighting coefficients, random perturbation is applied to generate M sets of incremental weighting coefficients. These incremental weighting coefficients represent different levels of perturbation, simulating the impact of measurement errors or environmental changes that may occur during the experiment on the data. The initial weighting coefficients can be perturbed using methods such as uniform distribution or normal distribution to obtain different incremental weighting coefficients. This process aims to increase data diversity.

[0041] Using the obtained M sets of incremental weighting coefficients, the K original three-dimensional fluorescence spectral data are weighted and linearly superimposed. The specific operation is as follows: Y ij =X i ×W j , where X i W represents the i-th raw spectral data. j Y represents the j-th incremental weight coefficient. ij This represents the intermediate spectral data after the i-th spectral data is perturbed by the j-th weighting coefficient. Through this weighted linear superposition method, K×M sets of intermediate three-dimensional fluorescence spectral data are finally obtained, which represent different combinations of pollution source concentrations and their impact on the spectrum.

[0042] Aquatic environmental background noise is injected into the K×M sets of intermediate three-dimensional fluorescence spectral data to simulate the influence of water on fluorescence signals in a real environment. This background noise originates from: environmental noise, such as stray light and electromagnetic interference in the water; measurement noise, such as sensor errors and light source instability; and water characteristics, such as turbidity and dissolved substances. Injecting noise into each set of spectral data makes the generated spectral data more closely resemble reality. After noise injection, K×M sets of enhanced three-dimensional fluorescence spectral data are obtained. These data reflect the three-dimensional fluorescence performance of pollution sources under the influence of environmental noise, simulating measurement data in a real environment. Ultimately, these data constitute the first enhanced three-dimensional fluorescence spectral dataset, serving as the data source for subsequent model training.

[0043] For example, the enhancement method is as follows: ; Represents the original sample data. This represents the enhanced sample. This represents the randomly assigned weight coefficients. This is a small random noise term, which can be Gaussian noise or environmental noise introduced from the water background to simulate experimental uncertainty. This method not only preserves the inherent fluorescence properties of each class but also introduces variability, thereby expanding the training dataset, enhancing the model's generalization ability, and reducing the risk of overfitting.

[0044] Furthermore, based on the multiple original three-dimensional fluorescence spectral datasets, the reconstruction error of the multiple enhanced three-dimensional fluorescence spectral datasets is calculated to filter and obtain multiple valid three-dimensional fluorescence spectral datasets. The method includes: A221: Construct K original fluorescence spectral data matrices based on the K original three-dimensional fluorescence spectral data, using the excitation wavelength as the row index and the emission wavelength as the column index; A222: Perform weighted linear superposition of the K original fluorescence spectral data matrices according to the M sets of incremental weighting coefficients to obtain a K×M set of reference superposition matrices; A223: Construct a K×M set of enhanced fluorescence spectral data matrices based on the K×M set of enhanced three-dimensional fluorescence spectral data, using the excitation wavelength as the row index and the emission wavelength as the column index; A224: Calculate the point-by-point mean square error of the K×M set of reference superposition matrices and the K×M set of enhanced fluorescence spectral data matrices to obtain the K×M set of reconstruction error values; A225: Use a preset error threshold to iterate through the K×M set of reconstruction error values ​​to filter and obtain the first effective three-dimensional fluorescence spectral dataset.

[0045] Three-dimensional fluorescence spectroscopy data has two main dimensions: excitation wavelength and emission wavelength. The excitation wavelength is the wavelength range used to excite fluorescence, and the emission wavelength is the wavelength range used to detect fluorescence emission. Each raw three-dimensional fluorescence spectroscopy data point can be viewed as a two-dimensional matrix, where rows represent different excitation wavelengths, columns represent different emission wavelengths, and each element in the matrix represents the fluorescence intensity under the corresponding excitation-emission wavelength combination. Therefore, based on K raw three-dimensional fluorescence spectroscopy data points, K raw fluorescence spectroscopy data matrices are constructed, with each matrix having a size equal to the number of excitation and emission wavelengths.

[0046] Using M sets of incremental weighting coefficients, K original fluorescence spectral data matrices are linearly superimposed with weights. The weighting process is similar to the weighted linear superposition process described above. Through weighted linear superposition, K×M sets of baseline superposition matrices are obtained, each matrix representing an enhanced spectral dataset after different combinations of weighting coefficients. The purpose of weighted linear superposition is to generate different spectral data matrices by changing the weighting coefficients, simulating fluorescence performance under different pollution source concentrations. Multiple superpositions enhance the diversity of the data, enabling subsequent model training to better adapt to different environmental factors and changes.

[0047] Based on the obtained K×M sets of enhanced three-dimensional fluorescence spectral data, a K×M set of enhanced fluorescence spectral data matrix is ​​constructed. Each enhanced spectral data is transformed into a two-dimensional matrix again after noise injection and weighted superposition. Similarly, the enhanced fluorescence spectral data matrix is ​​represented in the format of excitation wavelength as row index and emission wavelength as column index. The rows of each matrix represent the excitation wavelength, the columns represent the emission wavelength, and each element in the matrix represents the fluorescence intensity under different wavelength combinations. The K×M set of enhanced three-dimensional fluorescence spectral data matrix is ​​the set of two-dimensional matrices corresponding to each enhanced spectral data, totaling K sets, with M different weighted and enhanced combinations in each set.

[0048] For the K×M group reference superposition matrix and the K×M group enhanced fluorescence spectral data matrix, the mean square error between each point (i.e., each excitation-emission wavelength combination) is calculated. A smaller mean square error indicates that the enhanced spectral data is closer to the original data, indicating that the enhancement process is effective and accurate. The K×M group reconstruction error value is obtained after calculation. The reconstruction error value reflects the difference between the reference matrix and the enhancement matrix at each point, that is, the degree of deviation between the incremental weighted superposition and the actual data. Through these reconstruction error values, it is possible to determine which enhanced spectral data are effective.

[0049] For each K×M group of reconstruction error values, they are compared with a preset error threshold. If a group's reconstruction error value is less than or equal to the threshold, the enhanced spectral data is considered valid. If a group's error value is greater than the threshold, the enhanced data is considered inaccurate or invalid. Finally, the datasets that meet the criteria are selected and called the first valid three-dimensional fluorescence spectral dataset.

[0050] Furthermore, the method involves destructively cross-mixing the multiple pollution source water sample sets at multiple particle sizes to obtain a mixed pollution source water sample set, the method comprising: A310: After retrieving pollution events using the multiple pollution source tags as pollution source category retrieval keys, pollution scene aggregation is performed to obtain mixed pollution source scenes; A320: Multiple gradient mixing ratios for the mixed pollution source scenes are preset; A330: Destructive multi-granularity cross-mixing is performed on the multiple pollution source water sample sets according to the multiple gradient mixing ratios to obtain multiple mixed polluted water samples; A340: Multiple pollution source concentration vector tags are constructed based on the mixed pollution source scenes and multiple gradient mixing ratios; A350: The multiple pollution source concentration vector tags and multiple mixed polluted water samples are bound together to generate the mixed pollution source water sample set.

[0051] Using multiple pollution source tags as search keys, relevant pollution event information is extracted from known datasets. These pollution events are related to different pollution source types, pollution concentrations, and pollution locations. After retrieving the pollution events, they are aggregated to form mixed pollution source scenarios. This means that multiple pollution source events are merged into a comprehensive pollution scenario based on certain specific criteria, such as pollution source type, pollution concentration, and time period. This aggregation helps simulate and analyze the environmental impact of multiple pollution sources interacting. Through this aggregation, the resulting mixed pollution source scenarios can be used to simulate different types of polluted environments, thereby supporting more complex pollution source identification and analysis.

[0052] Based on different pollution source types, such as industrial pollution and agricultural pollution, and different pollution environmental conditions, such as water pollution and soil pollution, multiple pollution source mixed scenarios are preset. For example, one scenario may be dominated by industrial wastewater, another by agricultural waste, or a mixture of both. Within each pollution source mixed scenario, a gradient mixing ratio is set. These ratios control the degree of influence of different pollution sources on the pollution scenario. For example, in one mixed scenario, pollution source A accounts for 70% and pollution source B accounts for 30%; in another mixed scenario, pollution source A accounts for 50% and pollution source B accounts for 50%. By setting different ratios, environmental scenarios with the relative contributions of different pollution sources can be simulated, covering various possible pollution environmental conditions.

[0053] Based on multiple gradient mixing ratios, multiple pollutant source water samples are mixed at different particle sizes. Particle size can be defined as the concentration, type, or physical characteristics of different pollutants, such as particle size. For example, different concentration values ​​of pollutant source A and pollutant source B represent their impact in different environments. The term "destructive" is used here to emphasize the intensity and complexity of the mixing. Destructive mixing refers to simulating the interactions and reactions of pollutants in the natural environment while preserving the overall characteristics of the water samples. By cross-mixing different water samples, complex polluted water samples are generated, increasing the realism and diversity of the environment. Multiple sets of mixed polluted water samples are generated according to different mixing ratios, each set containing different pollutant concentrations and characteristics, simulating different pollution scenarios.

[0054] Each group of polluted water samples has a corresponding pollution source concentration vector label, representing the concentration values ​​of different pollution sources. For example, for a mixed water sample containing pollution source A and pollution source B, its pollution source concentration vector is [0.7, 0.3], indicating that the concentration of pollution source A is 70% and the concentration of pollution source B is 30%. The pollution source concentration vector of each mixed polluted water sample is constructed according to the set gradient mixing ratio. For example, if the mixing ratio of a certain scenario is set to 60% for pollution source A and 40% for pollution source B, the generated concentration vector label will be [0.6, 0.4]. For different mixing scenarios, multiple concentration vector labels are generated, each label corresponding to a different pollution source contribution and concentration. These labels will provide labeled data for subsequent model training.

[0055] Multiple sets of pollution source concentration vector labels are bound to multiple sets of mixed polluted water samples. In this way, each mixed polluted water sample has a corresponding pollution source concentration vector label. This label accurately reflects the relative concentration of each pollution source in the water sample. The bound water samples and labels are combined into a complete mixed pollution source water sample set, which contains multiple mixed water samples with different pollution source concentrations. Each water sample corresponds to a pollution source concentration vector label.

[0056] Furthermore, after performing three-dimensional fluorescence scanning on the mixed pollution source water sample set to obtain the mixed pollution spectral set, spectral multi-source decoupling calibration is performed to obtain a label vector training set. The method includes: A410: Perform three-dimensional fluorescence scanning on the multiple sets of mixed polluted water samples in the mixed pollution source water sample set to obtain multiple sets of original mixed pollution spectral data; A420: Perturb and enhance the multiple sets of original mixed pollution spectral data to obtain multiple sets of enhanced mixed pollution spectral data; A430: Based on the multiple sets of original mixed pollution spectral data, calculate the reconstruction error of the multiple sets of enhanced mixed pollution spectral data to screen and obtain multiple sets of effective three-dimensional fluorescence spectral data; A440: After identifying the multiple sets of effective three-dimensional fluorescence spectral data using the multiple sets of pollution source concentration vector labels, perform group splitting to construct the label vector training set.

[0057] Three-dimensional fluorescence scanning was performed on multiple sets of mixed polluted water samples to record the fluorescence signal between the excitation and emission wavelengths, generating multiple sets of original mixed pollutant spectral data. The spectral data of each water sample represents the fluorescence characteristics of the pollution source in the water sample. These data will be used for pollution source identification and analysis in subsequent processing.

[0058] Perturbation enhancement involves introducing random perturbations into multiple sets of original mixed pollution spectral data to simulate noise or changes that may occur in the natural environment. For example, slight random changes are made to the intensity, wavelength, and other characteristics of the spectral data to simulate factors such as sensor errors and interference from the external environment. After perturbation, multiple sets of enhanced mixed pollution spectral data are obtained. These enhanced data are very similar to the original data in content, but with the addition of certain perturbations, making the data more diverse.

[0059] Using multiple sets of original mixed pollution spectral data as a benchmark, multiple sets of enhanced mixed pollution spectral data are compared with multiple sets of original mixed pollution spectral data. By comparing the differences between the original data and the enhanced data, the reconstruction error is calculated. Based on the set error threshold, data with smaller reconstruction errors are selected as multiple sets of valid three-dimensional fluorescence spectral data. Valid data refers to data whose differences from the original data are within the allowable range, representing the rationality and accuracy of the enhanced spectral data.

[0060] Multiple sets of pollution source concentration vector labels are mapped to corresponding sets of valid 3D fluorescence spectral data, ensuring that each spectral data point carries clear pollution source category information. Through this mapping, the model can learn the pollution source concentration characteristics corresponding to each spectral data point. After the label mapping is completed, the valid 3D fluorescence spectral data containing labels are split into groups. This splitting method helps to further classify the data based on the characteristics of the pollution sources. Each group contains the concentration information of a specific pollution source, representing a specific pollution source category. A label vector training set is constructed based on the split data. This training set contains the concentration label of each pollution source and its corresponding 3D fluorescence spectral data.

[0061] Furthermore, the M sets of incremental weighting coefficients include zero-value weights to achieve selective combination of some samples of the K original three-dimensional fluorescence spectral data.

[0062] In the perturbation enhancement step, the original three-dimensional fluorescence spectral data are weighted and superimposed using M sets of incremental weight coefficients. Among them, the M sets of incremental weight coefficients include zero-value weights. This is to selectively combine the original spectral data. In other words, some data in the K original three-dimensional fluorescence spectral data will not participate in the weighted superposition during the enhancement process. By assigning zero-value weights to some samples, these samples are actually excluded during the training process to avoid their influence on the enhanced data, thereby selectively constructing the perturbation-enhanced dataset.

[0063] Furthermore, by removing the K original three-dimensional fluorescence spectral data from the first original three-dimensional fluorescence spectral dataset, an independent test set is obtained as single-pollution test data for single-round training and verification of the pollution source identification model.

[0064] From the first original 3D fluorescence spectrum dataset, K original 3D fluorescence spectrum data points that have already been used for enhancement are removed. These samples will no longer participate in the subsequent training process. The dataset after removal becomes the independent test set, which contains data samples that were not used for enhancement or perturbation. Using the independent test set, a single-round training validation of the pollution source identification model is performed. This part of the data contains only one pollution source and is used to verify the model's accuracy in identifying a single pollution source. In this way, the model's performance under specific pollution source conditions can be evaluated more accurately.

[0065] Example 2, based on the same inventive concept as the pollution source identification model optimization method based on three-dimensional fluorescence data enhancement in the foregoing examples, such as... Figure 2 As shown in the embodiments of this application, a pollution source identification model optimization system based on three-dimensional fluorescence data enhancement is provided. The system includes: The system comprises the following modules: a 3D fluorescence scanning module 10, which collects multiple pollution source water sample sets across time and space based on sampling rules, performs 3D fluorescence scanning to obtain multiple original 3D fluorescence spectral datasets; a spectral homology enhancement module 20, which enhances the spectral homology of the multiple original 3D fluorescence spectral datasets to obtain multiple labeled training sets; a multi-granularity cross-mixing module 30, which performs destructive multi-granularity cross-mixing on the multiple pollution source water sample sets to obtain a mixed pollution source water sample set, wherein each mixed pollution source water sample in the mixed pollution source water sample set carries a multi-pollution source concentration vector; a spectral multi-source decoupling calibration module 40, which performs 3D fluorescence scanning on the mixed pollution source water sample set to obtain a mixed pollution spectral set, and then performs spectral multi-source decoupling calibration to obtain a label vector training set; and a model hierarchical iterative training module 50, which uses the multiple labeled training sets as single pollution training samples, the multiple original 3D fluorescence spectral datasets as single pollution test data, the label vector training set as mixed pollution training samples, and the mixed pollution spectral set as mixed pollution test data to perform hierarchical iterative training output of the pollution source component identification model.

[0066] Furthermore, the model hierarchical iterative training module 50 is used to perform the following operation steps: The pollution source identification model is pre-constructed based on a gradient boosting machine. Multiple labeled training sets are used as single-pollution training samples, and multiple original three-dimensional fluorescence spectral datasets are used as single-pollution test data. Iterative training and verification of the pollution source identification model are performed until the model's average accuracy and variance meet the first set of preset thresholds, at which point a single-class pollution source model is output. The labeled vector training set is used as mixed-pollution training samples, and the mixed-pollution spectral set is used as mixed-pollution test data. Iterative training and verification of the single-class pollution source model are performed until the model's average accuracy and variance meet the second set of preset thresholds, at which point the pollution source component identification model is output.

[0067] Furthermore, the spectral homology enhancement module 20 is used to perform the following operation steps: The multiple original three-dimensional fluorescence spectral datasets are perturbed to obtain multiple enhanced three-dimensional fluorescence spectral datasets; based on the multiple original three-dimensional fluorescence spectral datasets, the reconstruction error of the multiple enhanced three-dimensional fluorescence spectral datasets is calculated to screen and obtain multiple effective three-dimensional fluorescence spectral datasets; the multiple effective three-dimensional fluorescence spectral datasets are identified by multiple pollution source labels of the multiple pollution source water sample sets to obtain the multiple labeled training sets.

[0068] Furthermore, the spectral homology enhancement module 20 is used to perform the following operation steps: K original three-dimensional fluorescence spectral data are randomly selected from the first original three-dimensional fluorescence spectral dataset of the first pollution source; after setting M initial weighting coefficients, random perturbation is performed to obtain M sets of incremental weighting coefficients; the K original three-dimensional fluorescence spectral data are weighted and linearly superimposed using the M sets of incremental weighting coefficients to output K×M sets of intermediate three-dimensional fluorescence spectral data; water environment background noise is injected into the K×M sets of intermediate three-dimensional fluorescence spectral data to obtain K×M sets of enhanced three-dimensional fluorescence spectral data, which constitute the first enhanced three-dimensional fluorescence spectral dataset.

[0069] Furthermore, the spectral homology enhancement module 20 is used to perform the following operation steps: Using excitation wavelength as the row index and emission wavelength as the column index, K original fluorescence spectral data matrices are constructed based on the K original three-dimensional fluorescence spectral data. The K original fluorescence spectral data matrices are then weighted linearly superimposed according to the M sets of incremental weighting coefficients to obtain a K×M set of reference superposition matrices. Using excitation wavelength as the row index and emission wavelength as the column index, a K×M set of enhanced three-dimensional fluorescence spectral data matrices are constructed based on the K×M set of enhanced three-dimensional fluorescence spectral data. The mean square error of the K×M set of reference superposition matrices and the K×M set of enhanced fluorescence spectral data matrices is calculated point-by-point to obtain the K×M set of reconstruction error values. A preset error threshold is used to iterate through the K×M set of reconstruction error values ​​to select the first effective three-dimensional fluorescence spectral dataset.

[0070] Furthermore, the multi-granularity cross-mixing module 30 is used to perform the following operation steps: After retrieving pollution events using the multiple pollution source tags as pollution source category search keys, pollution scene aggregation is performed to obtain mixed pollution source scenes. Multiple gradient mixing ratios for the mixed pollution source scenes are preset. Based on the multiple gradient mixing ratios, the multiple pollution source water sample sets are subjected to destructive multi-granularity cross-mixing to obtain multiple sets of mixed polluted water samples. Multiple pollution source concentration vector tags are constructed based on the multiple pollution source mixed scenes and multiple gradient mixing ratios. The multiple pollution source concentration vector tags and multiple sets of mixed polluted water samples are bound together to generate the mixed pollution source water sample set.

[0071] Furthermore, the spectral multi-source decoupling calibration module 40 is used to perform the following operation steps: Three-dimensional fluorescence scanning is performed on multiple sets of mixed polluted water samples from the mixed pollution source set to obtain multiple sets of original mixed pollution spectral data; perturbation enhancement is applied to the multiple sets of original mixed pollution spectral data to obtain multiple sets of enhanced mixed pollution spectral data; based on the multiple sets of original mixed pollution spectral data, the reconstruction error of the multiple sets of enhanced mixed pollution spectral data is calculated to screen out multiple sets of effective three-dimensional fluorescence spectral data; after identifying the multiple sets of effective three-dimensional fluorescence spectral data using the multiple sets of pollution source concentration vector labels, group splitting is performed to construct the label vector training set.

[0072] Furthermore, the M sets of incremental weighting coefficients include zero-value weights to achieve selective combination of some samples of the K original three-dimensional fluorescence spectral data.

[0073] Furthermore, by removing the K original three-dimensional fluorescence spectral data from the first original three-dimensional fluorescence spectral dataset, an independent test set is obtained as single-pollution test data for single-round training and verification of the pollution source identification model.

[0074] Through the foregoing detailed description of the pollution source identification model optimization method based on three-dimensional fluorescence data enhancement, those skilled in the art can clearly understand the pollution source identification model optimization system based on three-dimensional fluorescence data enhancement in this embodiment. Since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and relevant parts can be referred to the method section.

[0075] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any way. Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some modifications or alterations to the above-disclosed technical content to create equivalent embodiments without departing from the scope of the present invention. Any modifications, equivalent changes, and alterations made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the scope of the present invention.

Claims

1. A method for optimizing a pollution source identification model based on three-dimensional fluorescence data enhancement, characterized in that, The method includes: After collecting water samples from multiple pollution sources across time and space based on sampling rules, three-dimensional fluorescence scanning was performed to obtain multiple raw three-dimensional fluorescence spectrum datasets. The multiple original three-dimensional fluorescence spectral datasets were subjected to spectral homology enhancement to obtain multiple labeled training sets. The multiple pollution source water sample sets are subjected to destructive multi-granularity cross-mixing to obtain a mixed pollution source water sample set, wherein each mixed pollution source water sample in the mixed pollution source water sample set carries a multi-pollution source concentration vector; Three-dimensional fluorescence scanning was performed on the mixed pollution source water sample set to obtain the mixed pollution spectrum set. Then, multi-source spectral decoupling calibration was performed to obtain the label vector training set. The multiple labeled training sets are used as single-pollution training samples, the multiple original three-dimensional fluorescence spectrum datasets are used as single-pollution test data, the label vector training set is used as mixed-pollution training samples, and the mixed-pollution spectrum set is used as mixed-pollution test data to perform hierarchical iterative training output of the pollution source component identification model.

2. The method for optimizing a pollution source identification model based on three-dimensional fluorescence data enhancement as described in claim 1, characterized in that, The method further includes: The pollution source identification model is pre-constructed based on a gradient boosting machine. The multiple labeled training sets are used as single pollution training samples, and the multiple original three-dimensional fluorescence spectrum datasets are used as single pollution test data. Iterative training and verification of the pollution source identification model are carried out until the average accuracy and variance of the model meet the first set of preset thresholds, and the single-class pollution source model is output. The label vector training set is used as the mixed pollution training sample, and the mixed pollution spectrum set is used as the mixed pollution test data. Iterative training and verification of the single-class pollution source model are carried out until the average accuracy and variance of the model meet the second set of preset thresholds, and the pollution source component identification model is output.

3. The method for optimizing a pollution source identification model based on three-dimensional fluorescence data enhancement as described in claim 2, characterized in that, The method involves performing spectral homology enhancement on the multiple original three-dimensional fluorescence spectral datasets to obtain multiple labeled training sets. The original three-dimensional fluorescence spectral datasets are perturbed and enhanced to obtain multiple enhanced three-dimensional fluorescence spectral datasets. Based on the multiple original three-dimensional fluorescence spectral datasets, the reconstruction error of the multiple enhanced three-dimensional fluorescence spectral datasets is calculated to screen and obtain multiple effective three-dimensional fluorescence spectral datasets; The multiple effective three-dimensional fluorescence spectral datasets are identified by multiple pollution source labels of the multiple pollution source water sample sets, thus obtaining the multiple labeled training sets.

4. The method for optimizing a pollution source identification model based on three-dimensional fluorescence data enhancement as described in claim 3, characterized in that, The method involves perturbating and enhancing the multiple original three-dimensional fluorescence spectral datasets to obtain multiple enhanced three-dimensional fluorescence spectral datasets, the method comprising: K original three-dimensional fluorescence spectral data were randomly selected from the first original three-dimensional fluorescence spectral dataset of the first pollution source; After setting M initial weight coefficients, random perturbation is performed to obtain M sets of incremental weight coefficients; The K original three-dimensional fluorescence spectral data are weighted and linearly superimposed using the M sets of incremental weighting coefficients to output K×M sets of intermediate three-dimensional fluorescence spectral data. Aquatic environmental background noise is injected into the K×M group of intermediate three-dimensional fluorescence spectral data to obtain K×M group of enhanced three-dimensional fluorescence spectral data, which constitute the first enhanced three-dimensional fluorescence spectral dataset.

5. The method for optimizing a pollution source identification model based on three-dimensional fluorescence data enhancement as described in claim 4, characterized in that, Using the multiple original three-dimensional fluorescence spectral datasets as a benchmark, the reconstruction error of the multiple enhanced three-dimensional fluorescence spectral datasets is calculated to filter and obtain multiple valid three-dimensional fluorescence spectral datasets. The method includes: Using the excitation wavelength as the row index and the emission wavelength as the column index, a matrix of K original fluorescence spectra is constructed based on the K original three-dimensional fluorescence spectra. The K original fluorescence spectral data matrices are weighted and linearly superimposed based on the M sets of incremental weighting coefficients to obtain a K×M set of reference superposition matrices; Using the excitation wavelength as the row index and the emission wavelength as the column index, a K×M enhanced fluorescence spectral data matrix is ​​constructed based on the K×M groups of enhanced three-dimensional fluorescence spectral data. The point-by-point mean square error of the K×M group reference superposition matrix and the K×M group enhanced fluorescence spectral data matrix is ​​calculated to obtain the K×M group reconstruction error value; The K×M groups of reconstruction error values ​​are traversed using a preset error threshold to filter and obtain the first effective three-dimensional fluorescence spectrum dataset.

6. The method for optimizing a pollution source identification model based on three-dimensional fluorescence data enhancement as described in claim 3, characterized in that, The method involves destructively cross-mixing multiple pollutant source water sample sets at various particle sizes to obtain a mixed pollutant source water sample set. After retrieving pollution events using the multiple pollution source tags as pollution source category search keys, pollution scene aggregation is performed to obtain mixed pollution source scenes. Preset multiple gradient mixing ratios for the aforementioned mixed pollution source scenarios; Based on the aforementioned multiple gradient mixing ratios, the multiple sets of polluted water samples are subjected to destructive multi-particle cross-mixing to obtain multiple sets of mixed polluted water samples; Based on the aforementioned mixed pollution source scenarios and multiple sets of gradient mixing ratios, multiple sets of pollution source concentration vector labels are constructed. The mixed pollution source water sample set is generated by binding the multiple sets of pollution source concentration vector labels and the multiple sets of mixed pollution water samples.

7. The method for optimizing a pollution source identification model based on three-dimensional fluorescence data enhancement as described in claim 6, characterized in that, After performing three-dimensional fluorescence scanning on the mixed pollution source water sample set to obtain the mixed pollution spectral set, spectral multi-source decoupling calibration is performed to obtain a label vector training set. The method includes: Three-dimensional fluorescence scanning was performed on multiple sets of mixed pollutant water samples from the mixed pollution source collection to obtain multiple sets of original mixed pollution spectral data. The original mixed pollution spectral data of the multiple sets of raw mixed pollution spectral data are perturbed and enhanced to obtain multiple sets of enhanced mixed pollution spectral data; Based on the original mixed pollution spectral data, the reconstruction error of the enhanced mixed pollution spectral data is calculated to screen and obtain multiple sets of effective three-dimensional fluorescence spectral data. After identifying the multiple sets of effective three-dimensional fluorescence spectral data using the multiple sets of pollution source concentration vector labels, the data is split into groups to construct the label vector training set.

8. The method for optimizing a pollution source identification model based on three-dimensional fluorescence data enhancement as described in claim 4, characterized in that, The M groups of incremental weighting coefficients include zero-value weights to achieve selective combination of some samples from the K original three-dimensional fluorescence spectral data.

9. The method for optimizing a pollution source identification model based on three-dimensional fluorescence data enhancement as described in claim 4, characterized in that, The K original three-dimensional fluorescence spectral data are removed from the first original three-dimensional fluorescence spectral dataset to obtain an independent test set, which is used as single pollution test data for single-round training and verification of the pollution source identification model.

10. A pollution source identification model optimization system based on three-dimensional fluorescence data enhancement, characterized in that, The system is used to implement the pollution source identification model optimization method based on three-dimensional fluorescence data enhancement as described in any one of claims 1-9, the system comprising: The three-dimensional fluorescence scanning module is used to collect multiple water samples from pollution sources across time and space based on sampling rules, and then perform three-dimensional fluorescence scanning to obtain multiple raw three-dimensional fluorescence spectrum datasets. The spectral homology enhancement module is used to perform spectral homology enhancement on the multiple original three-dimensional fluorescence spectral datasets to obtain multiple labeled training sets. A multi-granularity cross-mixing module is used to perform destructive multi-granularity cross-mixing on the multiple pollution source water sample sets to obtain a mixed pollution source water sample set, wherein each mixed pollution source water sample in the mixed pollution source water sample set carries a multi-pollution source concentration vector; The spectral multi-source decoupling calibration module is used to perform three-dimensional fluorescence scanning on the mixed pollution source water sample set, obtain the mixed pollution spectral set, and then perform spectral multi-source decoupling calibration to obtain a label vector training set. The model hierarchical iterative training module is used to perform hierarchical iterative training output of the pollution source component identification model by using the multiple labeled training sets as single pollution training samples, the multiple original three-dimensional fluorescence spectrum datasets as single pollution test data, the label vector training set as mixed pollution training samples, and the mixed pollution spectrum set as mixed pollution test data.