Method for identification of risk factors for extracellular polymeric substances during uv / peroxyacetic acid disinfection process
By using Fourier transform ion cyclotron resonance mass spectrometry and machine learning models, the molecular transformation types and risk factors of extracellular polymers during UV/peracetic acid disinfection were identified. This solved the shortcomings of existing technologies in extracellular polymer risk identification and enabled efficient risk factor screening and molecular structure risk correlation analysis.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- TONGJI UNIV
- Filing Date
- 2026-05-20
- Publication Date
- 2026-07-21
Smart Images

Figure CN122430431A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of wastewater treatment technology, specifically to a method for identifying risk factors of extracellular polymers in the ultraviolet / peracetic acid disinfection process. Background Technology
[0002] In today's increasingly scarce water resources, wastewater disinfection is becoming increasingly important for ensuring the safety of reclaimed water. During the disinfection process, disinfectants inevitably come into contact with precursors that lead to the formation of byproducts. Among these, extracellular polymers (APIs) are important structural and reactive matrices in wastewater disinfection systems. They not only participate in multiple water treatment processes but also mediate interactions between microorganisms and their external environment. Due to their abundant binding and reactive sites, the transformation of APIs during disinfection is not merely a change in composition but also has direct environmental significance. Some traditional disinfectants, such as sodium hypochlorite, chloramine, and chlorine dioxide, may produce byproducts exhibiting cytotoxicity or neurotoxicity upon contact with APIs after the available chlorine generated during disinfection comes into contact with APIs. Therefore, the development of novel disinfection processes and the risk identification of APIs require special attention.
[0003] In recent years, peracetic acid (PAA) disinfection has attracted increasing attention due to its strong oxidizing power and limited disinfection byproduct generation compared to traditional disinfection processes. Among various PAA-based combined disinfection processes, ultraviolet / peracetic acid (UV / PAA) is particularly promising. This system combines direct photolysis, oxidant oxidation, and free radical oxidation, and has shown good potential for pathogen control, resistance reduction, and pollutant removal in real wastewater matrices. However, the environmental significance of UV / PAA cannot be judged solely by its disinfection efficiency. Highly efficient oxidation does not necessarily imply detoxification at the molecular level, especially for complex matrices with highly heterogeneous chemical compositions, such as extracellular polymers. Existing research mainly focuses on microbial inactivation, pollutant removal, or macroscopic spectral changes, while the molecular-level transformation patterns of extracellular polymers during UV / PAA disinfection and the resulting risk changes have received little attention.
[0004] In recent years, with the development of non-targeted mass spectrometry and reaction-process-based analytical methods, the specific changes in complex organic mixtures during transformation processes can be resolved at the molecular level. However, molecular characterization alone is insufficient to determine the direction of risk transformation in complex organic matter, let alone answer the contribution of transformation results to residual risk, or how to rapidly identify them. Against this backdrop, interpretable machine learning demonstrates unique advantages because it can integrate multidimensional molecular information and extract chemically meaningful risk screening and prioritization rules. Unfortunately, such methods are currently rarely used to integrate the structural transformation, toxicity evolution, and rapid risk identification of extracellular polymers during sterilization into a single analytical framework. Summary of the Invention
[0005] The present invention aims to provide a method for identifying risk factors of extracellular polymers in the ultraviolet / peracetic acid disinfection process, so as to solve the problems existing in the prior art.
[0006] Specifically, the method for identifying risk factors of extracellular polymers in the ultraviolet / peracetic acid disinfection process of the present invention includes the following steps: S1. Extract total extracellular polymers from the secondary effluent of a wastewater treatment plant to obtain an extracellular polymer sample; S2. The extracellular polymer sample obtained in step S1 was placed in a UV / peracetic acid system for sterilization, and samples were collected at different reaction time points; S3. Perform Fourier transform ion cyclotron resonance mass spectrometry analysis on the samples obtained at different reaction time points in step S2 to obtain molecular formula data at different reaction time points; S4. Based on the molecular formula data obtained in step S3, calculate the molecular parameters, match the molecular formula data to SMILES structural information, and further calculate the molecular structure descriptor; S5. Construct precursor-product molecular pairs based on the changes in molecular formula between adjacent reaction stages, and identify the molecular transformation type of extracellular polymers during UV / peracetic acid sterilization based on the mass difference relationship; S6. Based on the molecular parameters and molecular structure descriptors obtained in step S4, construct a binary risk label by combining toxicity prediction data, and establish a machine learning risk identification model. Further, screen key risk factors of extracellular polymers in the UV / peracetic acid disinfection process through model interpretability analysis.
[0007] Furthermore, in step S1, the total extracellular polymers are extracted using a modified heating-sodium carbonate combined method: the secondary effluent sample is centrifuged and enriched, the precipitate is washed with NaCl solution and resuspended, sodium carbonate is added to the resuspended solution and heated, centrifuged, and filtered to obtain a crude extract of total extracellular polymers, and then the crude extract of extracellular polymers is dialyzed to obtain a total extracellular polymer sample.
[0008] The secondary effluent was first centrifuged and enriched, then washed three times with 0.9% NaCl to reduce interference from dissolved background substances and maintain the stability of extracellular bound substances as much as possible. Sodium carbonate was then added to promote the detachment of extracellular polymers from the periplasmic environment and simultaneously stabilize active structures such as enzymes and proteins, thus obtaining a mixed solution containing sodium carbonate.
[0009] Furthermore, the NaCl solution has a mass concentration of 0.9%; the sodium carbonate concentration in the resuspension is 0.5 wt%; the heating temperature is 55–85 ℃, and the heating time is 20–35 min; the dialysis membrane used for dialysis has a molecular cutoff of 3500 Da.
[0010] Furthermore, the heating temperature was 80 ℃ and the heating time was 30 min; after heating, centrifugation was performed at 11000 g for 15 min, and the supernatant was collected; the obtained supernatant was filtered through a 0.22 μm filter membrane to obtain a crude extract of total extracellular polymeric substances.
[0011] Furthermore, the sample was dialyzed with deionized water to remove small molecule impurities, yielding a total extracellular polymeric sample.
[0012] Furthermore, the experimental conditions for sterilizing the extracellular polymeric samples in step S2 using a UV / peracetic acid system were set as follows: peracetic acid concentration of 10 mg / L, initial pH of 7.0 ± 0.1, and UV dose of 108 mJ / cm². 2 Step S3 involves different reaction time points: 0, 10, 30, and 60 min, divided into three reaction stages: 0-10 min, 10-30 min, and 30-60 min. This step simulates the structural changes and molecular transformations of extracellular polymers during UV / peracetic acid sterilization.
[0013] Further, the specific steps of step S3 are as follows: Solid-phase extraction is performed on the samples obtained at different reaction time points from step S2, and the samples are then tested to obtain molecular formula data for different reaction time points. Specifically, FT-ICR-MS technology is used to obtain the molecular formula data of extracellular polymers at different time points, which is used to characterize the changes in the molecular composition and molecular properties of extracellular polymers during the disinfection process. This is specifically divided into four time points: 0 min, 10 min, 30 min, and 60 min.
[0014] Furthermore, the molecular parameters in step S4 include O / C, H / C, N / C, DBE, NOSC, and AI. mod The descriptors include MolLogP, TPSA, FormalCharge, NumRotatableBonds, RingCount, BertzCT, and ExactMassformula.
[0015] Furthermore, in step S4, the O / C, H / C, and N / C parameters are used to characterize the elemental composition of the molecule, representing the ratios of oxygen, hydrogen, and nitrogen atoms to carbon atoms, respectively; DBE represents the double bond equivalent, reflecting the degree of unsaturation of the molecule, calculated as (2C + 2-H + N) / 2; NOSC represents the oxidation state of the saturated carbon in the molecule, reflecting the redox state of the molecule, calculated as 4 - [(4C + H-3 N-2O-2S) / C]; AI mod The corrosion index represents the molecular value and is used to characterize the aromaticity of a molecule. It is calculated as [1 + C - 0.5OS - 0.5(N + H)] / (C - 0.5O-S - N). These parameters can be used to describe the oxidation, cleavage, and structural reorganization characteristics of extracellular polymers during sterilization at the molecular level.
[0016] Furthermore, in step S4, the molecular formula data obtained in step S3 is matched with SMILES structural information, and the SMILES structural information is used as input value to calculate the molecular structure descriptor based on the code encapsulated in the RDKit package of Python.
[0017] Furthermore, in step S4, the descriptors MolLogP, TPSA, FormalCharge, NumRotatableBonds, RingCount, BertzCT, and ExactMassformula are used to characterize the molecule's hydrophobicity, polarity, formal charge, structural flexibility, number of ring structures, structural complexity, and precise mass, respectively, providing feature inputs for subsequent toxicity risk identification.
[0018] Furthermore, the conversion type in step S5 includes one or more of oxidation reactions, cracking reactions, dealkylation reactions, sulfur-containing reactions, and nitrogen-containing reactions.
[0019] Furthermore, in step S5, by comparing the molecular formulas at two adjacent time points of 0 min, 10 min, 30 min, and 60 min, a unique molecular formula is identified at each time point. For example, in the 0-10 min stage, molecular formulas present at 0 min but absent at 10 min are classified as precursors for this stage, while those absent at 0 min but present at 10 min are products for this stage. By comparing the mass difference relationship between these precursors and products, the molecular transformation type of this reaction stage can be determined. This step allows us to obtain the main molecular transformation pathways of extracellular polymers in different reaction stages.
[0020] Furthermore, in step S6, the toxicity prediction data is obtained by calculation using ECOSAR software, and a logarithmic operation is performed based on the concentration values in the toxicity prediction data. 10(Concentration) conversion, thereby constructing high-risk and low-risk binary risk labels.
[0021] Further, step S6 involves inputting the SMILES structural information obtained in step S4 into the ECOSAR software to obtain the corresponding toxicity prediction data. The toxicity prediction results are then processed and standardized, and based on logarithmic... 10 (Concentration) Construct molecular risk labels. Preferably, molecular risks are divided into two categories: high risk and low risk, thereby forming a binary risk dataset that can be used for machine learning modeling.
[0022] Furthermore, the SMILES structure information obtained in step S4 is input into ECOSAR software and calculated to obtain a preprocessed dataset. The preprocessed dataset is then processed using Python language through PyCharm software to obtain a binary risk dataset, specifically with two classification labels: high risk and low risk.
[0023] Furthermore, the binary risk label is constructed as follows: when log 10 When the concentration is ≤2, the corresponding molecule is marked as a high-risk molecule; when log 10 When the concentration is greater than 2, the corresponding molecule will be marked as a low-risk molecule.
[0024] Furthermore, in step S6, the machine learning risk identification model is an XGBoost model. The model uses the molecular parameters and molecular structure descriptors calculated in step S4 as input features and high-risk and low-risk binary risk labels as output labels to identify the risk category of extracellular polymer transformation products. Its evaluation indicators include one or more of AUC, PRAUC, accuracy, precision, recall, and F1 score.
[0025] Furthermore, the performance of the XGBoost model was improved by using hyperparameter tuning and random search methods for training and parameter tuning.
[0026] Furthermore, the prediction performance of the XGBoost model is evaluated using the five-fold method and cross-validation method.
[0027] Furthermore, in step S6, the model interpretability analysis adopts the SHAP method, which screens key risk factors by calculating the contribution value of different input features to the machine learning output results.
[0028] Compared with the prior art, the present invention has the following outstanding features and advantages: 1) This invention relates to a method for analyzing the toxicity transformation of extracellular polymers under UV / peracetic acid treatment and identifying key risk factors. A supervised machine learning model based on molecular descriptors and molecular composition can efficiently perform binary classification to identify the molecular risks of extracellular polymers during UV / peracetic acid disinfection. The model's AUC and PRAUC on the test set both reached 0.979. Furthermore, the accuracy, F1 score, precision, and recall were 0.934, 0.937, 0.926, and 0.948, respectively, indicating that this method has the ability to identify and compare high-risk and low-risk molecules.
[0029] 2) SHAP interpretability analysis revealed that this invention can identify major risk factors from multiple molecular descriptors. Based on toxicity correlation analysis, numerous factors associated with high toxicity were identified, and the interpretability results further highlighted this, revealing O / C, MolLogP, and TPSA as the three highest contributing features, with absolute SHAP values of 0.686, 0.624, and 0.400, respectively, corresponding to explanatory contributions of 10.5%, 9.5%, and 6.1%. This indicates that risk identification is primarily controlled by oxidation degree, hydrophobicity, and polarity. Furthermore, this invention can further identify changes in molecular risk by recognizing the transformation of risk factors. Results showed that risk reduction is highly correlated with molecular structure transitions towards higher oxidation degree, higher polarity, and lower hydrophobicity; increases in O / C, TPSA, and NOSC are generally associated with risk reduction, while increases in MolLogP are detrimental to risk reduction. Therefore, this method can identify key features of risk reduction in extracellular polymer transformation products at the molecular property level.
[0030] 3) This invention first focuses on the risk transformation of extracellular polymers during UV / peracetic acid disinfection. It was found that after 60 min of disinfection, the proportion of high-risk molecules decreased from 55.8% to 34.0%, while the proportion of low-risk molecules increased from 44.2% to 66.0%. Therefore, this method can be used to characterize the trend of extracellular polymer molecular risk transformation from high-risk to low-risk components during disinfection. Furthermore, this invention employs Fourier transform ion cyclotron resonance mass spectrometry to further analyze the changes in extracellular polymers at the molecular level. By analyzing the transformation reaction types of precursor-product pairs, the risk change mechanism of extracellular polymers can be further assessed. Specifically, oxidation reactions consistently correlated with risk reduction in all three stages, while cleavage reactions were often associated with increased risk or retention of high-risk molecules. Specifically, oxidation reactions reduced the number of high-risk molecules in the identified precursor-product pairs from 125 to 91, from 217 to 147, and from 123 to 74 in the three time periods, respectively. Attached Figure Description
[0031] Figure 1The graph shows the changes in molecular properties of extracellular polymers during different time periods in the UV / peracetic acid sterilization process. (a), (b), and (c) are the distribution diagrams comparing the corrosion index and carbon number of the precursor-product at 0-10 min, 10-30 min, and 30-60 min, respectively. Figure 2 The graph shows the changes in molecular properties of extracellular polymers during different time periods in the UV / peracetic acid sterilization process. (a), (b), and (c) are the distribution diagrams of the saturation and molecular weight of the precursor-product at 0-10 min, 10-30 min, and 30-60 min, respectively. Figure 3 The graph shows the changes in molecular properties of extracellular polymers during different time periods in the UV / peracetic acid sterilization process. (a), (b), and (c) are comparison distributions of the oxidation degree and saturation of precursor-products at 0-10 min, 10-30 min, and 30-60 min, respectively. Figure 4 Kendrick quality defect diagrams of precursor-product pairs of extracellular polymers at different time points during UV / peracetic acid sterilization are shown. (a), (b), and (c) are the COO Kendrick defect diagrams of the precursor-product at 0-10 min, 10-30 min, and 30-60 min, respectively. Figure 5 Kendrick quality defect diagrams of precursor-product pairs of extracellular polymers during different time periods in UV / peracetic acid sterilization are shown. (a), (b), and (c) are the CH2 Kendrick defect diagrams of precursor-products at 0-10 min, 10-30 min, and 30-60 min, respectively. Figure 6 Kendrick quality defect diagrams of precursor-product pairs of extracellular polymers at different time points during UV / peracetic acid sterilization are shown. (a), (b), and (c) are Kendrick defect diagrams of precursor-product O corresponding to 0-10 min, 10-30 min, and 30-60 min, respectively. Figure 7 A statistical graph showing the actual reaction type of the precursor-product pair of extracellular polymers at different time points during UV / peracetic acid sterilization. Figure 8The redistribution of toxicity risk categories of extracellular polymers during UV / peracetic acid disinfection at different time points is shown in the diagram. (a) shows a dumbbell plot, which shows the ratio of high-risk molecules to low-risk molecules in different reaction stages. (b) to (d) show Sankey plots of toxicity distribution of identified precursor-product pairs at different time points, illustrating the toxicity risk of the identified precursor molecules after different reaction types in the time periods of 0-10 min, 10-30 min, and 30-60 min. Figure 9 The distribution of the effects of changes in key molecular descriptors and molecular composition parameters on the changes in toxicity risk of extracellular polymers during UV / peracetic acid disinfection is shown, where (a) and (b) show the effects of the range of changes in MolLogP and O / C on the changes in toxicity risk, respectively. Figure 10 The distribution of the effects of changes in key molecular descriptors and molecular composition parameters on the changes in toxicity risk of extracellular polymers during UV / peracetic acid disinfection is shown, where (a) and (b) show the effects of the range of changes in TPSA and NOSC on the changes in toxicity risk, respectively. Figure 11 The distribution of the effects of changes in key molecular descriptors and molecular composition parameters on the changes in toxicity risk of extracellular polymers during UV / peracetic acid disinfection is shown in (a) and (b), respectively, where (a) and (b) show the effects of the range of changes in BertzCT and FormalCharge on the changes in toxicity risk. Figure 12 The relationship between EPS molecular toxicity and key molecular descriptors and molecular composition parameters is shown, where (a) shows a heatmap of the correlation between toxicity category and descriptor parameters; and (b) shows the correlation strength of key molecular descriptors associated with high risk. Figure 13 To evaluate the performance of the machine learning model for binary risk classification of EPS molecules, this figure shows the validation metrics and prediction performance results on the test set, where (a) is the probability stratification by observation category; (b) is the ROC curve; (c) is the calibration curve with probability distribution plot; and (d) is the precision-recall curve. Figure 14 This is a diagram showing the confusion matrix results of a machine learning model. Figure 15 The results of the SHAP interpretive analysis for machine learning are shown, where (a) is the SHAP value of the top 10 important feature parameters; and (b) is the average absolute value of the SHAP values of the top 10 important feature parameters. Detailed Implementation
[0032] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention. Where specific conditions are not specified in the embodiments, conventional conditions or conditions recommended by the manufacturer shall be followed. Reagents or instruments whose manufacturers are not specified are all conventional products that can be purchased commercially.
[0033] Example This embodiment provides a method for identifying risk factors of extracellular polymers during ultraviolet / peracetic acid disinfection, including the following steps: S1. Total extracellular polymers were extracted from the secondary effluent of a wastewater treatment plant in Jiading District, Shanghai. The extraction of total extracellular polymers was performed using a modified heating-sodium carbonate combined method. The specific steps are as follows: S11. Centrifuge 50 mL of secondary effluent sample at 8000 rpm for 5 min and collect the precipitate. S12. The precipitate was washed three times with a 0.9% NaCl solution. The washed precipitate was then resuspended in the 0.9% NaCl solution to obtain a resuspension. Sodium carbonate was added to the resuspension to make the sodium carbonate concentration in the mixed solution 0.5 wt%. The sodium carbonate-containing mixed solution was then heated at 80 ℃ for 30 min. After heating, the supernatant was collected by centrifugation at 11000 g for 15 min. S13. Filter the obtained supernatant through a 0.22 μm filter membrane to obtain a crude extract of total extracellular polymers; place the crude extract of total extracellular polymers in a dialysis bag with a molecular weight cutoff of 3500 Da and dialyze it with deionized water to obtain a total extracellular polymer sample.
[0034] The total extracellular polymeric substances obtained through the above steps are used for subsequent UV / peracetic acid disinfection reactions and molecular risk factor identification.
[0035] S2. The extracellular polymeric sample obtained in step S1 was placed in a UV / peracetic acid system for sterilization, and samples were collected at different reaction time points. The experimental conditions for sterilization in the UV / peracetic acid system were set as follows: peracetic acid concentration was 10 mg / L, initial pH was 7.0±0.1, and UV dose was 108 mJ / cm². 2 The different reaction time points were 0, 10, 30 and 60 min, and the reaction was divided into three stages: 0-10 min, 10-30 min and 30-60 min.
[0036] S3. Fourier transform ion cyclotron resonance mass spectrometry (FTIR) analysis was performed on the samples obtained at different reaction time points in step S2 to obtain molecular formula data at different reaction time points. Specifically, mixed samples from the UV / peracetic acid system at four time points (0 min, 10 min, 30 min, and 60 min) during the disinfection experiment were subjected to solid-phase extraction and tested. To avoid repeated sampling leading to changes in the volume of the reaction system, each reaction flask corresponded to only one sampling time point, and three parallel samples were set up for the experiment.
[0037] S4. Based on the molecular formula data obtained in step S3, calculate the molecular parameters, match the molecular formula data to SMILES structural information, and further calculate the molecular structure descriptor; specifically as follows: Step S41: Molecular parameters include O / C, H / C, N / C, DBE, NOSC, and AI. mod Kendrick; O / C, H / C, and N / C are used to characterize the elemental composition of a molecule, representing the ratios of oxygen, hydrogen, and nitrogen atoms to carbon atoms, respectively; DBE represents the double bond equivalent of a molecule, used to reflect the degree of unsaturation, calculated as (2C + 2-H + N) / 2; NOSC represents the oxidation state of saturated carbon in a molecule, used to reflect the redox state, calculated as 4 - [(4C + H-3 N-2O-2S) / C]; AI mod The corrosion index represents the molecular value and is used to characterize the aromaticity of molecules. It is calculated as [1 + C - 0.5O - 0.5(N + H)] / (C - 0.5O - S - N). Kendrick's index is used to characterize the transformation relationship of homologues. The above parameters can be used to describe the oxidation, cleavage and structural recombination characteristics of extracellular polymers during sterilization at the molecular level. Step S42: Match the molecular formula data obtained in step S3 to SMILES structural information. Use the SMILES structural information as input value and calculate the molecular structure descriptor according to the code encapsulated in the RDKit package in Python. The descriptors MolLogP, TPSA, FormalCharge, NumRotatableBonds, RingCount, BertzCT, and ExactMassformula are used to characterize the molecule's hydrophobicity, polarity, formal charge, structural flexibility, number of ring structures, structural complexity, and precise mass, respectively, providing feature input for subsequent toxicity risk identification.
[0038] S5. Construct precursor-product molecular pairs based on the molecular formula changes in three adjacent reaction stages: 0-10 min, 10-30 min, and 30-60 min. Identify the molecular transformation type of the extracellular polymer during UV / peracetic acid sterilization based on the mass difference relationship. For example, in the 0-10 min stage, molecular formulas present at 0 min but absent at 10 min are classified as precursors for this stage, while those absent at 0 min but present at 10 min are products for this stage. By comparing the mass difference relationships between these precursors and products, the molecular transformation type of this reaction stage can be determined. This step allows for the identification of the main molecular transformation pathways of the extracellular polymer in different reaction stages. These molecular transformation types include oxygen reactions, cleavage reactions, dealkylation reactions, sulfur-containing reactions, and nitrogen-containing reactions. The specific reaction type classifications are shown in Table 1. Table 1 O-3 oxygen reaction H-2O-2 oxygen reaction O-2 oxygen reaction H+2O-2 oxygen reaction O-1 oxygen reaction H+2O-1 oxygen reaction H-2O-1 oxygen reaction C + 2H + 4 Dealkylation reaction C + 2H + 2 Dealkylation reaction C+3H+6 Dealkylation reaction C+2H+6 Dealkylation reaction C+3H+4 Dealkylation reaction C+1O+2 Decarboxylation reaction C-1O-2 Decarboxylation reaction C + 2H + 2O + 2 Decarboxylation reaction C + 3H + 2O + 2 Decarboxylation reaction C + 4H + 2O + 4 Decarboxylation reaction C + 3H + 6O + 3 pyrolysis reaction C + 3H + 4O + 2 pyrolysis reaction C + 4H + 8O + 4 pyrolysis reaction C + 4H + 4O + 2 pyrolysis reaction C + 5H + 4O + 2 pyrolysis reaction C + 5H + 10O + 5 pyrolysis reaction C + 5H + 10O + 4 pyrolysis reaction C + 6H + 10O + 5 pyrolysis reaction S-1 Sulfur-containing reactions O-2S-1 Sulfur-containing reactions O-3S-1 Sulfur-containing reactions H-1N-1 Nitrogen-containing reactions H-3N-1 Nitrogen-containing reactions H+2O-2 Other reactions C + 7H + 8O + 2 Other reactions C + 6H + 8O + 2 Other reactions H-2 Other reactions H+2 Other reactions S6. Based on the molecular parameters and molecular structure descriptors obtained in step S4, a binary risk label is constructed by combining toxicity prediction data, and a machine learning risk identification model is established. Furthermore, key risk factors of extracellular polymers in the UV / peracetic acid disinfection process are screened through model interpretability analysis. The specific steps are as follows: S61. Input the SMILES structural information obtained in step S4 into ECOSAR software to obtain the corresponding toxicity prediction data. Then, perform data processing and standardization on the toxicity prediction results, and base them on log... 10 (Concentration) Constructing molecular risk labels. Molecular risks are divided into high-risk and low-risk categories, thus forming a binary risk dataset that can be used for machine learning modeling. The binary risk label is constructed as follows: when log... 10 When the value is ≤2, the corresponding molecule is marked as a high-risk molecule; when log 10 When the value is greater than 2, the corresponding molecule will be marked as a low-risk molecule; S62. The machine learning risk identification model is an XGBoost model. This model uses the molecular parameters and molecular structure descriptors calculated in step S4 as input features, and the high-risk and low-risk binary risk labels from S61 as output labels to identify the risk category of extracellular polymer transformation products. The evaluation metrics for the machine learning risk identification model include one or more of AUC, PR AUC, accuracy, precision, recall, and F1 score. The XGBoost model is trained and its parameters are tuned using hyperparameter optimization and random search methods to improve its performance. S63. Use the SHAP method to perform interpretability analysis on the machine learning model, calculate the contribution of different feature parameters to the model's risk identification results, and screen key risk factors based on the contribution value.
[0039] like Figure 1-3 As shown, Figure 1 The variations in different basic molecular parameters in the examples are shown. Figure 1 As shown in (a)-(c), AI mod The changes clearly reflect the transformation trend of aromatic structures. In the 0-10 min stage, EPS precursor molecules are mainly distributed in higher AI regions. mod The region indicates that the system still retains a significant amount of aromatic or aromatic-like structures; while the overall product molecule concentration decreases, but the magnitude of the change is relatively small. Within the 10-30 min range, the product's AI... mod The value decreased significantly, and the higher aromatic components almost completely disappeared. mod The average value decreased by 95.7%, and the distribution density shifted significantly to the low-value region, indicating that the aromatic ring structure underwent severe fracture. AI in the 30-60 min stage. mod The value tends to stabilize, indicating that the aromatic structure has been basically completely transformed.
[0040] Figure 2 (a)-(c) show the distribution relationship between H / C and Mass, indicating that the saturation of EPS molecules gradually increases over time.
[0041] Figure 3 (a)-(c) further analyzed the evolution of redox states, clearly demonstrating the saturation transformation pathway of EPS. The results showed that within 0-10 min, the proportion of reduced unsaturated compounds (region B) decreased by 7.9%, while the proportion of reduced saturated compounds (region C) increased by 7.8%, clearly indicating a trend of structural saturation dominating in the initial stage of the reaction, further confirming the reaction characteristic of preferential breakage of aromatic structures. Within the 10-30 min stage, the proportion of region C decreased significantly (-22.4%), while the proportion of oxidized saturated compounds (region D) increased significantly (+17.4%), indicating that the introduction of oxidized functional groups became the dominant change process in this stage. Within 30-60 min, region D only increased slightly (+3.2%), while regions A and B decreased by 1.7% and 15.9%, respectively, further illustrating that the changes in EPS tended to stabilize within this stage.
[0042] Figure 4-6 The Kendrick mass difference distribution of EPS at different stages in Example 4 is shown. The transformation of EPS molecules is preliminarily confirmed from three perspectives. Similar molecular transformation angles can be clearly observed between different homologues. Figure 4 (a)-(c) show the results of COO-related homologous transformations. The properties and background of the molecules changed over time in the three time stages. In the representative C19 to C23 homologous sequences, the 0 to 10 min stage showed that the relatively reduced precursor was gradually transformed into products with higher O / C and NOSC ratios, indicating that carboxyl-related functionalization had already occurred in the early stage of the reaction. Figure 5 CH2 in (a)-(c), and Figure 6 The Kendrick mass defect diagrams of O in (a)-(c) further indicate that this stage is not a single process, but rather a simultaneous occurrence of homologous chain rearrangement, initial oxidation, and the consumption of aromatic and unsaturated structures.
[0043] Figure 7 This shows the percentage of specific reaction types occurring in EPS molecules at different stages. The results indicate that oxidation reactions are dominant in the 0-10 min and 10-30 min stages, accounting for 21.3% and 22.4% of the total conversion, respectively. However, in the 30-60 min stage, the dominant reaction shifts to cracking, accounting for 24.7%. Specifically, the +2H-2O reaction is the largest in the 0-10 min and 10-30 min stages, at 5.8% and 7.2%, respectively; while the +C3H4 reaction is the largest in the 30-60 min stage, accounting for 5.2%.
[0044] Figure 8 This study illustrates the toxicity risk distribution of EPS molecules at different stages. The results show that within 0-10 min, the proportion of high-risk molecules decreased from 55.8% to 52.8%; within 10-30 min, it decreased from 52.8% to 41.7%; and within 30-60 min, it decreased from 41.7% to 34.0%. Conversely, the proportion of low-risk molecules increased by 2.9%, 11.1%, and 7.8%, respectively. This result indicates that UV / PAA does not simply alter the composition of EPS, but rather causes low-risk molecules to gradually accumulate within the residual molecular system. Sankey diagram analysis of the identified precursor-product pairs further supports this explanation. During the 0–10 min period, the oxidation reaction received 125 high-risk precursors and 74 low-risk precursors, but produced only 91 high-risk products and 108 low-risk products. During the 10–30 min period, it reduced the number of high-risk precursors from 217 to 147; and during the 30–60 min period, it further reduced the number of high-risk precursors from 123 to 74. These quantitative changes indicate that oxidation is the primary cause of the expansion of low-risk components during UV / peracetic acid treatment.
[0045] Figure 9-11This further demonstrates that changes in the molecular properties of EPS are closely related to a reduction in its toxicity. Specifically, the risk-reduced group typically exhibits increased O / C, TPSA, and NOSC, and decreased MolLogP, indicating that increased molecular oxidation, enhanced polarity, and decreased hydrophobicity are important molecular bases for the reduced risk of extracellular polymeric conversion products. Conversely, the risk-increased group shows no significant decrease in MolLogP, and TPSA generally shows a decreasing trend, suggesting that some conversion products may exhibit higher risk due to retained hydrophobicity or decreased polarity. BertzCT and FormalCharge show significant overlap among different risk change types, indicating their limited ability to distinguish independently and making them more suitable as auxiliary molecular descriptors in machine learning models. Overall, these results preliminarily demonstrate that O / C, MolLogP, and TPSA are the core factors for identifying the molecular risk of EPS during UV / peracetic acid disinfection. Figure 12 The findings further support this view. Correlation analysis showed that low-risk molecules were closely correlated with O / C, MolLogP, and TPSA. Overall, O / C showed the highest positive correlation (0.479), while MolLogP showed the highest negative correlation (-0.433). In conclusion, the detoxification process is not based on random changes, but rather on a directed promotion of the formation of more oxidizing and polar structures.
[0046] Figure 13 The performance of the machine learning model was demonstrated. The results showed that the AUC and PRAUC of the XGBoost model on the test set were both 0.979, and the accuracy, F1 score, precision and recall were 0.934, 0.937, 0.926 and 0.948, respectively, indicating that the method can reliably distinguish between high-risk and low-risk molecules. Figure 14 The conclusions further confirm that the model performs well, with its confusion matrix showing excellent performance, and it can effectively distinguish between different risk labels and their results.
[0047] Figure 15 The results show the key risk factors identified by the machine learning model. Among them, SHAP analysis shows that O / C, MolLogP and TPSA are the three features that contribute the most, with average absolute SHAP values of 0.686, 0.624 and 0.400, respectively, which explain 10.5%, 9.5% and 6.1% of the risk, respectively. This indicates that risk identification is mainly controlled by oxidation degree, hydrophobicity and polarity.
[0048] The embodiments described above are merely examples of several implementations of the present invention, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these modifications and improvements all fall within the scope of protection of the present invention.
Claims
1. A method for identifying risk factors of extracellular polymers during UV / peracetic acid disinfection, characterized in that, Includes the following steps: S1. Extract total extracellular polymers from the secondary effluent of a wastewater treatment plant to obtain an extracellular polymer sample; S2. The extracellular polymer sample obtained in step S1 was placed in a UV / peracetic acid system for sterilization, and samples were collected at different reaction time points; S3. Perform Fourier transform ion cyclotron resonance mass spectrometry analysis on the samples obtained at different reaction time points in step S2 to obtain molecular formula data at different reaction time points; S4. Based on the molecular formula data obtained in step S3, calculate the molecular parameters, match the molecular formula data to SMILES structural information, and further calculate the molecular structure descriptor; S5. Construct precursor-product molecular pairs based on the changes in molecular formula between adjacent reaction stages, and identify the molecular transformation type of extracellular polymers during UV / peracetic acid sterilization based on the mass difference relationship; S6. Based on the molecular parameters and molecular structure descriptors obtained in step S4, construct a binary risk label by combining toxicity prediction data, and establish a machine learning risk identification model. Further, screen key risk factors of extracellular polymers in the UV / peracetic acid disinfection process through model interpretability analysis.
2. The method according to claim 1, characterized in that, In step S1, the total extracellular polymers are extracted using a modified heating-sodium carbonate combined method: the secondary effluent sample is centrifuged and enriched, the precipitate is washed with NaCl solution and resuspended, sodium carbonate is added to the resuspended solution and heated, centrifuged and filtered to obtain a crude extract of total extracellular polymers, and then the crude extract of extracellular polymers is dialyzed to obtain a total extracellular polymer sample.
3. The method according to claim 2, characterized in that, The NaCl solution has a mass concentration of 0.9%; the sodium carbonate concentration in the resuspension is 0.5 wt%; the heating temperature is 55–85 ℃, and the heating time is 20–35 min; the dialysis membrane used for dialysis has a molecular cutoff of 3500 Da.
4. The method according to claim 1, characterized in that, The experimental conditions for step S2, where the extracellular polymeric sample was sterilized in a UV / peracetic acid system, were set as follows: peracetic acid concentration of 10 mg / L, initial pH of 7.0 ± 0.1, and UV dose of 108 mJ / cm². 2 The different reaction time points in step S3 are 0, 10, 30 and 60 min, and are divided into three reaction stages: 0-10 min, 10-30 min and 30-60 min.
5. The method according to claim 1, characterized in that, The molecular parameters in step S4 include O / C, H / C, N / C, DBE, NOSC, and AI. mod The descriptors include MolLogP, TPSA, FormalCharge, NumRotatableBonds, RingCount, BertzCT, and ExactMassformula.
6. The identification method steps according to claim 1, characterized in that, The conversion type in step S5 includes one or more of oxidation reactions, cracking reactions, dealkylation reactions, sulfur-containing reactions, and nitrogen-containing reactions.
7. The identification method according to claim 1, characterized in that, In step S6, the toxicity prediction data is calculated using ECOSAR software, and the concentration values in the toxicity prediction data are converted to log10 to construct high-risk and low-risk binary risk labels.
8. The identification method according to claim 7, characterized in that, The binary risk label is constructed as follows: when log 10 When the value is ≤2, the corresponding molecule is marked as a high-risk molecule; when log 10 When the value is greater than 2, the corresponding molecule is marked as a low-risk molecule.
9. The identification method according to claim 8, characterized in that, In step S6, the machine learning risk identification model is an XGBoost model. The model uses the molecular parameters and molecular structure descriptors calculated in step S4 as input features and high-risk and low-risk binary risk labels as output labels to identify the risk category of extracellular polymer transformation products. Its evaluation metrics include one or more of AUC, PR AUC, accuracy, precision, recall, and F1 score.
10. The identification method according to claim 1, characterized in that, In step S6, the model interpretability analysis uses the SHAP method to screen key risk factors by calculating the contribution of different input features to the machine learning output.