A rapid method for predicting the trichloroacetic acid formation potential during water disinfection.
By constructing an XGBoost model based on molecular reactivity and combining it with quantum chemical calculations to obtain key molecular characteristic parameters, the problem of time-consuming, labor-intensive, and low-accuracy prediction of trichloroacetic acid (TCA) generation potential in existing technologies has been solved. This enables rapid, accurate, and low-cost prediction of TCA generation potential, which is suitable for screening and risk warning of DOM precursors in the field of drinking water treatment.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-05
- Publication Date
- 2026-04-03
AI Technical Summary
Existing technologies for predicting the trichloroacetic acid (TCAAFP) generation potential during water disinfection are time-consuming, labor-intensive, costly, have low accuracy, and have limited applicability, making it difficult to meet the needs of real-time monitoring and large-scale screening of DOM precursors in drinking water treatment.
Using a molecular reactivity-based approach combined with high-precision machine learning algorithms, this study constructs an extreme gradient boosting (XGBoost) machine learning algorithm and a Bayesian optimization algorithm. By utilizing quantum chemical calculations to obtain key molecular characteristic parameters, a model for predicting the trichloroacetic acid generation potential is built, enabling rapid prediction without complex experimental operations or high-precision detection equipment.
It achieves rapid, accurate, and low-cost prediction of trichloroacetic acid (TCA) formation potential, has a wide range of applications, and can meet the needs of rapid screening of DOM precursors and risk warning of disinfection byproducts in the field of drinking water treatment. The model has excellent performance and strong stability.
Smart Images

Figure CN121459992B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of environmental engineering technology, specifically relating to a rapid prediction method for the trichloroacetic acid generation potential during water disinfection. Background Technology
[0002] Chlorine disinfection, due to its advantages of low cost, high sterilization efficiency, and ease of operation, has become the mainstream disinfection technology in the global drinking water treatment field. However, chlorine disinfectants react with dissolved organic matter (DOM) in water to generate disinfection byproducts (DBPs), among which trichloroacetic acid (TCAA) is one of the most significant byproducts. According to the "Standards for Drinking Water Quality" (GB5749-2022) and the classification of the International Agency for Research on Cancer (IARC), TCAA is classified as a Group 2B potential carcinogen, with clear carcinogenic, teratogenic, and mutagenic risks. Long-term ingestion of drinking water containing TCAA can cause irreversible damage to the liver, kidneys, and nervous system, seriously threatening public health. Therefore, it is crucial to rapidly and accurately predict the formation potential of TCAA during chlorine disinfection. FP This is a core technological requirement for optimizing drinking water disinfection process parameters, controlling DOM precursors, and providing early warnings for water quality safety. It is of great practical significance for ensuring the safety of drinking water at the point of use.
[0003] Currently, TCAA FPThe prediction methods for TCAA precursors are mainly divided into traditional experimental analysis methods and empirical model prediction methods. Both types of methods have significant technical defects and are difficult to meet the needs of practical applications. The limitations of traditional experimental analysis methods: TCAAFP determination requires a process involving water sample collection, DOM extraction, chlorine disinfection simulation, TCAA separation and purification (e.g., solid-phase extraction), and detection (e.g., gas chromatography-mass spectrometry). This process consumes large amounts of reagents (e.g., methanol, dichloromethane) and samples (≥1 L of water sample per experiment), is cumbersome (total time ≥48 h), and relies on high-precision detection equipment (equipment cost ≥500,000 RMB). It cannot meet the needs of real-time monitoring and rapid decision-making in drinking water treatment, and is especially unsuitable for large-scale DOM precursor screening scenarios. The limitations of empirical model prediction methods: Early empirical models were built based on limited datasets (sample size usually ≤100), relying only on static molecular structural parameters (e.g., atomic hydrogen-to-carbon ratio H:C, molecular weight), lacking consideration of the complex molecular structure of DOM and its reaction mechanism with chlorine, resulting in narrow model applicability and low prediction accuracy. For example, Luilo et al. (Luilo GB, Cabaniss S E. QSPR for predicting chloroform formation in drinking water disinfection. SAR and QSAR in Environmental Research, 2011, 22 (5-6): 489-504) used a multiple linear regression algorithm to construct TCAA using structural descriptors such as H:C. FP Predictive models, although externally validated coefficient of determination Q 2 =0.63 meets the minimum statistical requirement ( Q 2 >0.50), but root mean square error RMSE The value is 0.51 log units, and the predicted value deviates non-linearly from the experimental value. The core reason is that traditional structural descriptors cannot characterize the active sites and reaction barriers of the organic precursors reacting with chlorine, and are difficult to reflect the essential reactive characteristics of TCAA formation.
[0004] In recent years, the rapid development of computational chemistry and machine learning (ML) technologies has driven the TCAA FPUpgrades to the prediction model. Cordero et al. (Cordero JA, He K, Janya K, et al. Predicting formation of haloacetic acids by chlorination of organic compounds using machine-learning-assisted quantitative structure-activity relationships. Journal of Hazardous Materials, 2021, 408: 124466) used over 1800 two-dimensional and three-dimensional molecular descriptors generated by the open-source software Mordred. After rigorous screening, they constructed a prediction model using the Random Forest (RF) algorithm. Although this study expanded the dataset size to over 200 samples and improved the overall performance of the model, external validation performance metrics ( R 2 ext =0.65, RMSE ext =1.07) is still not ideal. This indicates that despite using a larger number of structural descriptors and more complex nonlinear ML algorithms, the improvement in model performance remains limited.
[0005] In summary, the existing TCAA FP Predictive techniques suffer from three major technical shortcomings: time-consuming and laborious experimental methods, insufficient accuracy of empirical models, and lack of reactive characterization. Overcoming these existing technical bottlenecks is a crucial issue we face today. Summary of the Invention
[0006] The main objective of this invention is to propose a rapid prediction method for the formation potential of trichloroacetic acid (TCAA) during water disinfection. This rapid prediction method is based on molecular reactivity and combines it with a high-precision machine learning algorithm. It eliminates the need for complex experimental procedures and high-precision detection equipment, allowing direct calculation of key parameters through molecular structure to determine the TCAA formation potential of organic compounds in chlorine disinfection reactions. FP It offers rapid prediction of values and boasts advantages such as ease of operation, low cost, high stability, and wide applicability, meeting the practical needs of rapid screening of DOM precursors and risk warning of disinfection byproducts in the field of drinking water treatment.
[0007] This invention provides a rapid method for predicting the trichloroacetic acid generation potential during water disinfection, comprising the following steps:
[0008] (1) Construct a dataset and divide the dataset into a training set and a test set, wherein the training set and the test set each independently contain multiple organic compounds and the trichloroacetic acid generation potential value and key molecular characteristic parameters corresponding to each organic compound;
[0009] (2) An extreme gradient boosting machine learning algorithm was adopted, and a trichloroacetic acid generation potential prediction model was constructed based on the trichloroacetic acid generation potential value and key molecular feature parameters of the organic compound. The accuracy of the trichloroacetic acid generation potential prediction model was verified.
[0010] (3) The Bayesian optimization algorithm was used to optimize the trichloroacetic acid generation potential prediction model to obtain the optimal trichloroacetic acid generation potential prediction model;
[0011] (4) Input the key molecular characteristic parameters of the target compound into the optimal prediction model for the trichloroacetic acid generation potential to obtain the trichloroacetic acid generation potential value corresponding to the target compound.
[0012] In some embodiments, the organic compound includes at least one of phenols, carboxylic acids, esters, halogenated compounds, amines, nitro compounds, phosphorus-containing compounds, aldehydes, ketones, ethers, and heterocyclic compounds.
[0013] In some embodiments, the number of organic compounds is 291.
[0014] In some embodiments, the determination conditions for the trichloroacetic acid generation potential of each organic compound are the same, including a chlorine dosage of ≥5 mg / L in the reaction system of the organic compound and chlorine disinfectant, a pH value of 7.0~8.0, and complete reaction of chlorine with the organic compound.
[0015] In some embodiments, the trichloroacetic acid formation potential value is distributed in the range of 0 mol / mol to 0.88 mol / mol.
[0016] In some embodiments, the key molecular characteristic parameter includes the square root of the ring activation index √ RAI The most positive electrostatic charge of hydrogen atoms in the molecule q H + The most negative electrostatic charge of atoms in a molecule q - The maximum positive static charge of carbon atoms in the molecule q C + The maximum negative electrostatic charge of carbon atoms in the molecule q C - and the highest occupied molecular orbital energy E HOMO .
[0017] In some embodiments, the key molecular feature parameters are selected based on the mechanism analysis of the reaction between each organic compound and chlorine disinfectant, combined with the recursive feature elimination method contributed by SHAP.
[0018] In some embodiments, step (2), the construction and accuracy verification of the trichloroacetic acid generation potential prediction model includes: training the model using the training set to establish the trichloroacetic acid generation potential prediction model; and verifying the accuracy of the trichloroacetic acid generation potential prediction model using the test set.
[0019] In some embodiments, the coefficient of determination for evaluating model accuracy is used. R 2 and root mean square error RMSE As a statistical indicator to characterize the model's fit performance, and using the square of the predicted correlation coefficient... Q 2 Characterize the predictive performance of the model; where, Q 2 >0.8.
[0020] In some embodiments, the ratio of the data volume of the training set to the data volume of the test set is set to 8:2, and the sum of the ratios of the data volume of the training set and the test set is always 1.
[0021] In some embodiments, in step (3), the key hyperparameters of the optimal prediction model for trichloroacetic acid generation potential include: tree depth = 9, learning rate = 0.27, subsample weight = 2, sample ratio = 0.82, feature ratio = 0.88, L1 regularization coefficient = 6.36, L2 regularization coefficient = 1.73, and minimum loss reduction required for node splitting = 0.019.
[0022] In some embodiments, the rapid prediction method further includes: comprehensively evaluating the fitting ability, stability, and predictive ability of the optimal prediction model for the trichloroacetic acid generation potential using three methods: fitting performance analysis, simulated external validation, and one-to-one cross-validation.
[0023] The rapid prediction method provided by this invention focuses on molecular reactivity and combines it with high-precision machine learning algorithms. It eliminates the need for complex experimental procedures and high-precision detection equipment, directly obtaining key parameters through molecular structure calculations to achieve TCAA (Total Chemical Acetate) in chlorination disinfection reactions of organic compounds. FP It offers rapid prediction of values and boasts advantages such as ease of operation, low cost, high stability, and wide applicability, meeting the practical needs of rapid screening of DOM precursors and risk warning of disinfection byproducts in the field of drinking water treatment.
[0024] This invention significantly surpasses existing TCAA techniques through a combination of "reactive feature parameters + XGBoost algorithm".FP Predicting technological bottlenecks has the following core beneficial effects:
[0025] 1. High prediction efficiency and low cost: No complex experiments are required; six key characteristic parameters can be obtained solely through molecular structure calculations, enabling rapid prediction of TCAA for tens of thousands of organic compounds. FP This significantly reduces detection costs (by more than 90% compared to traditional experimental methods), meeting the needs of large-scale DOM precursor screening.
[0026] 2. Molecular characteristic parameters are easy to obtain: q H + , q - , q C + , q C - , E HOMO It can be directly extracted using mature Gaussian16 software combined with the DFT algorithm; √ RAI It can be calculated based on clear scoring rules, and the parameter acquisition process is standardized and easy to operate, without relying on special equipment or complex calculations.
[0027] 3. Excellent model performance and strong stability: model fit R 2 =0.83、 RMSE =0.04, simulating external verification Q 2 =0.71、 RMSE =0.03, LOO-CV verification Q 2 CV =0.71、 RMSE CV =0.04, which shows statistically significant performance superior to existing models, with no systematic error and strong stability.
[0028] 4. Clearly defined application domain and broad applicability: The model is clearly applicable to multiple categories of organic compounds with ST ≥ 0.25, covering common DOM precursors in drinking water, and can meet the TCAA requirements in different scenarios. FP Forecast demand.
[0029] 5. Compliant with standards and supporting risk assessment: The model establishment and validation strictly follow the QSPR model development and use guidelines stipulated by the Organization for Economic Cooperation and Development (OECD). The accuracy and reliability of the prediction results are guaranteed, and they can serve as important basic data for the health risk assessment of organic compounds, providing core technical support for the optimization of drinking water disinfection processes and the early warning of disinfection byproduct risks.
[0030] The above description is merely an overview of the technical solution of the present invention. In order to better understand the technical means of the present invention and to implement it in accordance with the contents of the specification, and in order to make the above and other objects, features and advantages of the present invention more apparent and understandable, specific embodiments of the present invention are described below. Attached Figure Description
[0031] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the invention. In the drawings:
[0032] Figure 1 TCAA is the optimal prediction model for the trichloroacetic acid formation potential in the embodiments of the present invention. FP A graph showing the fit between predicted and experimental values;
[0033] Figure 2 The similarity threshold in the embodiments of the present invention ( S cutoff The similarity between compounds in the training and test sets and the delocalization of compounds in the test set when the similarity is 0.25. Detailed Implementation
[0034] Exemplary embodiments of the present invention will now be described in more detail with reference to specific examples. It should be understood that the present invention can be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this invention will be thorough and complete, and will fully convey the scope of the invention to those skilled in the art.
[0035] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains; the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the invention; the terms “comprising” and “having”, and any variations thereof, in the specification, claims, and foregoing description of the invention, are intended to cover non-exclusive inclusion.
[0036] In the description of the embodiments of the present invention, the technical terms "first" and "second" are used only to distinguish different objects and should not be construed as indicating or implying relative importance or implicitly indicating the number, specific order or primary and secondary relationship of the indicated technical features.
[0037] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of the invention. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0038] In the description of the embodiments of this invention, the term "and / or" is merely a description of the association relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists, A and B exist simultaneously, and B exists. Additionally, the character " / " in this document generally indicates that the preceding and following related objects have an "or" relationship.
[0039] In the description of the embodiments of the present invention, the term "multiple" refers to two or more (including two), similarly, "multiple groups" refers to two or more (including two groups), and "multiple pieces" refers to two or more (including two pieces).
[0040] In order to overcome the existing organic compound TCAA FP Addressing the technical shortcomings of traditional prediction methods, such as "long experimental time, high cost, low model accuracy, and narrow applicability," this invention utilizes quantum chemical calculations to obtain key electronic structure and energy parameters that accurately reflect the reactivity of organic molecules with chlorine, constructing a "reactivity-oriented" high-dimensional molecular feature system. Furthermore, it employs a descriptor screening strategy combining SHAP interpretive analysis and recursive feature elimination (SHAP-RFE), and constructs a TCAA based on the XGBoost algorithm. FP High-precision prediction model for organic compound TCAA FP Rapid and accurate prediction provides core technological support for ensuring drinking water safety.
[0041] This invention provides a rapid method for predicting the potential for trichloroacetic acid formation during water disinfection, which is carried out according to the following steps.
[0042] (a) Obtaining high-quality TCAA FP Data, building TCAA FP Dataset.
[0043] In this embodiment of the invention, multiple organic compounds and their respective trichloroacetic acid (TCAA) generation potential values were obtained. In other words, the constructed TCAA... FP The dataset contains multiple organic compounds, and each organic compound has a corresponding trichloroacetic acid generation potential value.
[0044] Generally speaking, high-quality experimental data with significant differences in compound structures are the core foundation for improving model prediction accuracy and expanding the model's application domain. Based on this, this invention screens TCAAs. FP Data Construction TCAA FP The core principles of the dataset are: authoritative data sources, standardized measurement conditions, and diverse compound types.
[0045] In this embodiment of the invention, TCAA measured using standard methods as reported in the literature is first selected. FP Experimental values ensure the authenticity and reliability of the data.
[0046] In some embodiments, the entire TCAA FP The dataset includes 291 organic compounds, covering typical structures such as phenols, carboxylic acids, aldehydes, ketones, esters, ethers, amines, halogenated compounds, phosphorus-containing compounds, heterocycles, and nitro compounds, as well as other structurally similar compounds.
[0047] In this invention, TCAA is modified to eliminate differences in experimental conditions. FP Data interference, TCAA of all compounds FP The experimental values were measured under the same conditions. Specifically, TCM FP The experimental values were determined under standard conditions: the chlorine dosage was ≥5 mg / L (excess chlorine) in the reaction system of organic compounds and chlorine disinfectant, the pH value was 7.0~8.0, and the chlorine and organic compounds reacted completely (ensuring that the reaction between chlorine and organic matter reached equilibrium).
[0048] In this embodiment of the invention, TCAA FP The experimental values ranged from 0 mol / mol (no TCAA formation) to 0.88 mol / mol (high TCAA formation potential).
[0049] (ii) Screening and calculation of key molecular characteristic parameters.
[0050] Existing technologies, whether open-source general-purpose structure descriptors or traditional empirical descriptors, primarily characterize the static structural features of molecules. Their ability to capture and quantify the key reactivity parameters that determine TCAA formation is severely lacking, resulting in models that cannot establish a "molecular feature-reactivity-TCAA" model. FP The direct correlation between "and" makes it difficult to improve prediction accuracy.
[0051] This invention overcomes the limitations of existing technologies that rely on "static molecular structure parameters." Based on the reaction mechanism of organic compounds with chlorine disinfectants and combined with the screening results of the SHAP-RFE algorithm (Shape-Based Recursive Feature Elimination), it ultimately determines key feature parameters that can accurately characterize molecular reactivity, specifically including the square root of the ring activation index.RAI The most positive electrostatic charge of hydrogen atoms in the molecule q H + The most negative electrostatic charge of atoms in a molecule q - The maximum positive static charge of carbon atoms in the molecule q C + The maximum negative electrostatic charge of carbon atoms in the molecule q C - and the highest occupied molecular orbital energy E HOMO A "reactivity-guided" molecular characteristic system was constructed, and the specific parameters and calculation methods are as follows:
[0052] Quantum chemical descriptor calculation method: The density functional theory (DFT) algorithm M062X / 6-311+G(d,p) of Gaussian 16 software is used to optimize the structure of the compound molecule. The most positive electrostatic charge of the hydrogen atom in the molecule can be directly extracted from the output file. q H + The most negative electrostatic charge of atoms in a molecule q - The maximum positive static charge of carbon atoms in the molecule q C + The maximum negative electrostatic charge of carbon atoms in the molecule q C - and the highest occupied molecular orbital energy E HOMO The data value.
[0053] The square root of the cyclic activation index √ RAI Calculation: Based on the scoring rules shown in Table 1, assign scores to the ring structures of the compound and calculate the score of each ring. RAI The value is then calculated to get √. RAI If the molecule contains multiple rings, first calculate the average RAI of each ring (average). RAI Then calculate √ according to formula (1). RAI value:
[0054] √ RAI =√(average) RAI (1)
[0055] Table 1. Conditions RAI Assignment
[0056]
[0057] Note: ED represents the number of strong electron-donating groups (such as -OH, -NH2, -OR) on the ring; AR represents the number of aromatic rings.
[0058] It should be noted that the molecular structure optimization in this embodiment of the invention must ensure that there are no imaginary frequencies. If imaginary frequencies occur, the optimization parameters of the Gaussian 16 software need to be adjusted (such as increasing the number of iterations or adjusting the convergence threshold) until a stable molecular structure is obtained and molecular characteristic parameters are acquired.
[0059] (III) Trichloroacetic acid formation potential (TCAA) FP The establishment and optimization of prediction models.
[0060] Based on the aforementioned key molecular characteristic parameters, this invention constructs TCAA using the Extreme Gradient Boosting (XGBoost) machine learning algorithm. FP Regression prediction models were established between key molecular feature parameters, and SHAP was used to evaluate the descriptors in the final model against TCAA. FP The contribution of the predicted values is characterized, and the specific process is executed using Python 3.12; in other words, the contribution of each feature parameter to TCAA is characterized through SHAP interpretive analysis. FP The contribution weights of the predicted values ensure that the model has both "high accuracy" and "interpretability".
[0061] Specifically, all calculations were implemented using Python 3.12, utilizing open-source libraries such as xgboost 2.1.4, shap0.46.0, rdkit 2024.03.5, and sklearn 1.6.1 to complete model building, feature importance analysis, and performance evaluation; the coefficient of determination was used. R 2 (Measure how well the model fits the data,) R 2 (The closer to 1 the better) and root mean square error RMSE (Measures the deviation between model predictions and experimental values) RMSE (The smaller the better) is used as a statistical indicator to characterize the model's fit performance; the square of the predictive correlation coefficient is used. Q 2 (Measure the model's predictive ability on unknown external data,) Q 2 A score >0.8 indicates an excellent model, representing the model's predictive performance.
[0062] It should be noted that model training and prediction must be performed in a Python 3.12 or equivalent open-source library environment. If the software version is changed, the model performance must be re-verified to avoid version compatibility issues that could lead to increased prediction errors.
[0063] In this embodiment of the invention, the Extreme Gradient Boosting (XGBoost) machine learning algorithm is employed, and a Trichloroacetic Acid (TCAA) generation potential is constructed based on the trichloroacetic acid generation potential values of all organic compounds in the overall dataset and key molecular feature parameters.FP The regression model between the trichloroacetic acid formation potential and the key molecular characteristic parameters was established, which is the trichloroacetic acid formation potential prediction model (also known as the XGBoost prediction model). The accuracy of the trichloroacetic acid formation potential prediction model was then verified.
[0064] The specific implementation process is as follows:
[0065] A dataset was constructed using the trichloroacetic acid (TCA) generation potential values and key molecular characteristic parameters of all obtained organic compounds (e.g., 291 organic compounds). This dataset was divided into a training set and a test set, with a data size ratio of 8:2, and the sum of the training and test set data sizes was always 1. Each training and test set independently contained multiple organic compounds and their corresponding TCA generation potential values and key molecular characteristic parameters; each organic compound corresponded to one TCA generation potential value and six key molecular characteristic parameters.
[0066] A model for predicting the trichloroacetic acid generation potential was established by training the model using the training set.
[0067] The accuracy of the trichloroacetic acid formation potential prediction model was validated using a test set. The coefficient of determination was used to evaluate the model accuracy. R 2 (See Equation (2)) and root mean square error RMSE (See Equation (3)) as a statistical indicator to characterize the model's fit performance, and the square of the predicted correlation coefficient is used. Q 2 (See Equation (4)) characterizes the predictive performance of the model.
[0068] (2)
[0069] (3)
[0070] (4)
[0071] In the formula, and These represent the predicted and experimental values of the trichloroacetic acid formation potential, respectively. This represents the average value of experimental values for the trichloroacetic acid formation potential. This represents the number of samples in the dataset. It can be understood that each sample corresponds to a key molecular feature in an organic compound as input and an experimental value for the trichloroacetic acid formation potential as output.
[0072] Generally speaking, the model R 2 >0.60, Q 2When the value is greater than 0.50, the model has an acceptable goodness of fit. Q 2 A score >0.8 indicates a model with superior performance.
[0073] (iv) Optimize to obtain the optimal prediction model for the trichloroacetic acid generation potential.
[0074] To avoid overfitting and improve generalization ability, a Bayesian optimization algorithm was used to optimize the key hyperparameters of the trichloroacetic acid generation potential prediction model, such as tree depth (max_depth), learning rate (eat), subsample weight (min_child_weight), sample ratio (subsample), feature ratio (colsample_bytree), L1 regularization coefficient (alpha), L2 regularization coefficient (lambda), and minimum loss decrease value required for node splitting (gamma). The optimal prediction model for trichloroacetic acid generation potential and the optimal combination of hyperparameters were obtained, as shown in Table 2.
[0075] Table 2. Hyperparameters of the trichloroacetic acid formation potential prediction model
[0076]
[0077] In this embodiment of the invention, the key molecular characteristic parameters of the target compound (the compound to be tested) are input into the optimal prediction model for the trichloroacetic acid formation potential to obtain the trichloroacetic acid formation potential value corresponding to the target compound.
[0078] (v) Performance evaluation and verification of the optimal prediction model for trichloroacetic acid formation potential.
[0079] The fitting ability, stability, and predictive ability of the optimal prediction model for trichloroacetic acid formation potential were comprehensively evaluated through three methods: fitting performance analysis, simulated external validation, and one-out-of-one cross-validation (LOO-CV).
[0080] Specifically, the fitting performance analysis was conducted using a dataset comprised of the trichloroacetic acid (TCA) formation potential values and key molecular characteristic parameters of 291 organic compounds. The key molecular characteristic parameters of these 291 organic compounds were used as input, and the model-predicted TCA formation potential values were used as output. These values were then compared with the corresponding experimental TCA formation potential values. The final model fitting result was... R 2 =0.75, RMSE =0.06, indicating that the model has good fitting ability, there is no systematic bias between the prediction error and the experimental value, and the statistical performance is significantly better than the existing technology.
[0081] External validation simulation: The dataset consisting of the trichloroacetic acid formation potential values and key molecular characteristic parameters of 291 organic compounds was randomly divided into a training set (containing 233 organic compounds) and a test set (containing 58 organic compounds) in an 8:2 ratio. It can be understood that the training set and test set used when modeling the trichloroacetic acid formation potential prediction model can be used to evaluate and validate the above-mentioned optimal prediction model for trichloroacetic acid formation potential.
[0082] See Figure 1 As shown, the fitting results validated using the training set are: R 2 =0.83 and RMSE =0.04; The fitting result validated using the test set is Q 2 =0.71 and RMSE =0.03. It is evident that the statistical performance of both the training and test sets is very close to that of the model based on the entire set, indicating that the optimal prediction model for trichloroacetic acid (TCAA) generation potential is based on TCAA. FP The essential correlation between the key molecular characteristic parameters and the established correlation is not accidental and exhibits excellent statistical stability.
[0083] One-out cross-validation: The core is iterative training + prediction. The training set (containing 233 organic compounds) is used for evaluation and validation. Each time, one sample is retained as the test set, and the remaining samples are used as a training subset. After the first sample is validated, it is deleted, and this process is repeated. Finally, the prediction results from all rounds are used to calculate... Q 2 and RMSE This represents the LOO-CV performance of the model on the training set. Q 2 CV =0.71, RMSE CV =0.04, further validating the model's stability and predictive consistency under different data distributions, ensuring the model's accuracy for the unknown compound TCAA. FP The reliability of the prediction.
[0084] It is worth mentioning that, in order to avoid applying the model to compounds with large structural differences and causing prediction bias, the optimal prediction model for trichloroacetic acid formation potential is also characterized by application domain characterization in this embodiment of the invention. For example, the application domain of the target compound can be determined before inputting the target compound into the optimal prediction model for trichloroacetic acid formation potential.
[0085] (vi) Model application domain representation.
[0086] In this embodiment of the invention, the "molecular fingerprint similarity threshold method" is used to characterize the applicability of the model, avoiding prediction bias caused by applying the model to compounds with excessively different structures. The specific operation method is as follows:
[0087] (1) Fingerprint calculation: The extended connectivity fingerprint (ECFP4) between the target compound and each compound in the training set is calculated using the rdFingerprintGenerator tool in the rdkit library. Based on this fingerprint, the Tanimoto similarity coefficient (ST) between the target compound and the compounds in the training set is calculated. The larger the ST value, the higher the structural similarity of the compounds.
[0088] (2) Threshold setting: Set the similarity threshold S cutoff =0.25, and a threshold for the number of compounds in the training set is set. N min =1 (to ensure that the target compound has at least one training set sample with a similar structure as a reference).
[0089] (3) Application domain determination: Statistically, the ST value of the target compound in the training set is greater than or equal to the target compound's ST value. S cutoff The number of compounds with ST ≥ 0.25, if this number ≥ N min If the value is ≥1, then the target compound is determined to be within the model application domain.
[0090] (4) Application domain scope: Combining the above determination methods, the TCAA constructed in this invention FP The application domain of the prediction model is defined as: organic compounds with a Tanimoto coefficient (ST) ≥ 0.25 based on the ECFP4 fingerprint of the training set compounds. Specifically, this includes carboxylic acids, phenols, amines, halogenated hydrocarbons, phosphorus-containing compounds, heterocycles, and other structurally similar compounds. The prediction model can predict the TCAA of compounds within this range. FP The value can be accurately predicted.
[0091] Unless otherwise defined, the technical terms used in the following embodiments have the same meaning as commonly understood by those skilled in the art. Unless otherwise specified, the experimental reagents used in the following embodiments are all conventional biochemical reagents; the raw materials, instruments, and equipment used in the following embodiments can all be obtained commercially or through existing methods; unless otherwise specified, the amounts of experimental reagents used are the amounts used in conventional experimental operations; unless otherwise specified, the experimental methods are conventional methods. It should be further noted that the following description is merely exemplary and not a specific limitation of the present invention.
[0092] Example 1: TCAA of 2-methoxyphenyl acetate FP predict
[0093] The 2D molecular structure of 2-methoxyphenyl acetate is shown in formula (I) below:
[0094]
[0095] Formula (1)
[0096] Based on the experimental values of trichloroacetic acid (TCAA) formation potential of 291 organic compounds obtained in this invention and the key molecular characteristic parameters, an optimal prediction model for TCAA formation potential was constructed, and the TCAA of acetic acid was then performed based on this optimal prediction model. FP The prediction process includes:
[0097] Calculation of key molecular characteristic parameters: The DFT M062X / 6-311+G(d, p) algorithm of Gaussian 16 software was used for structure optimization and frequency analysis (ensuring the optimized structure is at the lowest energy point with no imaginary frequencies). Key parameters were extracted from the output file: the most positive electrostatic charge of hydrogen atoms in the molecule. q H + =0.166, the most negative electrostatic charge of an atom in a molecule q - =-0.554, the maximum positive static charge of carbon atoms in the molecule q C + =0.598, the maximum negative static charge of carbon atoms in the molecule q C - = -0.431 and the highest occupied molecular orbital energy E HOMO =-0.289; Referring to Table 1 (Ring Activation Index Scoring Rules), 2-methoxyphenyl acetate contains one benzene ring, and the benzene ring is attached with a methoxy group (activating group, score 0.1), which meets the judgment condition R2. Since the molecule contains only one ring, the average... RAI =0.1, calculated according to formula (1) to obtain √ RAI =0.32.
[0098] TCAA FP Prediction: By inputting the above six key molecular characteristic parameters into the optimal prediction model for the formation potential of trichloroacetic acid, the model outputs the TCAA of 2-methoxyphenyl acetate. FP The predicted value is 0.09700 mol / mol.
[0099] Furthermore, in Example 1, the application domain of the model was determined: using the rdFingerprintGenerator tool from the RDKit library, the ECFP4 fingerprint Tanimoto similarity coefficient (ST) between 2-methoxyphenyl acetate and the training set compounds was calculated based on the SMILES of the compounds. The results showed that its maximum ST with the training set compounds was 0.59, satisfying ST≥0.25. N min The ≥1 criterion further proves that 2-methoxyphenylacetic acid ester is within the model application domain and can be accurately predicted.
[0100] Example 2: TCAA of 1-fluoro-1,1-dichloroethane FP predict
[0101] The 2D molecular structure of 1-fluoro-1,1-dichloroethane is shown in equation (II) below:
[0102]
[0103] Formula (II)
[0104] Based on the experimental values of trichloroacetic acid (TCAA) formation potential of 291 organic compounds obtained in this invention and the key molecular characteristic parameters, an optimal prediction model for TCAA formation potential was constructed, and the TCAA of acetic acid was then performed based on this optimal prediction model. FP The prediction process includes:
[0105] Calculation of key molecular characteristic parameters: The DFT M062X / 6-311+G(d, p) algorithm of Gaussian 16 software was used for structure optimization and frequency analysis (ensuring that the optimized structure is at the lowest energy point and has no imaginary frequencies). Key parameters were extracted from the output file: the most positive electrostatic charge of hydrogen atoms in the molecule. q H + =0.166, the most negative electrostatic charge of an atom in a molecule q - =-0.554, the maximum positive static charge of carbon atoms in the molecule q C + =0.598, the maximum negative static charge of carbon atoms in the molecule q C - = -0.431 and the highest occupied molecular orbital energy E HOMO =-0.289; Referring to Table 1 (Ring Activation Index Scoring Rules), 1-fluoro-1,1-dichloroethane does not contain a benzene ring, therefore all judgment conditions are not met, and the average... RAI =0.00, calculated according to formula (1) √ RAI =0.00.
[0106] TCAA FP Prediction: By inputting the above six key molecular characteristic parameters into the optimal prediction model for the formation potential of trichloroacetic acid, the model outputs the TCAA of 1-fluoro-1,1-dichloroethane. FP The predicted value is 0.00006 mol / mol. This also indicates that the compound produces almost no TCAA during chlorination, consistent with the actual experimental results for "TCAA from haloalkanes". FP The conclusion of "low" is consistent.
[0107] Furthermore, Example 2 also determined the application domain of the model: using the rdFingerprintGenerator tool from the RDKit library, the ECFP4 fingerprint Tanimoto similarity coefficient (ST) between 1-fluoro-1,1-dichloroethane and the compounds in the training set was calculated based on the SMILES of the compounds. The results showed that its maximum ST with the compounds in the training set was 0.77, satisfying ST≥0.25. N min The ≥1 criterion further proves that 1-fluoro-1,1-dichloroethane is within the model's application domain and can be accurately predicted.
[0108] Example 3: TCAA of styrene FP predict
[0109] The 2D molecular structure of styrene is shown in equation (III) below:
[0110]
[0111] Formula (3)
[0112] Based on the experimental values of trichloroacetic acid (TCAA) formation potential of 291 organic compounds obtained in this invention and the key molecular characteristic parameters, an optimal prediction model for TCAA formation potential was constructed, and the TCAA of acetic acid was then performed based on this optimal prediction model. FP The prediction process includes:
[0113] Calculation of key molecular characteristic parameters: The DFT M062X / 6-311+G(d, p) algorithm of Gaussian 16 software was used for structure optimization and frequency analysis (ensuring that the optimized structure is at the lowest energy point and has no imaginary frequencies). Key parameters were extracted from the output file: the most positive electrostatic charge of hydrogen atoms in the molecule. q H + =0.126, the most negative electrostatic charge of an atom in a molecule q - =-0.291, the maximum positive static charge of carbon atoms in the molecule. q C + =0.079, the maximum negative static charge of carbon atoms in the molecule q C- = -0.291 and the highest occupied molecular orbital energy E HOMO =-0.273; Referring to Table 1 (Ring Activation Index Scoring Rules), styrene contains only one benzene ring, with no -OH, -NH2, -OR, or -COOH on the ring, nor any ortho-carboxylic acid or -OH / NH2 ortho-para structures. Therefore, all judgment conditions are not met, and the average score is -0.273. RAI =0.00, calculated according to formula (1) √ RAI =0.00.
[0114] TCAA FP Prediction: By inputting the above six key molecular characteristic parameters into the optimal prediction model for the trichloroacetic acid (TCAA) formation potential, the model outputs the TCAA of styrene. FP The predicted value is 0.00175 mol / mol. This result is consistent with the observation that "alkene compounds are easily oxidized due to the double bonds, TCAA..." FP The prediction results are reliable, showing a pattern of "slightly higher than halogenated alkanes, but lower than phenols".
[0115] Furthermore, Example 3 also determined the application domain of the model: the rdFingerprintGenerator tool from the RDKit library was used to calculate the Tanimoto similarity coefficient (ST) of the ECFP4 fingerprint of styrene and the compounds in the training set based on the SMILES of the compounds. The results showed that its maximum ST with the compounds in the training set was 0.35, satisfying ST≥0.25. N min The ≥1 criterion further proves that styrene is within the model's application domain and can be accurately predicted.
[0116] Example 4: TCAA of 2,3-dimethoxyphenol FP predict
[0117] The 2D molecular structure of 2,3-dimethoxyphenol is shown in formula (iv):
[0118]
[0119] Formula (IV)
[0120] Based on the experimental values of trichloroacetic acid (TCAA) formation potential of 291 organic compounds obtained in this invention and the key molecular characteristic parameters, an optimal prediction model for TCAA formation potential was constructed, and the TCAA of acetic acid was then performed based on this optimal prediction model. FP The prediction process includes:
[0121] Calculation of key molecular characteristic parameters: The DFT M062X / 6-311+G(d, p) algorithm of Gaussian 16 software was used for structure optimization and frequency analysis (ensuring that the optimized structure is at the lowest energy point and has no imaginary frequencies). Key parameters were extracted from the output file: the most positive electrostatic charge of hydrogen atoms in the molecule. q H + =0.339, the most negative electrostatic charge of an atom in a molecule q - =-0.576, the maximum positive static charge of carbon atoms in the molecule q C + =0.328, the maximum negative static charge of carbon atoms in the molecule q C - = -0.172 and the highest occupied molecular orbital energy E HOMO =-0.266; Referring to Table 1 (Ring Activation Index Scoring Rules), 2,3-dimethoxyphenol contains only one benzene ring, with -OH and -OR (alkoxy groups) on the ring, no -NH2, no ortho-carboxylic acid, and no ortho / para combination of -OH and -NH2. Therefore, the judgment condition R5 is used for calculation, and ED:AR=3 corresponds to... RAI =0.5, calculated according to formula (1) to obtain √ RAI =0.71.
[0122] TCAA FP Prediction: By inputting the above six key molecular characteristic parameters into the optimal prediction model for the formation potential of trichloroacetic acid, the model outputs the TCAA of 2,3-dimethoxyphenol. FP The predicted value is 0.10900 mol / mol. This predicted value is higher than that of 2-methoxyphenyl acetate in Example 1 because the dimethoxy group has a stronger activating effect on the benzene ring, and the more activating groups there are, the higher the TCAA concentration. FP The reaction mechanism is consistent with that of "higher" and verifies the rationality of the model.
[0123] Furthermore, in Example 4, the application domain of the model was determined: the rdFingerprintGenerator tool from the RDKit library was used to calculate the ECFP4 fingerprint Tanimoto similarity coefficient (ST) between 2,3-dimethoxyphenol and the compounds in the training set based on the SMILES of the compounds. The results showed that its maximum ST with the compounds in the training set was 0.49, satisfying ST≥0.25. N min The ≥1 criterion further proves that 2,3-dimethoxyphenol is within the model's application domain and can be accurately predicted.
[0124] This invention, by quantifying molecular feature parameters and combining them with machine learning modeling, can efficiently predict compounds in the training set (based on the Tanimoto coefficient similarity of ECFP4 fingerprints). S cutoff (i.e., ST>0.25) TCAA, a multi-class organic compound with similar structures FP The organic compounds mentioned include carboxylic acids, phenols, esters, amines, halogenated hydrocarbons, nitro compounds, aldehydes, ketones, ethers, phosphorus-containing compounds, and heterocyclic compounds, and are applicable to scenarios such as drinking water treatment process optimization, disinfection by-product risk warning, and water quality safety assessment.
[0125] The above description is merely a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A rapid prediction method for the trichloroacetic acid generation potential during water disinfection, characterized in that, Includes the following steps: (1) Construct a dataset and divide the dataset into a training set and a test set. The training set and the test set each independently contain multiple organic compounds and the trichloroacetic acid generation potential value and key molecular characteristic parameters corresponding to each organic compound. The key molecular characteristic parameters include the square root of the ring activation index √ RAI The most positive electrostatic charge of hydrogen atoms in the molecule q H + The most negative electrostatic charge of atoms in a molecule q - The maximum positive static charge of carbon atoms in the molecule q C + The maximum negative electrostatic charge of carbon atoms in the molecule q C - and the highest occupied molecular orbital energy E HOMO ; (2) An extreme gradient boosting machine learning algorithm was adopted, and a trichloroacetic acid generation potential prediction model was constructed based on the trichloroacetic acid generation potential value and key molecular feature parameters of the organic compound. The accuracy of the trichloroacetic acid generation potential prediction model was verified. (3) The Bayesian optimization algorithm was used to optimize the trichloroacetic acid generation potential prediction model to obtain the optimal trichloroacetic acid generation potential prediction model; (4) Input the key molecular characteristic parameters of the target compound into the optimal prediction model for the trichloroacetic acid generation potential to obtain the trichloroacetic acid generation potential value corresponding to the target compound.
2. The rapid prediction method as described in claim 1, characterized in that, The organic compounds include at least one of phenols, carboxylic acids, esters, halogenated compounds, amines, nitro compounds, phosphorus-containing compounds, aldehydes, ketones, ethers, and heterocyclic compounds; and / or, The number of organic compounds is 291.
3. The rapid prediction method as described in claim 1, characterized in that, The trichloroacetic acid generation potential of each of the organic compounds was determined under the same conditions, including a chlorine dosage of ≥5 mg / L in the reaction system of the organic compound and chlorine disinfectant, a pH value of 7.0~8.0, and complete reaction of chlorine with the organic compound.
4. The rapid prediction method as described in claim 1, characterized in that, The data distribution range of the trichloroacetic acid formation potential value is 0 mol / mol to 0.88 mol / mol.
5. The rapid prediction method as described in claim 1, characterized in that, Based on the mechanistic analysis of the reaction between each organic compound and chlorine disinfectant, and combined with the recursive feature elimination method contributed by SHAP, the key molecular feature parameters were selected and obtained.
6. The rapid prediction method as described in claim 1, characterized in that, In step (2), the construction and accuracy verification of the trichloroacetic acid formation potential prediction model include: The training set is used to train the model and establish the trichloroacetic acid generation potential prediction model. The accuracy of the trichloroacetic acid formation potential prediction model was verified using the test set.
7. The rapid prediction method as described in claim 6, characterized in that, The coefficient of determination is used to evaluate the accuracy of the model. R 2 and root mean square error RMSE As a statistical indicator to characterize the model's fit performance, and using the square of the predicted correlation coefficient... Q 2 Characterize the predictive performance of the model; where, Q 2 >0.8; And / or, the ratio of the data volume of the training set to the data volume of the test set is set to 8:2, and the sum of the ratios of the data volume of the training set to the data volume of the test set is always 1.
8. The rapid prediction method as described in claim 6, characterized in that, In step (3), the key hyperparameters of the optimal prediction model for the trichloroacetic acid generation potential include: tree depth = 9, learning rate = 0.27, subsample weight = 2, sample ratio = 0.82, feature ratio = 0.88, L1 regularization coefficient = 6.36, L2 regularization coefficient = 1.73, and minimum loss reduction required for node splitting = 0.
019.
9. The rapid prediction method as described in claim 1, characterized in that, The rapid prediction method also includes: The fitting ability, stability, and predictive ability of the optimal prediction model for the trichloroacetic acid formation potential were comprehensively evaluated using three methods: fitting performance analysis, simulated external validation, and cross-validation with one-to-one elimination.
Citation Information
Patent Citations
Method for predicting concentration of disinfection by-product haloacetic acid in water supply system
CN110287651A
Drinking water disinfection by-product generation prediction method based on reaction kinetic model
CN120833862A