A rapid method for predicting the potential for chloroform formation during water disinfection.
By constructing a trichloromethane generation potential prediction model based on molecular characteristic parameters and the XGBoost algorithm, the problems of time-consuming, labor-intensive, and narrowly applicable traditional methods are solved. This model enables rapid and accurate prediction of trichloromethane generation potential during water treatment and is applied to the field of environmental engineering technology. It relates to a method for rapid prediction of trichloromethane generation potential during water disinfection.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-05
- Publication Date
- 2026-04-03
AI Technical Summary
Existing methods for predicting trichloromethane formation potential (TCMFP) rely on traditional experimental analysis and empirical models, which are time-consuming, labor-intensive, have small datasets, and lack a complete descriptor system. These methods are insufficient to achieve rapid and accurate TCMFP prediction and cannot meet the needs of water treatment process optimization and drinking water safety assurance.
A trichloromethane (TCMFP) formation potential prediction model was constructed using molecular feature parameters and machine learning methods, combining the Extreme Gradient Boosting (XGBoost) machine learning algorithm with SHAP contribution analysis. The model was used to predict TCMFP by obtaining key molecular feature parameters of organic compounds, including the most negative electrostatic potential of molecular surface, chemical hardness, carbonyl index, and aromaticity index. The model was then optimized to improve prediction accuracy.
It enables rapid and accurate prediction of the chloroform generation potential of organic compounds during water disinfection, reduces detection costs, expands the scope of application, meets the needs of large-scale DOM precursor screening, significantly improves model stability and applicability, conforms to QSPR model specifications, and supports drinking water disinfection process optimization and risk assessment.
Smart Images

Figure CN121459991B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of environmental engineering technology, and specifically relates to a rapid prediction method for the trichloromethane generation potential during water disinfection. Background Technology
[0002] Natural water contains various microorganisms, including pathogenic bacteria and viruses, due to pollution from domestic sewage and industrial wastewater. Disinfection kills these pathogens, prevents the spread of waterborne diseases, and protects human health; it is an essential step in drinking water treatment. Chlorine disinfection is widely used globally due to its cost-effectiveness and efficiency. However, while killing pathogens, chlorine disinfectants react with dissolved organic matter (DOM) in water to generate a series of disinfection byproducts (DBPs), among which trichloromethane (TCM) is one of the most significant. TCM poses potential carcinogenic, teratogenic, and mutagenic risks, seriously threatening human health. Furthermore, the amount of TCM generated is closely related to the generation of controlled DBPs such as haloacetic acids and highly toxic nitrogenous DBPs such as haloacetonits. Therefore, rapidly and accurately predicting the TCM generation potential of organic compounds during chlorination disinfection is crucial. FP This is of great significance for optimizing water treatment processes and ensuring drinking water safety. Currently, TCM... FP Prediction methods primarily rely on traditional experimental analysis and empirical models. While these methods can provide some information related to TCM generation, they have several limitations. On the one hand, experimental analysis requires a large number of samples and complex procedures, which is time-consuming and labor-intensive, making it difficult to meet the needs of real-time monitoring and rapid decision-making. On the other hand, empirical models are usually based on limited datasets and lack a deep understanding of the complex molecular structures and reaction mechanisms of dissolved organic compounds, thus limiting the accuracy and applicability of predictions.
[0003] Quantitative structure-property relationships (QSPR) can predict the properties and environmental behavior parameters of compounds based on molecular structure information, providing a basis for developing rapid TCM prediction methods. FP The method provided a good idea. However, current reports on TCM... FPThe QSPR model still has problems in terms of the scope of application of compounds, the predictive ability of the model, and the operability of practical applications. Luilo et al. (Luilo GB, Cabaniss S E. QSPR for predicting chloroformformation in drinking water disinfection. SAR and QSAR in Environmental Research, 2011, 22(5-6): 489-504) based on TCM of 90 compounds FP Experimental measurements were used to construct the TCM using multiple linear regression. FP The prediction model has good fitting performance. Q 2 =0.91, RMSE The model has a low accuracy (0.08), but its small dataset limits its ability to predict tens of thousands of organic compounds. Furthermore, the descriptor system used in this model does not consider electronic effects parameters such as molecular orbital energy levels and electrostatic potential distribution, leading to a prediction bias exceeding 40% for precursors containing conjugated double bonds or electron-withdrawing groups. Bond et al. (Bond T, Graham N. Predicting chloroform production from organic precursors. Water Research, 2017, 124: 167-176) expanded the dataset to 211 compounds and introduced substituent steric effect parameters (…). Q 2 =0.90, RMSE =0.08), but the electron transfer mechanism during the chlorine substitution process was not quantified, which limited the applicability of the model in complex aquatic matrices.
[0004] This shows that the existing TCM FP Predictive models (including traditional empirical models and QSPR models) are limited by issues such as dataset size and incomplete descriptor systems (lack of electronic effect parameters or unquantified electron transfer mechanisms in chlorination processes). This not only limits their applicability but also makes it difficult to simultaneously achieve TCM prediction for organic compounds. FP The ability to accurately predict and quickly assess values cannot meet the practical needs of optimizing water treatment processes and ensuring drinking water safety. Summary of the Invention
[0005] The purpose of this invention is to provide a rapid prediction method for the formation potential of chloroform during water disinfection. This rapid prediction method is based on molecular characteristic parameters and machine learning, and can quickly predict the formation potential of chloroform during the reaction of organic compounds with chlorine disinfectants in water disinfection, thus achieving TCM (trichloromethane formation potential) of organic compounds. FP It offers rapid and accurate value prediction with advantages such as convenience, low cost, ease of use, and high stability.
[0006] This invention provides a rapid method for predicting the trichloromethane formation potential during water disinfection, comprising the following steps:
[0007] (1) Obtain multiple organic compounds and their respective trichloromethane generation potential values;
[0008] (2) Obtain the key molecular characteristic parameters for accurately characterizing the molecular reactivity of each organic compound;
[0009] (3) An extreme gradient boosting machine learning algorithm is adopted, and a chloroform generation potential prediction model is constructed based on the chloroform generation potential value and key molecular characteristic parameters of the organic compound. The accuracy of the chloroform generation potential prediction model is verified.
[0010] (4) The Bayesian optimization algorithm is used to optimize the trichloromethane generation potential prediction model to obtain the optimal trichloromethane generation potential prediction model;
[0011] (5) Input the key molecular characteristic parameters of the target compound into the optimal prediction model for chloroform generation potential to obtain the chloroform generation potential value of the target compound.
[0012] In some embodiments, the organic compound includes at least one of phenols, carboxylic acids, esters, halogenated hydrocarbons, amines, nitro compounds, amides, aldehydes, ketones, ethers, alcohols, heterocyclic compounds, sulfides, cyanides, sugars, and polyhydroxy compounds.
[0013] In some embodiments, the number of organic compounds is 222.
[0014] In some embodiments, the determination conditions for the trichloromethane generation potential of each organic compound are the same, including a chlorine dosage of ≥5 mg / L in the reaction system of the organic compound and chlorine disinfectant, a pH value of 7.0~8.0, and complete reaction of chlorine with the organic compound.
[0015] In some embodiments, the data distribution range of the trichloromethane generation potential value is 0 mol / mol to 0.99 mol / mol.
[0016] In some embodiments, in step (2), the key molecular characteristic parameters include the most negative electrostatic potential of the compound's molecular surface. V s,min Chemical hardness η Carbonyl index CI olefin index Enolizable Score and aromaticity index Aromatic Score .
[0017] In some embodiments, in step (2), the key molecular feature parameters are selected based on the mechanism analysis of the reaction between each organic compound and the chlorine disinfectant and the results of recursive feature elimination based on SHAP contribution.
[0018] In some embodiments, step (3) of constructing and verifying the trichloromethane generation potential prediction model includes: dividing the trichloromethane generation potential values and key molecular characteristic parameter data of all the obtained organic compounds into a training set and a test set; using the training set to train the model and establish the trichloromethane generation potential prediction model; and using the test set to verify the model accuracy of the trichloromethane generation potential prediction model.
[0019] In some embodiments, the coefficient of determination for evaluating model accuracy is used. R 2 and root mean square error RMSE As a statistical indicator to characterize the model's fit performance, and using the square of the predicted correlation coefficient... Q 2 Characterize the predictive performance of the model; where, Q 2 >0.8.
[0020] In some embodiments, the ratio of the data volume of the training set to the data volume of the test set is set to 8:2, and the sum of the ratios of the data volume of the training set and the test set is always 1.
[0021] In some embodiments, in step (4), the key hyperparameters of the optimal prediction model for chloroform generation potential include: tree depth = 3, learning rate = 0.097, subsample weight = 3, sample ratio = 0.65, feature ratio = 0.80, L1 regularization coefficient = 0.014, L2 regularization coefficient = 2.12, and minimum loss reduction required for node splitting = 0.029.
[0022] In some embodiments, the rapid prediction method further includes: comprehensively evaluating the fitting ability, stability, and prediction ability of the optimal prediction model for chloroform generation potential using three methods: fitting performance analysis, simulated external validation, and one-to-one cross-validation.
[0023] The rapid prediction method provided by this invention is based on molecular feature parameters and machine learning methods, which can quickly predict the formation potential of trichloromethane during the reaction of organic compounds with chlorine disinfectants in water disinfection, and realize the TCM of organic compounds. FP It offers rapid and accurate value prediction with advantages such as convenience, low cost, ease of use, and high stability.
[0024] This invention significantly surpasses existing TCMs through a combination of "reactive feature parameters + XGBoost algorithm". FP Predicting technological bottlenecks has the following core beneficial effects:
[0025] 1. High prediction efficiency and low cost: No complex experiments are required; five key molecular characteristic parameters can be obtained solely through molecular structure calculations, enabling rapid prediction of the TCM of tens of thousands of organic compounds. FP This significantly reduces detection costs (by more than 90% compared to traditional experimental methods), meeting the needs of large-scale DOM precursor screening.
[0026] 2. Molecular characteristic parameters are easy to obtain: V s,min and η It can be directly extracted using the mature Gaussian 16 software combined with the DFT algorithm; CI , Enolizable Score and Aromatic Score It can be calculated based on clear scoring rules, and the parameter acquisition process is standardized and easy to operate, without relying on special equipment or complex calculations.
[0027] 3. Excellent model performance and strong stability: model fit R 2 =0.95、 RMSE =0.06, simulating external verification Q 2 =0.94、 RMSE =0.06, LOO-CV verification Q 2 CV =0.81、 RMSE CV =0.09, which shows statistically significant performance superior to existing models, with no systematic error and strong stability.
[0028] 4. Clear application domain and wide applicability: The model is clearly applicable to multiple categories of organic compounds with ST ≥ 0.25, covering common DOM precursors in drinking water, and can meet the TCM requirements in different scenarios. FP Forecast demand.
[0029] 5. Compliant with standards and supporting risk assessment: The model establishment and validation strictly follow the QSPR model development and use guidelines stipulated by the Organization for Economic Cooperation and Development (OECD). The accuracy and reliability of the prediction results are guaranteed, and they can serve as important basic data for the health risk assessment of organic compounds, providing core technical support for the optimization of drinking water disinfection processes and the early warning of disinfection byproduct risks.
[0030] The above description is merely an overview of the technical solution of the present invention. In order to better understand the technical means of the present invention and to implement it in accordance with the contents of the specification, and in order to make the above and other objects, features and advantages of the present invention more apparent and understandable, specific embodiments of the present invention are described below. Attached Figure Description
[0031] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0032] Figure 1 The TCM of the optimal prediction model for trichloromethane formation potential in the embodiments of the present invention FP A graph showing the fit between experimental and predicted values;
[0033] Figure 2 The similarity threshold in the embodiments of the present invention ( S cutoff The similarity between compounds in the training and test sets and the delocalization of compounds in the test set when the similarity is 0.25. Detailed Implementation
[0034] Exemplary embodiments of the invention will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the invention are shown in the drawings, it should be understood that the invention can be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided to enable a more thorough understanding of the invention and to fully convey the scope of the invention to those skilled in the art.
[0035] It should be understood that the terminology used herein is for the purpose of describing particular exemplary embodiments only and is not intended to be limiting. Unless the context clearly indicates otherwise, the singular forms “a,” “an,” and “described” as used herein may also include the plural forms. The terms “comprising,” “including,” “containing,” and “having” are inclusive and therefore indicate the presence of the stated features, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, elements, components, and / or combinations thereof. The method steps, processes, and operations described herein are not construed as requiring them to be performed in a particular order described or illustrated unless the order of performance is explicitly indicated. It should also be understood that additional or alternative steps may be used.
[0036] In the description of the embodiments of this invention, technical terms such as "first" and "second" are used only to distinguish different objects and should not be construed as indicating or implying relative importance or implicitly specifying the number, specific order, or primary and secondary relationship of the indicated technical features. In the description of the embodiments of this invention, "multiple" means two or more, unless otherwise explicitly defined.
[0037] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of the invention. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0038] In the description of the embodiments of this invention, the term "and / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Additionally, the character " / " in this document generally indicates that the preceding and following related objects have an "or" relationship.
[0039] In the description of the embodiments of the present invention, the term "multiple" refers to two or more (including two), similarly, "multiple groups" refers to two or more (including two groups), and "multiple pieces" refers to two or more (including two pieces).
[0040] With the continuous development of computer software and artificial intelligence, machine learning methods are increasingly widely used in model building. For example, a series of high-performance QSPR models have been developed using methods such as Support Vector Machines, Random Forests, and Deep Learning, significantly improving the fitting and prediction capabilities of linear models. Among them, the XGBoost algorithm, based on iterative accumulation of gradient boosting decision trees, has shown outstanding performance. However, the internal workings of machine learning models, including XGBoost, remain largely "black boxes," with the relative contribution of each feature unclear, hindering the interpretation of the model's mechanism. The SHAP method provides an important tool for solving this problem. By calculating the contribution of each feature to the prediction result, it reflects the contribution of each feature to the target feature, thereby opening the "black box" of machine learning.
[0041] This invention utilizes SHAP nested XGBoost technology to develop a method for rapidly and accurately predicting the TCM of organic compounds. FP A new method overcomes the shortcomings of existing technologies, making it applicable for predicting the TCM of a wide range of organic compounds. FP These values will effectively compensate for the lack of property parameters for organic compounds and provide necessary basic data for health risk assessment and drinking water safety assessment of organic compounds.
[0042] This invention provides a rapid method for predicting the potential for chloroform formation during water disinfection, which is carried out according to the following steps.
[0043] (a) Obtaining high-quality TCM FP Data, building TCM FP Dataset.
[0044] In this embodiment of the invention, multiple organic compounds and their respective trichloromethane formation potential values were obtained. In other words, the constructed TCM... FP The dataset contains multiple organic compounds, and each organic compound has a corresponding chloroform generation potential value.
[0045] Generally speaking, high-quality experimental data with significant differences in compound structures are the core foundation for improving model prediction accuracy and expanding the model's application domain. Based on this, this invention constructs a TCM. FP The core principles of the dataset are: authoritative data sources, standardized measurement conditions, and diverse compound types.
[0046] In this embodiment of the invention, the TCM measured using standard methods as reported in the literature is first selected. FP Experimental values ensure the authenticity and reliability of the data.
[0047] In some embodiments, the entire TCM FPThe dataset includes 222 organic compounds, covering typical structures such as phenols, carboxylic acids, esters, halogenated hydrocarbons, amines, nitro compounds, amides, aldehydes, ketones, ethers, alcohols, heterocycles, sulfides, cyanides, sugars, and polyhydroxy compounds, as well as other compounds with similar structures.
[0048] In this invention, TCM is adjusted to eliminate differences in experimental conditions. FP Data interference, TCM of all compounds FP The experimental values were measured under the same conditions. Specifically, TCM FP The experimental values were determined under standard conditions: the chlorine dosage was ≥5 mg / L (excess chlorine), the pH value was 7.0~8.0, and the chlorine and organic compounds reacted completely (ensuring that the reaction between chlorine and organic matter reached equilibrium).
[0049] In this embodiment of the invention, TCM FP The experimental values ranged from 0 mol / mol (no TCM generation) to 0.99 mol / mol (high TCM generation potential).
[0050] (ii) Obtaining key molecular characteristic parameters of the compound.
[0051] In this embodiment of the invention, key molecular characteristic parameters for accurately characterizing the molecular reactivity of each organic compound are obtained.
[0052] In this embodiment of the invention, molecular structure descriptors and quantum chemical descriptors are used to characterize the structure, electronic and energy properties of organic compounds, respectively.
[0053] The optimal descriptor subset for the model is determined based on SHAP values, while the effect of each molecular feature parameter on TCM is quantified. FP . contributions.
[0054] In this embodiment of the invention, based on the reaction mechanism of organic compounds with chlorine disinfectants and the results of recursive feature elimination based on SHAP (SHapley Additive exPlanations) contributions (SHAP-REF), five key molecular characteristic parameters for accurately characterizing molecular reactivity are selected, in order of importance: the most negative electrostatic potential of the compound's molecular surface. V s,min Chemical hardness of compounds η Carbonyl index of compounds CI , compound alkenizability index Enolizable Score Aromaticity index of compounds Aromatic Score Participate in subsequent TCM FP Construction of prediction models.
[0055] Specifically, the quantum chemical descriptor calculation method uses the density functional theory (DFT) algorithm M062X / 6-311+G(d,p) of Gaussian 16 software to perform structure optimization and frequency analysis on all organic compound molecules (e.g., 222 organic compounds), and extracts the highest occupied molecular orbitals from the output file. E HOMO ) and the lowest unoccupied molecular orbital ( E LUMO The chemical hardness of the compound was calculated. η The value is calculated using the formula shown in equation (1) below.
[0056] (1)
[0057] Next, GsGrid 1.7 was used to calculate the most negative electrostatic potential of the compound's molecular surface on the output file. V s,min value.
[0058] It should be noted that the molecular structure optimization in this embodiment of the invention must ensure that there are no imaginary frequencies. If imaginary frequencies occur, the optimization parameters of the Gaussian 16 software need to be adjusted (such as increasing the number of iteration steps, adjusting the convergence threshold, etc.) until a stable molecular structure is obtained and molecular characteristic parameters are obtained.
[0059] Carbonyl index of a compound CI Calculation method: According to the scoring rules in Table 1, assign scores to the carbon atoms in all compounds that meet the conditions, and add them together to obtain the carbonyl index of the compound. CI Value. In other words, the carbonyl index of a compound. CI The value is obtained by adding the various carbon atoms in the compound according to the scoring rules.
[0060] Table 1. Values assigned to various carbon atoms
[0061]
[0062] olefination index of compounds Enolizable Scores Calculation method: Multiply the empirical constants of the three substituents R1, R2, and R3 around the alkeneizable proton of the compound, and list the empirical constant values of different substituents in Table 2 (ignoring the proton itself) to obtain the alkeneability index of the compound. Enolizable Scores Value. In other words, the alkenizability index of a compound. Enolizable Scores The value is obtained by multiplying the empirical constants of the three substituents (R1, R2, and R3) around the olefinic proton, with the proton itself being negligible.
[0063] Table 2. Used for calculating the alkenylability index Enolizable Score Empirical constants of each substituent
[0064]
[0065] Aromaticity index of cyclic compounds Aromatic Score Calculation method:
[0066] (1) Set the baseline score: Assign an initial baseline score according to the type of aromatic ring, with 1.0 for benzene ring and 0.2 for heterocyclic ring.
[0067] (2) Determine the substituent score and C1: Identify all substituents on the ring and assign them independent scores (see Table 3). Select the carbon atom containing the substituent with the highest score as C1. If multiple substituents have the same score, the system with the higher initial calculated score should be selected first.
[0068] (3) Substituent numbering rules: The adjacent substituent is preferentially numbered as C2 (clockwise adjacent to C1), and the next adjacent substituent is numbered as C3 (preferred to C5); if the numbering ambiguity is not resolved, the arrangement that maximizes the interaction product is selected.
[0069] (4) Correction of interaction: Starting from C1, traverse the substituents in a clockwise direction, calculate the interaction correction factor between adjacent substituents (such as the electron-donating-electron-withdrawing cooperating effect coefficient) and multiply them together; additionally include the inter-substituent correction factor that does not involve C1.
[0070] (5) Output comprehensive score: Multiply the benchmark score, the product of the independent scores of substituents, and the total product of the correction factor to obtain the aromaticity index. Aromatic Score If there are no substituents, the baseline score is output directly.
[0071] Table 3. Used for calculation Aromatic Score empirical constants of each substituent
[0072]
[0073] (III) Trichloromethane generation potential (TCM) FP The establishment and optimization of prediction models.
[0074] Based on the aforementioned key molecular characteristic parameters, this invention constructs a TCM using the Extreme Gradient Boosting (XGBoost) machine learning algorithm. FP The regression prediction model was constructed, and SHAP was used to evaluate the TCM pairs of descriptors in the final model. FP The contribution of the predicted values is characterized, and the specific process is executed using Python 3.12; in other words, the contribution of each feature parameter to the TCM is characterized through SHAP interpretive analysis. FP The contribution weights of the predicted values ensure that the model has both "high accuracy" and "interpretability".
[0075] Specifically, all calculations were implemented using Python 3.12, utilizing open-source libraries such as xgboost 2.1.4, shap0.46.0, rdkit 2024.03.5, and sklearn 1.6.1 to complete model building, feature importance analysis, and performance evaluation; the coefficient of determination was used to evaluate model accuracy. R 2 (Measures how well the model fits the data) R 2 (The closer to 1 the better) and root mean square error RMSE (Measures the deviation between model predictions and experimental values) RMSE (The smaller the better) is used as a statistical indicator to characterize the model's fit performance; the square of the predictive correlation coefficient is used. Q 2 (Measure the model's predictive ability on unknown external data,) Q 2 A score >0.8 indicates an excellent model, representing the model's predictive performance.
[0076] It should be noted that model training and prediction must be performed in a Python 3.12 or equivalent open-source library environment. If the software version is changed, the model performance must be re-verified to avoid version compatibility issues that could lead to increased prediction errors.
[0077] In this embodiment of the invention, the Extreme Gradient Boosting (XGBoost) machine learning algorithm is employed, and a Trichloromethane Generation Potential (TCM) is constructed based on the trichloromethane generation potential values of all organic compounds in the overall dataset and key molecular feature parameters. FP The regression model between the trichloromethane formation potential and the key molecular characteristic parameters is the trichloromethane formation potential prediction model (also known as the XGBoost prediction model). Then, the accuracy of the trichloromethane formation potential prediction model is verified.
[0078] The specific implementation process is as follows:
[0079] The dataset, consisting of the chloroform generation potential values and key molecular characteristic parameters of all obtained organic compounds (e.g., 222 organic compounds), was divided into a training set and a test set. The ratio of the training set to the test set was set to 8:2, and the sum of the ratios of the training set and the test set was always 1. Each training set and the test set independently contained the chloroform generation potential values and molecular characteristic parameter data of multiple organic compounds, with each organic compound corresponding to one chloroform generation potential value and five key molecular characteristic parameter data.
[0080] A model for predicting the formation potential of chloroform was established by training the model using the training set.
[0081] The accuracy of the trichloromethane formation potential prediction model was validated using a test set. The coefficient of determination was used to evaluate the model accuracy. R 2 (See Equation (2)) and root mean square error RMSE (See Equation (3)) as a statistical indicator to characterize the model's fit performance, and the square of the predicted correlation coefficient is used. Q 2 (See Equation (4)) characterizes the predictive performance of the model.
[0082] (2)
[0083] (3)
[0084] (4)
[0085] In the formula, and These represent the predicted and experimental values of the chloroform formation potential, respectively. This represents the average value of experimental values for the trichloromethane formation potential. This represents the number of samples in the dataset. It can be understood that each sample corresponds to a key molecular feature in an organic compound as input and an experimental value for the trichloromethane formation potential as output.
[0086] Generally speaking, the model R 2 >0.60, Q 2 When the value is greater than 0.50, the model has an acceptable goodness of fit. Q 2 A score >0.8 indicates a model with superior performance.
[0087] (iv) Optimize to obtain the optimal prediction model for the trichloromethane generation potential.
[0088] To avoid overfitting and improve generalization ability, a Bayesian optimization algorithm was used to optimize the key hyperparameters of the trichloromethane generation potential prediction model, such as tree depth (max_depth), learning rate (eat), subsample weight (min_child_weight), sample ratio (subsample), feature ratio (colsample_bytree), L1 regularization coefficient (alpha), L2 regularization coefficient (lambda), and minimum loss decrease value required for node splitting (gamma). The optimal prediction model for trichloromethane generation potential and the optimal combination of hyperparameters were obtained, as shown in Table 4.
[0089] Table 4. Hyperparameters of the optimal prediction model for chloroform formation potential
[0090]
[0091] In this embodiment of the invention, the key molecular characteristic parameters of the target compound (the compound to be tested) are input into the optimal prediction model for the trichloromethane formation potential to obtain the trichloromethane formation potential value corresponding to the target compound.
[0092] (v) Performance evaluation and verification of the optimal prediction model for chloroform generation potential.
[0093] The fitting ability, stability, and predictive ability of the optimal prediction model for chloroform generation potential were comprehensively evaluated through three methods: fitting performance analysis, simulated external validation, and one-out-of-one cross-validation (LOO-CV).
[0094] Specifically, the fitting performance analysis was conducted using a dataset comprised of the trichloromethane formation potential values and key molecular characteristic parameters of 222 organic compounds. The key molecular characteristic parameters of these 222 organic compounds were used as input, and the model-predicted trichloromethane formation potential values were used as output. These values were then compared with the corresponding experimental values of the trichloromethane formation potential. The final model fitting result was... R 2 =0.94, RMSE =0.06, indicating that the model has good fitting ability, there is no systematic bias between the prediction error and the experimental value, and the statistical performance is significantly better than the existing technology.
[0095] External validation simulation: The overall dataset consisting of the trichloromethane generation potential values and key molecular characteristic parameters of 222 organic compounds was randomly divided into a training set (containing 178 organic compounds) and a test set (containing 44 organic compounds) in an 8:2 ratio. It can be understood that the training set and test set used when modeling the trichloromethane generation potential prediction model can be used to evaluate and validate the above-mentioned optimal trichloromethane generation potential prediction model.
[0096] See Figure 1 As shown, the fitting results validated using the training set are: R 2 =0.95 and RMSE =0.06; The fitting result validated using the test set is Q 2 =0.94 and RMSE =0.06. It is evident that the statistical performance of both the training and test sets is very close to that of the model based on the entire set, indicating that the optimal prediction model for chloroform generation potential is based on TCM. FP The essential correlation between the key molecular characteristic parameters and the established correlation is not accidental and exhibits excellent statistical stability.
[0097] One-out cross-validation: The core is iterative training + prediction. The training set (containing 178 organic compounds) is used for evaluation and validation. Each time, one sample is retained as the test set, and the remaining samples are used as the training subset. After the first sample is validated, it is deleted, and this process is repeated. Finally, the prediction results from all rounds are used to calculate... Q 2 and RMSE This represents the model's LOO-CV performance on the training set. Among them, LOO... -Q 2 CV =0.81, RMSE CV =0.09, further verifying the model's stability and predictive consistency under different data distributions, ensuring the model's ability to predict unknown compounds TCM. FP The reliability of the prediction.
[0098] It is worth mentioning that, in order to avoid applying the model to compounds with large structural differences and causing prediction bias, the optimal prediction model for chloroform generation potential is also characterized by application domain characterization in this embodiment of the invention. For example, the application domain of the target compound can be determined before inputting the target compound into the optimal prediction model for chloroform generation potential.
[0099] (vi) Model application domain representation.
[0100] In this embodiment of the invention, the "molecular fingerprint similarity threshold method" is used to characterize the applicability of the model, avoiding prediction bias caused by applying the model to compounds with excessively different structures. The specific operation method is as follows:
[0101] (1) Fingerprint calculation: The extended connectivity fingerprint (ECFP4) between the target compound and each compound in the training set is calculated using the rdFingerprintGenerator tool in the rdkit library. Based on this fingerprint, the Tanimoto similarity coefficient (ST) between the target compound and the compounds in the training set is calculated. The larger the ST value, the higher the structural similarity of the compounds.
[0102] (2) Threshold setting: Set similarity threshold S cutoff =0.25, and a threshold for the number of compounds in the training set is set. N min =1 (to ensure that the target compound has at least one structurally similar training set sample as a reference).
[0103] (3) Application domain determination: Statistically, the ST value of the target compound in the training set is greater than or equal to the target compound's ST value. S cutoff The number of compounds with ST ≥ 0.25, if this number ≥ Nmin If the value is ≥1, then the target compound is determined to be within the model application domain.
[0104] (4) Application domain scope: Combining the above determination methods, the TCM constructed in this invention FP The application domain of the prediction model is defined as: organic compounds with a Tanimoto coefficient ST ≥ 0.25 based on the ECFP4 fingerprint of the training set compounds. (See [link to relevant documentation]). Figure 2 As shown, the specific categories include carboxylic acids (such as aromatic carboxylic acids / aliphatic carboxylic acids), esters (such as phenolic esters / sugar esters), ketones (such as aromatic ketones / aliphatic ketones), aldehydes (such as aromatic aldehydes), alcohols / phenols (such as polyphenols / sugar alcohols), amines (such as amino acids / aromatic amines), ethers (such as aromatic ethers), heterocyclic compounds (such as rings containing N / O / S) (pyrimidines / purines / imidazolium), sulfonic acids / sulfonamides (such as sulfonyl ketones), halogenated hydrocarbons (such as chlorophenols), nitro compounds (such as nitroaromatic hydrocarbons), sugars (such as polysaccharides / glycosides), and other structurally similar compounds. This prediction model can predict the TCM of compounds within this range. FP The value can be accurately predicted.
[0105] Unless otherwise defined, the technical terms used in the following embodiments have the same meaning as commonly understood by those skilled in the art. Unless otherwise specified, the experimental reagents used in the following embodiments are all conventional biochemical reagents; the raw materials, instruments, and equipment used in the following embodiments can all be obtained commercially or through existing methods; unless otherwise specified, the amounts of experimental reagents used are the amounts used in conventional experimental operations; unless otherwise specified, the experimental methods are conventional methods. It should be further noted that the following description is merely exemplary and not a specific limitation of the present invention.
[0106] Example 1: TCM of acetic acid FP predict
[0107] The 2D molecular structure of acetic acid is shown in equation (I) below:
[0108]
[0109] Formula (1)
[0110] Based on the experimental values of chloroform generation potential of 222 organic compounds obtained in this invention and the key molecular characteristic parameters, an optimal prediction model for chloroform generation potential was constructed, and TCM of acetic acid was performed based on this optimal prediction model. FP The prediction process includes:
[0111] Calculation of key molecular characteristic parameters: Structure optimization and frequency analysis were performed using the DFT M062X / 6-311+G(d, p) algorithm in Gaussian 16 software. The results were obtained from the output file.E HOMO and E LUMO The chemical hardness of the compound was calculated according to formula (1). η =0.206 eV; further calculations using GsGrid 1.7 yielded the most negative electrostatic potential at the molecular surface. V s,min =-0.075 eV; The carbonyl index was calculated according to the parameter calculation rules in Tables 1 to 3. CI =0, olefination index Enolizable Score =0.1 and aromaticity index Aromatic Score =0.
[0112] TCM FP Prediction: By inputting the above five key molecular characteristic parameters into the optimal prediction model for chloroform formation potential, the model outputs the TCM of acetic acid. FP The predicted value is 0.003 mol / mol.
[0113] Furthermore, Example 1 also determined the application domain of the model: using the rdFingerprintGenerator tool from the RDKit library, the ECFP4 fingerprint Tanimoto similarity coefficient (ST) between acetic acid and compounds in the training set was calculated based on the SMILES of the compounds. The results showed that its maximum ST with the compounds in the training set was 0.61, satisfying ST≥0.25. N min The ≥1 criterion further proves that acetic acid is within the model's application domain and can be accurately predicted.
[0114] Example 2: TCM of 1,3-acetone dicarboxylic acid FP predict
[0115] The 2D molecular structure of 1,3-propanone dicarboxylic acid is shown in formula (II) below:
[0116]
[0117] Formula (II)
[0118] Based on the experimental values of chloroform generation potential of 222 organic compounds obtained in this invention and the key molecular characteristic parameters, an optimal prediction model for chloroform generation potential was constructed. Based on this optimal prediction model, TCM of 1,3-acetone dicarboxylic acid was performed. FP The prediction process includes:
[0119] Molecular characteristic parameter calculation: The DFT M062X / 6-311+G(d, p) algorithm of Gaussian 16 software was used for structure optimization and frequency analysis. The results were obtained from the output file.E HOMO and E LUMO The chemical hardness of the compound was calculated according to formula (1). η =0.172 eV; further calculations using GsGrid 1.7 yielded the most negative electrostatic potential at the molecular surface. V s,min =-0.073eV; According to the parameter calculation rules in Tables 1 to 3, the carbonyl index is calculated respectively. CI =3.000, olefination index Enolizable Score =4.680 and aromaticity index Aromatic Score =0.000.
[0120] TCM FP Prediction: By inputting the above five key molecular characteristic parameters into the optimal prediction model for chloroform formation potential, the model outputs the TCM of 1,3-acetone dicarboxylic acid. FP The predicted value is 0.600 mol / mol.
[0121] Furthermore, in Example 2, the application domain of the model was determined: using the rdFingerprintGenerator tool from the RDKit library, the ECFP4 fingerprint Tanimoto similarity coefficient (ST) between 1,3-acetone dicarboxylic acid and the compounds in the training set was calculated based on the SMILES of the compounds. The results showed that its maximum ST with the compounds in the training set was 0.58, satisfying ST≥0.25. N min The ≥1 criterion further proves that 1,3-acetone dicarboxylic acid is within the model's application domain and can be accurately predicted.
[0122] Example 3: TCM of 3-hydroxyacetophenone FP predict
[0123] The 2D molecular structure of 3-hydroxyacetophenone is shown in formula (III) below:
[0124]
[0125] Formula (3)
[0126] Based on the experimental values of chloroform generation potential of 222 organic compounds obtained in this invention and the key molecular characteristic parameters, an optimal prediction model for chloroform generation potential was constructed, and TCM of 3-hydroxyacetophenone was performed based on this optimal prediction model. FP The prediction process includes:
[0127] Molecular characteristic parameter calculation: The DFT M062X / 6-311+G(d, p) algorithm of Gaussian 16 software was used for structure optimization and frequency analysis. The results were obtained from the output file. E HOMO and E LUMO The chemical hardness of the compound was calculated according to formula (1). η =0.134 eV; further calculations using GsGrid 1.7 yielded the most negative electrostatic potential at the molecular surface. V s,min =-0.080eV; According to the parameter calculation rules in Tables 1 to 3, the carbonyl index is calculated respectively. CI =0.000, olefination index Enolizable Score =1.000 and aromaticity index Aromatic Score =0.700.
[0128] TCM FP Prediction: By inputting the above five key molecular characteristic parameters into the optimal prediction model for chloroform generation potential, the model outputs the TCM of 3-hydroxyacetophenone. FP The predicted value is 0.448 mol / mol.
[0129] Furthermore, Example 3 also determined the application domain of the model: using the rdFingerprintGenerator tool from the RDKit library, the ECFP4 fingerprint Tanimoto similarity coefficient (ST) between 3-hydroxyacetophenone and the compounds in the training set was calculated based on the SMILES of the compounds. The results showed that its maximum ST with the compounds in the training set was 0.74, satisfying ST≥0.25. N min The ≥1 criterion further proves that 3-hydroxyacetophenone is within the model's application domain and can be accurately predicted.
[0130] Example 4: TCM of glucose FP predict
[0131] The 2D molecular structure of glucose is shown in equation (iv):
[0132]
[0133] Formula (IV)
[0134] Based on the experimental values of chloroform generation potential of 222 organic compounds obtained in this invention and the key molecular characteristic parameters, an optimal prediction model for chloroform generation potential was constructed, and TCM of glucose was performed based on this optimal prediction model. FP The prediction process includes:
[0135] Molecular characteristic parameter calculation: The DFT M062X / 6-311+G(d, p) algorithm of Gaussian 16 software was used for structure optimization and frequency analysis. The results were obtained from the output file. E HOMO and E LUMO The chemical hardness of the compound was calculated according to formula (1). η =0.205 eV; further calculations using GsGrid 1.7 yielded the most negative electrostatic potential at the molecular surface. V s,min =-0.082eV; According to the parameter calculation rules in Tables 1 to 3, the carbonyl index is calculated respectively. CI =0.000, alkeneability index Enolizable Score =0.140 and aromaticity index Aromatic Score =0.000.
[0136] TCM FP Prediction: Inputting the above 5 molecular characteristic parameters into the optimal prediction model for chloroform formation potential, the model outputs the TCM of glucose. FP The predicted value is 0.015 mol / mol.
[0137] Furthermore, Example 4 also determined the application domain of the model: using the rdFingerprintGenerator tool from the RDKit library, the ECFP4 fingerprint Tanimoto similarity coefficient (ST) between glucose and the compounds in the training set was calculated based on the SMILES of the compounds. The results showed that the maximum ST between glucose and the compounds in the training set was 0.45, satisfying ST≥0.25. N min The ≥1 criterion further proves that glucose is within the model's application domain and can be accurately predicted.
[0138] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A rapid prediction method for the trichloromethane generation potential during water disinfection, characterized in that, Includes the following steps: (1) Obtain multiple organic compounds and their respective trichloromethane generation potential values; (2) Obtain the key molecular characteristic parameters for accurately characterizing the molecular reactivity of each organic compound; the key molecular characteristic parameters include the most negative electrostatic potential of the molecular surface of the compound. V s,min Chemical hardness η Carbonyl index CI olefin index Enolizable Score and aromaticity index Aromatic Score ; (3) An extreme gradient boosting machine learning algorithm is adopted, and a chloroform generation potential prediction model is constructed based on the chloroform generation potential value and key molecular characteristic parameters of the organic compound. The accuracy of the chloroform generation potential prediction model is verified. (4) The Bayesian optimization algorithm is used to optimize the trichloromethane generation potential prediction model to obtain the optimal trichloromethane generation potential prediction model; (5) Input the key molecular characteristic parameters of the target compound into the optimal prediction model for chloroform generation potential to obtain the chloroform generation potential value of the target compound.
2. The rapid prediction method as described in claim 1, characterized in that, The organic compounds include at least one of phenols, carboxylic acids, esters, halogenated hydrocarbons, amines, nitro compounds, amides, aldehydes, ketones, ethers, alcohols, heterocyclic compounds, sulfides, cyanides, sugars, and polyhydroxy compounds; and / or, The number of organic compounds is 222.
3. The rapid prediction method as described in claim 1, characterized in that, The determination conditions for the trichloromethane generation potential of each organic compound are the same, including a chlorine dosage of ≥5 mg / L in the reaction system of the organic compound and chlorine disinfectant, a pH value of 7.0~8.0, and complete reaction of chlorine with the organic compound.
4. The rapid prediction method as described in claim 1, characterized in that, The data distribution range of the trichloromethane generation potential value is 0 mol / mol to 0.99 mol / mol.
5. The rapid prediction method as described in claim 1, characterized in that, In step (2), the key molecular feature parameters are selected based on the mechanism analysis of the reaction between each organic compound and the chlorine disinfectant and the results of recursive feature elimination based on SHAP contribution.
6. The rapid prediction method as described in claim 1, characterized in that, In step (3), the construction and accuracy verification of the trichloromethane generation potential prediction model include: The trichloromethane generation potential values and key molecular characteristic parameter data of all the obtained organic compounds were divided into training set and test set; The training set is used to train the model and establish the trichloromethane generation potential prediction model. The accuracy of the trichloromethane generation potential prediction model was verified using the test set.
7. The rapid prediction method as described in claim 6, characterized in that, The coefficient of determination is used to evaluate the accuracy of the model. R 2 and root mean square error RMSE As a statistical indicator to characterize the model's fit performance, and using the square of the predicted correlation coefficient... Q 2 Characterize the predictive performance of the model; where, Q 2 >0.8; And / or, the ratio of the data volume of the training set to the data volume of the test set is set to 8:2, and the sum of the ratios of the data volume of the training set to the data volume of the test set is always 1.
8. The rapid prediction method as described in claim 6, characterized in that, In step (4), the key hyperparameters of the optimal prediction model for chloroform generation potential include: tree depth = 3, learning rate = 0.097, subsample weight = 3, sample ratio = 0.65, feature ratio = 0.80, L1 regularization coefficient = 0.014, L2 regularization coefficient = 2.12, and minimum loss reduction required for node splitting = 0.
029.
9. The rapid prediction method as described in claim 1, characterized in that, The rapid prediction method also includes: The fitting ability, stability, and predictive ability of the optimal prediction model for chloroform generation potential were comprehensively evaluated using three methods: fitting performance analysis, simulated external validation, and cross-validation with one-to-one elimination.
Citation Information
Patent Citations
Method for predicting concentration of trihalomethane in water supply system based on radial neural network and grey relational analysis modeling
CN110287652A
Drinking water disinfection by-product generation prediction method based on reaction kinetic model
CN120833862A