A method for predicting the cytotoxicity of disinfection by-products
By establishing a cytotoxicity database of DBPs and machine learning algorithms, the CHO cytotoxicity of disinfection byproducts is predicted, and the problem of lack of prediction methods in the existing technology is solved, and rapid and accurate toxicity assessment is achieved, reducing experimental costs.
Patent Information
- Application Number
- CN202210680219.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-16
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2042-06-16
AI Technical Summary
The prior art lacks effective methods to predict the toxicity of disinfection by-products (DBPs) to CHO cells, resulting in difficulties in assessing and controlling environmental risks.
Establish a cytotoxic database of DBPs, obtain the SMILES expression of DBPs, calculate the molecular fingerprint and preprocess it, and use machine learning algorithms to build a toxicity prediction model, and directly output the cytotoxicity value of DBPs by inputting the SMILES expression.
It achieves rapid and accurate prediction of CHO cytotoxicity of DBPs, reduces the demand for biological experiments, saves manpower and material resources, and provides guidance for scientific research and drinking water risk assessment.
Smart Images

Figure CN114974460B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of environmental risk assessment, and particularly relates to a method for predicting the cytotoxicity of disinfection by-products. Background Art
[0002] Drinking water disinfection is an important public health measure, which helps to inactivate pathogenic microorganisms and thus prevent water-borne diseases. However, disinfectants (such as chlorine, chloramine, chlorine dioxide, etc.) may inadvertently react with natural organic matter and halogens in source water to generate disinfection by-products (DBPs). Many DBPs have cytotoxicity, genotoxicity, mutagenicity, teratogenicity or carcinogenicity. These characteristics that have adverse effects on organisms are of great guiding significance for environmental risk assessment and control. At present, the number of DBPs in the environment is huge and growing rapidly. Conducting experiments on all DBPs is laborious and costly. Therefore, it is particularly important to understand the cytotoxicity of DBPs that have not been experimented and to pre-screen the toxicity of DBPs before conducting experiments.
[0003] Cytotoxicity is the determination of the toxic effects of exogenous compounds or other factors in the environment on cell structure and function. Generally, cytotoxicity experiments are carried out with in vitro cell culture. In vitro cell culture refers to the culture technique in which cells grow and proliferate under suitable conditions in vitro. Chinese hamster ovary cells (CHO) are widely used in toxicology research. The concentration for 50% of maximal effect (EC 50 ) refers to the concentration that can cause 50% of the maximum effect. Using the EC 50 of CHO cells as an index to measure cytotoxicity is very common in research and has important reference significance for environmental risk assessment and control. By the method of this patent, the accurate EC 50 of DBPs can be predicted, which can reduce the large amount of manpower, material resources, financial resources and time required for biological experiments.
[0004] The Chinese patent document with the publication number CN114171137 discloses a method for predicting the environmental hazard of compounds based on machine learning. Based on the molecular structure of compounds, a prediction model is established according to the relationship between the compound structure and its PMT properties (persistence, mobility and toxicity) or vPvM properties (high persistence and high mobility) to predict the PMT properties or vPvM properties of compounds, including the following steps: (1) establishing screening criteria for the environmental hazard of compounds; (2) extracting some compounds from the compound database as samples, and taking the SMILES expressions of these exported samples as sample data; (3) constructing a prediction model based on machine learning algorithms and optimizing the parameters of the prediction model; (4) finally using the optimized prediction model to predict whether a new molecule has environmental hazard. The Chinese patent document with the publication number CN110890137A discloses a method for modeling a compound toxicity prediction model, including: (1) establishing classification labels for the toxicity of compounds; (2) providing molecular descriptors of each candidate modeling compound; (3) providing target protein descriptors of each candidate modeling compound; (4) providing quantitative high-throughput screening analysis descriptors of each candidate modeling compound; (5) constructing and training a compound toxicity prediction model and being able to make predictions.
[0005] However, so far, there is a lack of corresponding technologies in the field of predicting the cytotoxicity of DBPs in CHO cells. Summary of the Invention
[0006] 1. Problems to be Solved
[0007] The purpose of the present invention is to provide a method for predicting the cytotoxicity of disinfection by-products in the field of predicting the cytotoxicity of DBPs in CHO cells.
[0008] 2. Technical Solutions
[0009] In order to solve the above problems, the technical solutions adopted by the present invention are as follows:
[0010] A method for predicting the cytotoxicity of disinfection by-products includes at least the following steps:
[0011] (1) Establishing a cytotoxicity database of DBPs;
[0012] (2) Obtaining the SMILES of DBPs samples;
[0013] (3) Calculating the molecular fingerprints of DBPs samples and preprocessing the sample data;
[0014] (4) Constructing a toxicity prediction model based on machine learning algorithms: retaining the samples that simultaneously have all descriptors and cytotoxicity values to construct a data set, calculating relevant parameters for model evaluation, and screening the model;
[0015] (5) After inputting the SMILES expression of the DBPs to be measured, the molecular fingerprint of the DBPs to be measured is automatically calculated and then input into the prediction model to predict the cytotoxicity value of the DBPs to be measured;
[0016] Among them, the cytotoxicity refers to the EC of CHO cells 50 value.
[0017] Furthermore, the sources of the disinfection by-product data are as follows:
[0018] Published literature. Illustratively, the literature can be relevant literature in JOURNAL OF ENVIRONMENT SCIENCES (Acta Scientiae Circumstantiae) and WATER RESEARCH;
[0019] Widely recognized public databases. Illustratively, the public databases can be ToxCast, PubChem, etc.;
[0020] Standardized and scientific biological experiments. Illustratively, the biological experiments can be experimental data from the State Key Laboratory of Pollution Control and Resource Reuse of Nanjing University.
[0021] Furthermore, in step (2), during the process of obtaining the SMILES expression of DBPs, substances that cannot be converted into SMILES are excluded.
[0022] Furthermore, in step (3), the molecular fingerprint needs to be the 166-bit molecular fingerprint of MACCS; and / or, the 1024-bit extended connectivity fingerprint of ECFP_4; and / or, the 1024-bit functional group type fingerprint of FCFP_4.
[0023] Furthermore, in step (3), the preprocessing methods include standardization and normalization;
[0024] The standardization processes the data according to the columns of the feature matrix, converting the feature values of the samples to the same dimension;
[0025] The normalization processes the data according to the rows of the feature matrix, mapping the data to a specified range.
[0026] Furthermore, in step (4), the machine learning algorithm is selected from the random forest algorithm, support vector machine algorithm, naive Bayes algorithm, and artificial neural network algorithm.
[0027] Furthermore, in step (4), the data set constructed from the samples is divided into a training set and a test set;
[0028] The prediction model is trained using the training set;
[0029] Evaluate the goodness of the prediction model using the test set and optimize the parameters of the prediction model.
[0030] Further, the training set and the test set are divided according to the ratio of (8 - 7):(2 - 3). Schematically, they are divided according to the ratio of 8:2 or 7:3.
[0031] Further, in step (4), the model is screened by calculating the regression coefficient and the mean square error.
[0032] Further, select the R 2 The model closest to 1 and with the minimum MSE is the optimal model.
[0033] Further, the calculation formula of MSE is:
[0034] where n is the number of samples, Y i is the true value of the sample, is the predicted value of the sample.
[0035] 3. Beneficial effects
[0036] Compared with the prior art, the method for predicting the cytotoxicity of disinfection by-products provided by the present invention:
[0037] 1) Fills the gap in the field of predicting the cytotoxicity of CHO cells of DBPs in the current technology.
[0038] 2) Regression based on the machine learning method can perform quantitative prediction, that is, predict the specific toxicity value, and the accuracy is higher than that of the traditional regression method.
[0039] Different from the traditional known classification based on the machine learning method, it can only perform qualitative prediction, such as whether a compound is toxic or non-toxic.
[0040] 3) A large number of complicated biological experiments can be omitted. Just input the SMILES expression of the substance to be tested, and the predicted value of the cytotoxicity can be directly output, which has the advantages of batch, rapidity, and accuracy, saving manpower, time, and economic costs. It can be used to narrow the toxicity screening range of disinfection by-products and provide guidance for scientific research work and the risk assessment and control of drinking water. Brief description of the drawings
[0041] Figure 1 is the flow chart of the method for predicting the cytotoxicity of disinfection by-products based on machine learning of the present invention;
[0042] Figure 2 is the comparison chart of the predicted value and the true value of the random forest;
[0043] Figure 3It is a comparison chart of the predicted values and the true values of the artificial neural network. Detailed implementation manners
[0044] The method for predicting the cytotoxicity of disinfection by-products based on machine learning provided by the present invention is based on the molecular structure of DBPs, establishes a prediction model according to the relationship between the molecular structure of DBPs and their cytotoxicity, uses the SMILES expression of the molecule to be measured as the input, calculates the molecular fingerprint of the molecule to be measured, and finally outputs the predicted toxicity value of the substance to be measured.
[0045] On the basis of the foregoing [2. Technical solution], the more specific steps include:
[0046] (1) Collect the cytotoxicity values of DBPs from published academic journals, and / or widely recognized public databases, and / or standardized and scientific biological experiments, and establish a database;
[0047] (2) Provide the SMILES (Simplified Molecular Input Line Entry Specification) of all DBPs samples in step (1);
[0048] (3) Calculate the molecular fingerprints of all DBPs samples in step (1), and preprocess the sample data;
[0049] (4) Construct a toxicity prediction model based on machine learning algorithms: retain the samples with all descriptors and cytotoxicity values at the same time to construct a data set, calculate the relevant parameters for model evaluation, and screen the model;
[0050] (5) After inputting the SMILES expression of the DBP to be measured, automatically calculate the molecular fingerprint of the DBP to be measured, and then input it into the prediction model with optimized parameters to predict the cytotoxicity value of the DBP to be measured;
[0051] As described herein, the "cytotoxicity" in the method refers to the half-maximal effective concentration of Chinese hamster ovary cells, that is, the cytotoxicity refers to the EC 50 value of CHO cells. Among them, CHO cells are widely used in toxicology research, and EC 50 refers to the concentration that can cause 50% of the maximum effect. Its values are all from published journal articles and are measured through standardized biological experiments and scientific and rigorous statistical calculations.
[0052] As described herein, the "molecular fingerprint" in the method is a MACCS fingerprint (molecular access system, MACCS) or a Morgan molecular fingerprint. MACCS is a typical molecular fingerprint, and each of its 166 bits encodes specific structural features, such as whether the number of methyl groups in the molecule is greater than 1; whether the molecule is aromatic, etc. This fingerprint is derived from a chemical structure database developed by MDL (Molecular Design LTD). MDL is well-known for cheminformatics, and the molecular fingerprints it developed are widely used and highly recognized in related fields. ECFP (Extended Connectivity Fingerprints, ECFP) / FCFP (Functional-Class Fingerprints, FCFPs) both belong to Morgan Fingerprints. Morgan Fingerprints is a circular fingerprint and also belongs to topological fingerprints, which is obtained by modifying the standard Morgan algorithm. Since its definition requires setting a radius n (i.e., the number of iterations), and then calculating each atomic environment identifier. When n = 2, it is ECFP_4, and in this method, the length of 1024 bits is taken. All three fingerprints can be extracted using the RDkit toolkit.
[0053] As described herein, the "preprocessing method" in the method includes standardization and normalization. Standardization processes data according to the columns of the feature matrix, converting the feature values of the samples to the same dimension; normalization processes data according to the rows of the feature matrix, mapping the data to a specified range. Usually, this interval is [0, 1].
[0054] As described herein, the "machine learning algorithms" in the method include Random Forest algorithm, Support Vector Machines algorithm, Naive Bayesian algorithm, and Artificial neural networks algorithm. Specifically, the parameters of the machine learning algorithms are set according to the size of the database. For example, if the database has 90 data, then 100 trees can be set for the Random Forest method, and 3 hidden layers can be set for the Artificial Neural Network, with 80 neurons corresponding to each layer, etc.
[0055] The model goodness-of-fit evaluation indicators include: the regression coefficient R of each model regression 2 and the mean squared error MSE. The calculation formula of MSE is:
[0056] where n is the number of samples, and Y i is the true value of the sample. is the predicted value of the sample.
[0057] In summary, the present invention provides a method for predicting the cytotoxicity of DBPs by means of the molecular structure and physicochemical properties of compounds and using machine learning algorithms. The method process includes: collecting the cytotoxicity values of DBPs and establishing a database; converting all DBPs into SMILES; calculating the molecular fingerprints of all DBP samples, and normalizing and standardizing the sample data; constructing a toxicity prediction model based on multiple machine learning algorithms and selecting the optimal model; after inputting the SMILES expression of the DBP to be measured, directly outputting the predicted cytotoxicity value of the DBP to be measured.
[0058] The essential features and remarkable effects of the present invention can be reflected from the following embodiments. The described embodiments are part of the embodiments of the present invention, rather than all embodiments. Therefore, they do not impose any limitations on the present invention. Those skilled in the art make some non-essential improvements and adjustments based on the content of the present invention, which all fall within the protection scope of the present invention.
[0059] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present invention belongs; the terms and / or include any and all combinations of one or more related listed items.
[0060] The following further illustrates the present invention with specific embodiments, but the embodiments do not impose any form of limitation on the present invention. Unless otherwise specified, the reagents, methods, and equipment used in the present invention are conventional reagents, methods, and equipment in the technical field.
[0061] Example 1
[0062] The flow chart of the method for predicting the cytotoxicity of disinfection by-products based on machine learning in the embodiments of the present invention is as Figure 1 shown, including 5 steps as Figure 1 shown:
[0063] (1) Collect the cytotoxicity values of DBPs from published academic journals and establish a database;
[0064] Collect the cytotoxicity values of DBPs from published academic journals and establish a database. In the present invention, cytotoxicity refers to the EC 50 value. Each DBP corresponds to an EC 50 , and there are 90 DBPs in total.
[0065] (2) Provide the SMILES of all DBP samples in step (1).
[0066] Use chemical professional software such as ChemDraw to convert the names of all DBP samples found in step (1) into SMILES expressions.
[0067] (3) Calculate the molecular fingerprints of all DBP samples in step (1) and preprocess the sample data;
[0068] Use the RDKit toolkit to extract the specific structural features of the samples. Take the SMILES of the samples as input and the ECFP_4 molecular fingerprint as output. Each column of data corresponds to a molecular fingerprint, and finally obtain 1024 columns of molecular fingerprints. Adding the predicted value EC 50 As a column, it becomes a feature matrix of 90 rows and 1025 columns. Use the StandardScaler function in the sklearn.preprocessing toolkit to standardize and normalize the sample dataset.
[0069] (4) Build a toxicity prediction model based on machine learning algorithms: Retain the samples that have both all descriptors and cytotoxicity values to construct a dataset, calculate the relevant parameters for model evaluation, and screen the models.
[0070] Use the random forest method to perform regression modeling on the data obtained in step (3), set the parameters (n_estimators = 85, random_state = 0), and then use the artificial neural network for regression modeling, set the parameters (solver = 'lbfgs', alpha = 0.5e -5 , hidden_layer_sizes=(80, 80, 80, 80, 50), random_state = 1). Obtain two prediction models, and then calculate the regression coefficients R 2 And the mean squared error MSE of the two models respectively, and compare them. Select the model with R 2 Closest to 1 and the smallest MSE as the optimal model.
[0071] (5) After inputting the SMILES expression of the DBP to be measured, automatically calculate the molecular fingerprint of the DBP to be measured, and then input it into the prediction model with optimized parameters to predict the cytotoxicity value of the DBP to be measured.
[0072] After inputting the SMILES expression of the DBP to be measured, automatically call the ECFP_4 molecular fingerprint calculation function in the RDkit toolkit to calculate the molecular fingerprint of the DBP to be measured, and then input it into the prediction model with optimized parameters to predict the cytotoxicity value of the DBP. In the specific implementation, this step can be completed by a computer program. Therefore, only by inputting the SMILES expression, the predicted cytotoxicity value of the DBP to be measured will be directly output and displayed. The comparison chart of the predicted results and the true results of 90 kinds of DBPs is asFigure 2 , 3 as shown.
Claims
1. A method for predicting the cytotoxicity of disinfection by-products, characterized in that it at least includes the following steps: (1) Establish a cytotoxicity database of DBPs; (2) Obtain the SMILES of the DBP samples. During the process of obtaining the SMILES expressions of the DBPs, substances that cannot be converted into SMILES are excluded; (3) Calculate the molecular fingerprints of the DBP samples and preprocess the sample data; The molecular fingerprints are: the 166-bit molecular fingerprint of MACCS, and / or the 1024-bit extended connectivity fingerprint of ECFP_4, and / or the 1024-bit functional group type fingerprint of FCFP_4; The preprocessing method includes standardization and normalization; The standardization is to process the data according to the columns of the feature matrix, and convert the feature values of the samples to the same dimension; The normalization is to process the data according to the rows of the feature matrix and map the data to a specified range; (4) Construct a toxicity prediction model based on a machine learning algorithm: retain the samples with all descriptors and cytotoxicity values at the same time to construct a data set, calculate the relevant parameters for model evaluation, and screen the model; (5) Automatically calculate the molecular fingerprints of the DBP to be measured after inputting the SMILES expression of the DBP to be measured, and then input them into the prediction model to predict the cytotoxicity value of the DBP to be measured; wherein, The cytotoxicity refers to the EC of CHO cells 50 value.
2. The method for predicting the cytotoxicity of disinfection by-products according to claim 1, characterized in that in step (4), the machine learning algorithm is selected from any one of the random forest algorithm, support vector machine algorithm, naive Bayes algorithm and artificial neural network algorithm.
3. The method for predicting the cytotoxicity of disinfection by-products according to any one of claims 1 to 2, characterized in that in step (4), the data set constructed from the samples is divided into a training set and a test set; Use the training set to train the prediction model; Use the test set to evaluate the goodness of the prediction model and optimize the parameters of the prediction model.
4. The method for predicting the cytotoxicity of disinfection by-products according to claim 3, characterized in that the training set and the test set are divided in a ratio of (8-7):(2-3).
5. The method for predicting the cytotoxicity of disinfection by-products according to claim 3, characterized in that in step (4), the model is screened by calculating the regression coefficient and the mean square error.
6. The method for predicting the cytotoxicity of disinfection by-products according to claim 5, characterized in that Select R 2 The model with the value closest to 1 and the minimum MSE is the optimal model.
7. The method for predicting the cytotoxicity of disinfection by-products according to claim 6, characterized in that The calculation formula of MSE is: where n is the number of samples, is the true value of the sample, is the predicted value of the sample.
Citation Information
Patent Citations
Modeling method and device of compound toxicity prediction model and application of compound toxicity prediction model
CN110890137A
Method and device for identifying high-risk disinfection by-products in water and application thereof
CN112147244A
Method for predicting environmental harmfulness of compounds based on machine learning
CN114171137A