Construction method of micro-plastic toxicity prediction model based on transcriptomics and QSAR (Quantitative Synthetic Aperture Radar) model

By combining transcriptomics and QSAR models, a microplastic toxicity prediction model is constructed, which solves the problems of low efficiency, high cost and insufficient prediction performance of microplastics in the prior art, and achieves high-precision toxicity prediction and toxicity mechanism analysis, reduces experimental costs, and has broad applicability and significant benefits.

CN120015125APending Publication Date: 2025-05-16NANJING TECH UNIV

Patent Information

Application Number
CN202510111373.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The prior art has problems of low efficiency, high cost and insufficient prediction performance in microplastic toxicity prediction. In particular, traditional QSAR models ignore gene expression data and are difficult to meet the needs of microplastic toxicity prediction.

Method used

By combining transcriptomics and QSAR models, a prediction model of microplastic toxicity was constructed. The specific steps include selecting a variety of microplastics, measuring their physical and chemical characteristics and intracellular gene changes, screening out genes with common differential expression, selecting theoremological and gene expression descriptors, establishing training and test sets, using a variety of machine learning algorithms to build a QSAR model, and determining the scope of application of the model through performance verification and application domain analysis.

Benefits of technology

It significantly improves the accuracy of microplastic toxicity prediction, reveals the toxicity mechanism of microplastics, reduces experimental costs, reduces dependence on experimental animals, complies with the 3R principle, and has broad applicability and significant social and economic benefits.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120015125A_ABST
    Figure CN120015125A_ABST
Patent Text Reader

Abstract

The invention discloses a construction method of a microplastic toxicity prediction model based on transcriptomics and a QSAR (Quantitative Synthetic Aperture Radar) model, which comprises the following steps: measuring actual observation values of cell survival rates of different microplastic samples after contamination and changes of genes in the cells, and screening out the genes as specific genes; selecting a physicochemical descriptor and then determining a gene expression descriptor of the specific gene; establishing a training set and a test set; a plurality of machine learning algorithms are utilized to construct QSAR models respectively, and an optimal QSAR model is selected according to a performance evaluation result; and carrying out performance verification on the optimal QSAR model, and screening out a microplastic toxicity prediction range suitable for the optimal QSAR model. The construction method of the microplastic toxicity prediction model can realize reliable prediction of toxicity, and can effectively support safety evaluation and biotoxicity mechanism research of microplastics.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of bioinformatics and computational toxicology, and in particular to a method for constructing a microplastic toxicity prediction model based on transcriptomics and QSAR models. Background Art

[0002] In the field of environmental science and toxicology, the pollution of microplastics has become a global focus. Microplastics refer to plastic particles with a diameter of less than 5 mm. They are widely distributed in the ocean, soil, air and various organisms, posing a potential threat to ecosystems and human health. In order to evaluate the biological toxicity of microplastics, traditional methods mainly rely on laboratory cytotoxicity experiments and animal experiments. Although these methods are direct and effective, they have disadvantages such as being time-consuming, costly and inefficient.

[0003] In order to improve the efficiency and accuracy of microplastic toxicity prediction, quantitative structure-activity relationship (QSAR) models are widely used to predict the biological toxicity of compounds. The QSAR model predicts the biological toxicity of unknown compounds by establishing a mathematical relationship between the compound structure and biological activity. However, traditional QSAR models mainly consider the physical and chemical properties of the substance itself and ignore biological information such as gene expression data, which limits the improvement of the model's prediction performance.

[0004] In recent years, the development of transcriptomics technology has provided a new perspective for studying the biological effects of compounds. Transcriptomics can comprehensively analyze the expression changes of all genes in cells under specific conditions, providing rich molecular information for understanding the toxic mechanism of compounds. Combining transcriptomics data with QSAR models can more comprehensively describe the toxic effects of compounds and improve the accuracy and reliability of prediction models.

[0005] Although some studies have attempted to apply transcriptomics data to QSAR models, most of these studies are limited to specific compounds or biomarkers and lack systematic research on microplastics as a special pollutant. In addition, existing studies are relatively conservative in the selection of descriptors and modeling methods, which makes it difficult to meet the needs of microplastic toxicity prediction. Summary of the invention

[0006] The purpose of the invention is to provide a method for constructing a microplastic toxicity prediction model based on transcriptomics and QSAR models, which can achieve reliable prediction of toxicity and effectively support the safety assessment of microplastics and the study of biological toxicity mechanisms.

[0007] Technical solution: The method for constructing a microplastic toxicity prediction model based on transcriptomics and QSAR model described in the present invention comprises the following steps:

[0008] Step 1: Select a variety of microplastics commonly found in the environment, and measure the physical and chemical characteristics of various microplastics. Apply various microplastic samples to cells respectively, measure the actual observed values ​​of cell survival rates and changes in genes inside cells after exposure to different types of microplastic samples, and select genes with common differential expressions as specific genes based on changes in genes inside cells;

[0009] Step 2, selecting the physicochemical descriptors for the QSAR model according to the degree of cytotoxicity, then determining the gene expression descriptors of specific genes in the cells after exposure, and then normalizing the physicochemical descriptors and gene expression descriptors;

[0010] Step 3, using the actual observed values ​​of the cell survival rates of cells after exposure to a portion of microplastic samples and the normalized physicochemical descriptors and gene expression descriptors to establish a training set, and using the actual observed values ​​of the cell survival rates of cells after exposure to the remaining microplastic samples and the normalized physicochemical descriptors and gene expression descriptors to establish a test set;

[0011] Step 4: Use a variety of machine learning algorithms to build QSAR models respectively, use the normalized physicochemical descriptors and gene expression descriptors as independent variables of the QSAR model, use the model prediction value of cell survival rate as the dependent variable of the QSAR model, use the training set to train each QSAR model, and then use the test set to evaluate the performance of each trained QSAR model, and select the optimal QSAR model according to the performance evaluation results;

[0012] Step 5: Perform performance verification on the selected optimal QSAR model. After the performance verification is qualified, analyze the application domain of the optimal QSAR model to screen out the microplastic toxicity prediction range applicable to the optimal QSAR model.

[0013] Furthermore, in step 1, the specific steps of screening out the commonly differentially expressed genes as specific genes are as follows:

[0014] Step 1.1, extracting total RNA from the infected cells and further enriching mRNA;

[0015] Step 1.2, reverse transcribe the extracted mRNA into cDNA, and then sequence the cDNA library using high-throughput sequencing technology;

[0016] In step 1.3, the differentially expressed genes were screened out using the criteria of |log2FC| ≥ 1.5 and p-value < 0.05, and the differentially expressed genes common to each group were screened out using the Venn diagram, and the common differentially expressed genes were taken as specific genes.

[0017] Furthermore, in step 2, the physicochemical descriptors selected for the QSAR model include the concentration, particle size, potential, type and shape of microplastics according to the degree of cytotoxicity.

[0018] Furthermore, in step 2, the specific steps of determining the gene expression descriptors of specific genes in cells after infection are as follows:

[0019] Step 2.1, blow the cells that have been poisoned by microplastics off the culture dish, centrifuge at low temperature and transfer the upper aqueous phase;

[0020] Step 2.2, add an equal volume of isopropanol, place at -20°C for 12 h, and repeat centrifugation until a white precipitate appears;

[0021] Step 2.3, wash the RNA precipitate, repeat centrifugation, retain the RNA precipitate, and test the purity and concentration of the RNA;

[0022] Step 2.4, adding primers of specific genes to the extracted RNA for PCR amplification, and then determining the expression level of the specific gene.

[0023] Furthermore, in step 4, when using the test set to evaluate the performance of each trained QSAR model:

[0024] First, the test set is input into each trained QSAR model for prediction, and each trained QSAR model outputs the corresponding model prediction value;

[0025] Then, the difference evaluation index used to characterize the difference between the model prediction value and the actual observation value of each trained QSAR model is calculated;

[0026] Finally, each trained QSAR model is evaluated according to the difference evaluation index, and the QSAR model with the best difference evaluation index is selected as the optimal QSAR model.

[0027] Furthermore, in step 4, the difference evaluation indicators include mean absolute error, multiple correlation coefficient, root mean square error and standard error. When the difference evaluation indicators are optimal, the mean absolute error between the model prediction value and the actual observation value of the QSAR model is minimized, the multiple correlation coefficient is maximized, the root mean square error is minimized, and the standard error is minimized.

[0028] Furthermore, in step 4, the machine learning algorithms used to construct the QSAR model include a random forest algorithm, a support vector machine algorithm, a decision tree algorithm, a gradient boosting decision tree algorithm, and an extreme gradient boosting algorithm.

[0029] Furthermore, in step 5, when the performance of the selected optimal QSAR model is verified:

[0030] First, the training set validation is performed: the physicochemical descriptors and gene expression descriptors of each microplastic sample in the training set are substituted into the selected optimal QSAR model to obtain the model-predicted value of the cell survival rate of each microplastic sample in the training set; then the training set validation index used to characterize the difference between the model-predicted value and the actual observed value of the optimal QSAR model is calculated; if the training set validation index meets the corresponding index validation requirements, the training set validation of the optimal QSAR model is determined to be qualified;

[0031] Then, the test set validation is performed: the physicochemical descriptors and gene expression descriptors of each microplastic sample in the test set are substituted into the selected optimal QSAR model to obtain the model-predicted value of the cell survival rate of each microplastic sample in the test set; the test set validation index used to characterize the difference between the model-predicted value and the actual observed value of the optimal QSAR model is then calculated; if the test set validation index meets the corresponding index validation requirements, the test set validation of the optimal QSAR model is determined to be qualified;

[0032] Finally, a judgment is made: if both the training set verification and the test set verification are qualified, the performance verification is judged to be qualified.

[0033] Furthermore, in step 5, the training set validation indicators include mean absolute error, multiple correlation coefficient, root mean square error, standard error, and robustness Q 2 LOO , and when the mean absolute error is less than 0.4, the multiple correlation coefficient is greater than 0.6, the root mean square error is less than 0.4, the standard error is less than 0.4, and the robustness Q 2 LOO When it is greater than 0.5, the training set validation of the optimal QSAR model is considered qualified; the validation indicators of the test set include the prediction ability Q 2 ext , and when the prediction ability Q 2 ext When the value is greater than 0.5, the test set of the optimal QSAR model is judged to be qualified.

[0034] Furthermore, in step 5, the specific steps for screening out the microplastic toxicity prediction range applicable to the optimal QSAR model are as follows:

[0035] First, the mean and standard deviation of the physicochemical descriptors and the mean and standard deviation of the gene expression descriptors of each microplastic sample in the training set were calculated;

[0036] Then, the optimal QSAR model is used to predict the microplastic samples in the training set and the test set to obtain the model-predicted value of the cell survival rate of each microplastic sample, and then the residual between the model-predicted value and the actual observed value of the cell survival rate of each microplastic sample is calculated, and the residual is normalized to obtain the normalized residual res;

[0037] Then a Williams diagram is drawn, the horizontal axis of the Williams diagram is the leverage value hi of the microplastic sample, and the vertical axis of the Williams diagram is the normalized residual res of the microplastic sample. Then the warning value h* of the microplastic sample is calculated based on the leverage value hi and the normalized residual res. The leverage value hi is used to evaluate the impact of the sample on the optimal QSAR model, and the normalized residual res is used to evaluate the deviation between the model prediction value and the actual observation value.

[0038] Then, in the Williams diagram, the area with the horizontal axis between 0 and the warning value h* and the vertical axis between the normalized residual res-3 and the normalized residual res+3 is defined as the application domain of the optimal QSAR model;

[0039] Finally, it is determined whether the model-predicted value and the actual observed value of the cell survival rate of each microplastic sample fall within the application domain. If they do, it indicates that the prediction of the optimal QSAR model is reliable; otherwise, it indicates that the prediction of the optimal QSAR model is unreliable.

[0040] Compared with the prior art, the present invention has the following beneficial effects: (1) Improving prediction accuracy: By combining the physicochemical properties of microplastics with the expression levels of specific genes, the QSAR model of the present invention significantly improves the accuracy of predicting the toxicity of microplastics; (2) Revealing the toxicity mechanism: The present invention reveals the toxicity mechanism of microplastics by integrating gene expression data. The analysis of this mechanism provides new insights into the toxicity research of microplastics; (3) Reducing experimental costs: Compared with traditional in vivo and in vitro experiments, the QSAR model of the present invention can quickly evaluate the toxicity of microplastics through computational methods, reducing the time and cost required for experiments. The development of this prediction model provides an efficient and economical technical tool for research in the field of environmental toxicology, which can greatly reduce the dependence on experimental animals. It complies with the 3R (replacement, reduction, optimization) principle; (4) Wide applicability: The QSAR model constructed by the present invention has wide applicability. The application domain of the model is determined by the Williams diagram method, and its good predictive performance for different types of microplastics and under different experimental conditions is verified, making this method not only suitable for the toxicity assessment of microplastics, but also can be extended to the toxicity prediction of other types of pollutants; (5) Social and economic benefits: The present invention helps to improve the early warning capability of microplastic pollution in the environment, and supports the government and industry in providing a scientific basis when formulating relevant environmental protection standards and policies. In addition, the present invention can reduce the cost of toxicity assessment and promote the development of more efficient and low-cost environmental risk assessment research, which has significant social and economic benefits. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 is a flow chart of the method of the present invention;

[0042] Figure 2This is a diagram showing the prediction results of the five models of the present invention for predicting the toxicity of microplastics to BEAS-2B cells;

[0043] Figure 3 It is the Williams graph of the random forest model of the present invention. DETAILED DESCRIPTION

[0044] The technical solution of the present invention is described in detail below with reference to the accompanying drawings, but the protection scope of the present invention is not limited to the embodiments.

[0045] like Figure 1 As shown, the method for constructing a microplastic toxicity prediction model based on transcriptomics and QSAR models disclosed in the present invention comprises the following steps:

[0046] Step 1, select a variety of microplastics commonly found in the environment, and measure the physical and chemical characteristics of various microplastics. Apply various microplastic samples to cells respectively, and measure the actual observed values ​​of cell survival rate and changes in genes inside cells after exposure to different types of microplastic samples. The cell survival rate is measured after 24 hours of CCK-8 experiment, and the changes in genes inside cells are measured using RNA-seq (transcriptomics sequencing technology), and the genes with common differential expression are screened as specific genes based on the changes in genes inside cells;

[0047] Step 2: Select the physicochemical descriptors for the QSAR model according to the degree of cytotoxicity, then determine the gene expression descriptors of specific genes in the cells after exposure, and then normalize the physicochemical descriptors and gene expression descriptors to make the data ranges of different descriptors consistent to avoid deviations in the model;

[0048] Step 3: Use the actual observed values ​​of cell survival rates of cells after exposure to some microplastic samples and the normalized physicochemical descriptors and gene expression descriptors to establish a training set, and use the actual observed values ​​of cell survival rates of cells after exposure to the remaining microplastic samples and the normalized physicochemical descriptors and gene expression descriptors to establish a test set. The ratio of microplastic samples in the training set and the test set is 3:1 to ensure the generalization ability and prediction performance of the model.

[0049] Step 4: Use a variety of machine learning algorithms to build QSAR models respectively, use the normalized physicochemical descriptors and gene expression descriptors as independent variables of the QSAR model, use the model prediction value of cell survival rate as the dependent variable of the QSAR model, use the training set to train each QSAR model, and then use the test set to evaluate the performance of each trained QSAR model, and select the optimal QSAR model according to the performance evaluation results;

[0050] Step 5: Perform performance verification on the selected optimal QSAR model. After the performance verification is qualified, analyze the application domain of the optimal QSAR model to screen out the microplastic toxicity prediction range applicable to the optimal QSAR model.

[0051] Furthermore, in step 1, the specific steps of screening out the commonly differentially expressed genes as specific genes are as follows:

[0052] Step 1.1, extracting total RNA from the infected cells and further enriching mRNA;

[0053] Step 1.2, reverse transcribe the extracted mRNA into cDNA, and then sequence the cDNA library using high-throughput sequencing technology;

[0054] In step 1.3, the differentially expressed genes were screened out using the criteria of |log2FC| ≥ 1.5 and p-value < 0.05, and the differentially expressed genes common to each group were screened out using the Venn diagram, and the common differentially expressed genes were taken as specific genes.

[0055] Furthermore, in step 2, the physicochemical descriptors used in the QSAR model are selected based on the degree of cytotoxicity, including the concentration, particle size, potential, type and shape of microplastics, which are selected by comprehensively considering the physical and chemical characteristics of microplastics.

[0056] Furthermore, in step 2, when determining the gene expression descriptor of a specific gene in the infected cells, RT-PCR is used to determine the expression level of the specific gene in the infected cells, and the specific steps are as follows:

[0057] Step 2.1, blow the cells that have been poisoned by microplastics off the culture dish, centrifuge at low temperature and transfer the upper aqueous phase;

[0058] Step 2.2, add an equal volume of isopropanol, place at -20°C for 12 h, and repeat centrifugation until a white precipitate appears;

[0059] Step 2.3, wash the RNA precipitate, repeat centrifugation, retain the RNA precipitate, and test the purity and concentration of the RNA to ensure subsequent experiments;

[0060] Step 2.4, adding primers of specific genes to the extracted RNA for PCR amplification, and then determining the expression level of the specific gene.

[0061] Furthermore, in step 4, when using the test set to evaluate the performance of each trained QSAR model:

[0062] First, the test set is input into each trained QSAR model for prediction, and each trained QSAR model outputs the corresponding model prediction value;

[0063] Then, the difference evaluation index used to characterize the difference between the model prediction value and the actual observation value of each trained QSAR model is calculated;

[0064] Finally, each trained QSAR model is evaluated according to the difference evaluation index, and the QSAR model with the best difference evaluation index is selected as the optimal QSAR model.

[0065] Furthermore, in step 4, the difference evaluation indicators include mean absolute error (AAE), multiple correlation coefficient (R 2 ), root mean square error (RMSE) and standard error (SE) to quantitatively characterize the difference between the model prediction value and the actual observation value. When the difference evaluation index is optimal, the average absolute error (AAE) between the model prediction value and the actual observation value of the QSAR model is the smallest, and the multiple correlation coefficient (R 2 ), the root mean square error (RMSE) is minimized, and the standard error (SE) is minimized.

[0066] Furthermore, in step 4, the machine learning algorithms used to construct the QSAR model include a random forest algorithm, a support vector machine algorithm, a decision tree algorithm, a gradient boosting decision tree algorithm, and an extreme gradient boosting algorithm.

[0067] Furthermore, in step 5, when the performance of the selected optimal QSAR model is verified:

[0068] First, the training set validation is performed: the physicochemical descriptors and gene expression descriptors of each microplastic sample in the training set are substituted into the selected optimal QSAR model to obtain the model-predicted value of the cell survival rate of each microplastic sample in the training set; then the training set validation index used to characterize the difference between the model-predicted value and the actual observed value of the optimal QSAR model is calculated; if the training set validation index meets the corresponding index validation requirements, the training set validation of the optimal QSAR model is determined to be qualified;

[0069] Then, the test set validation is performed: the physicochemical descriptors and gene expression descriptors of each microplastic sample in the test set are substituted into the selected optimal QSAR model to obtain the model-predicted value of the cell survival rate of each microplastic sample in the test set; the test set validation index used to characterize the difference between the model-predicted value and the actual observed value of the optimal QSAR model is then calculated; if the test set validation index meets the corresponding index validation requirements, the test set validation of the optimal QSAR model is determined to be qualified;

[0070] Finally, a judgment is made: if both the training set verification and the test set verification are qualified, the performance verification is judged to be qualified.

[0071] Furthermore, in step 5, the training set validation indicators include mean absolute error, multiple correlation coefficient, root mean square error, standard error, and robustness Q 2 LOO , and when the mean absolute error is less than 0.4, the multiple correlation coefficient is greater than 0.6, the root mean square error is less than 0.4, the standard error is less than 0.4, and the robustness Q 2 LOO When it is greater than 0.5, the training set validation of the optimal QSAR model is considered qualified; the validation indicators of the test set include the prediction ability Q 2 ext , and when the prediction ability Q 2 ext When the value is greater than 0.5, the test set of the optimal QSAR model is judged to be qualified.

[0072] Furthermore, in step 5, the specific steps for screening out the microplastic toxicity prediction range applicable to the optimal QSAR model are as follows:

[0073] First, the mean and standard deviation of the physicochemical descriptors and the mean and standard deviation of the gene expression descriptors of each microplastic sample in the training set were calculated;

[0074] Then, the optimal QSAR model is used to predict the microplastic samples in the training set and the test set to obtain the model-predicted value of the cell survival rate of each microplastic sample, and then the residual between the model-predicted value and the actual observed value of the cell survival rate of each microplastic sample is calculated, and the residual is normalized to obtain the normalized residual res;

[0075] Then a Williams diagram is drawn. The horizontal axis of the Williams diagram is the leverage value hi of the microplastic sample, and the vertical axis of the Williams diagram is the normalized residual res of the microplastic sample. Then the warning value h* of the microplastic sample is calculated based on the leverage value hi and the normalized residual res. The leverage value hi is used to evaluate the impact of the sample on the optimal QSAR model, and the normalized residual res is used to evaluate the deviation between the model prediction value and the actual observation value. The calculation formulas of the leverage value hi and the warning value h* are:

[0076]

[0077] h * =h i +2*res

[0078] In the formula, x i is the descriptor of the i-th sample, X is the descriptor matrix of the training set samples, and res is the normalized residual;

[0079] Then, in the Williams diagram, the area with the horizontal axis between 0 and the warning value h* and the vertical axis between the normalized residual res-3 and the normalized residual res+3 is defined as the application domain of the optimal QSAR model;

[0080] Finally, it is determined whether the model-predicted value and the actual observed value of the cell survival rate of each microplastic sample fall within the application domain. If they do, it indicates that the prediction of the optimal QSAR model is reliable; otherwise, it indicates that the prediction of the optimal QSAR model is unreliable.

[0081] In order to verify the reliability of the construction method of the microplastic toxicity prediction model of the present invention, the toxicity prediction experiment is verified by taking BEAS-2B cells as an example:

[0082] 1. Data preparation

[0083] First, the physical and chemical characteristics of five types of microplastics were collected, including PVC, PE, PS, PP and PET, including hydrated particle size (AS), Zeta potential (ZP), concentration (C) and shape (Sh). At the same time, BEAS-2B cells were treated with microplastics, their cell survival rate was measured, and the expression data of the SGK3 gene log2FC was obtained through RT-PCR experiments.

[0084] 2. Feature Selection

[0085] The physicochemical characteristics data of microplastics and the gene expression data of SGK3 gene were selected as the input features of the model. The selected physicochemical characteristics data were AS, ZP, C and Sh. These input features were used to describe the physicochemical properties and biological responses of microplastics.

[0086] 3. Model construction

[0087] Five machine learning algorithms, random forest (RF), support vector machine (SVM), decision tree (DT), gradient boosted decision tree (GBDT) and extreme gradient boosting (XGBOOST), were used to construct QSAR models, which associated input features with the viability (CV%) of BEAS-2B cells.

[0088] 4. Model training and verification

[0089] Using the collected data set, the data was divided into a training set and a test set in a ratio of 3:1. The QSAR model was trained using the training set data, and the performance of the QSAR model was evaluated using the test set data. The running results of the five QSAR models are shown in Figure 2. Figure 2 The performance of the five QSAR models was evaluated by the multiple correlation coefficient (R 2), root mean square error (RMSE), mean absolute error (AAE) and standard error (SE). The evaluation results are shown in Table 1. The random forest model has the highest prediction accuracy, R2=0.8947, RMSE=0.0301, which is better than other QSAR models.

[0090] Table 1

[0091] Model <![CDATA[R 2 ]]> RMSE AAE (%) SE RF 0.8947 0.0301 0.0238 0.0591 SVM 0.3304 0.0697 0.0570 0.0193 DT 0.3304 0.0697 0.0599 0.0893 GBDT 0.8781 0.0311 0.0242 0.0706 XGBOOST 0.8939 0.0302 0.0240 0.0858

[0092] By R 2 , Q 2 LOO , RMSE, AAE, SE and Q 2 ext The performance of the random forest model was evaluated by using indicators such as , and the QSAR model was further verified. The evaluation results are shown in Table 2.

[0093]

[0094] As can be seen from Table 2, the random forest model has good stability and good predictive ability, and can pass external verification.

[0095] 5. Use Williams diagram to visualize the application range of the optimal QSAR model (RF) with the best performance, and check the outliers of the test set to evaluate the stability of the optimal QSAR model. The horizontal and vertical axes represent the leverage value hi and the normalized residual res of the sample. Figure 3 As shown, the optimal QSAR model constructed in this example has good stability.

[0096] Abbreviation or code description:

[0097] -QSAR: Quantitative Structure-Activity Relationship

[0098] -RF: Random Forest

[0099] -SVM: Support Vector Machine

[0100] -DT: Decision Tree

[0101] -GBDT: Gradient Boosting Decision Tree

[0102] -XGBOOST: Extreme Gradient Boosting

[0103] -AS: Hydrated Particle Size

[0104] -ZP:Zeta Potential

[0105] -C: Concentration

[0106] -Sh: Shape

[0107] -log2FC: Log2 Fold Change

[0108] -CV%: Cell Viability Percentage

[0109] -R 2 :Coefficient of Determination

[0110] -RMSE: Root Mean Square Error

[0111] -AAE: Average Absolute Error

[0112] -SE: Standard Error

[0113] -SGK3: serum and glucocorticoid-inducible kinase 3

[0114] As described above, although the present invention has been shown and described with reference to specific preferred embodiments, it should not be construed as limiting the present invention itself. Various changes in form and details may be made without departing from the spirit and scope of the present invention as defined in the appended claims.

Claims

1. A method for constructing a microplastic toxicity prediction model based on transcriptomics and QSAR models, characterized in that: The steps include: Step 1: Select a variety of microplastics commonly found in the environment, and measure the physical and chemical characteristics of various microplastics. Apply various microplastic samples to cells respectively, measure the actual observed values ​​of cell survival rates and changes in genes inside cells after exposure to different types of microplastic samples, and select genes with common differential expressions as specific genes based on changes in genes inside cells; Step 2, selecting the physicochemical descriptors for the QSAR model according to the degree of cytotoxicity, then determining the gene expression descriptors of specific genes in the cells after exposure, and then normalizing the physicochemical descriptors and gene expression descriptors; Step 3, using the actual observed values ​​of the cell survival rates of cells after exposure to a portion of microplastic samples and the normalized physicochemical descriptors and gene expression descriptors to establish a training set, and using the actual observed values ​​of the cell survival rates of cells after exposure to the remaining microplastic samples and the normalized physicochemical descriptors and gene expression descriptors to establish a test set; Step 4: Use a variety of machine learning algorithms to build QSAR models respectively, use the normalized physicochemical descriptors and gene expression descriptors as independent variables of the QSAR model, use the model prediction value of cell survival rate as the dependent variable of the QSAR model, use the training set to train each QSAR model, and then use the test set to evaluate the performance of each trained QSAR model, and select the optimal QSAR model according to the performance evaluation results; Step 5: Perform performance verification on the selected optimal QSAR model. After the performance verification is qualified, analyze the application domain of the optimal QSAR model to screen out the microplastic toxicity prediction range applicable to the optimal QSAR model.

2. The method for constructing a microplastic toxicity prediction model based on transcriptomics and QSAR model according to claim 1, characterized in that: In step 1, the specific steps of screening out the commonly differentially expressed genes as specific genes are as follows: Step 1.1, extracting total RNA from the infected cells and further enriching mRNA; Step 1.2, reverse transcribe the extracted mRNA into cDNA, and then sequence the cDNA library using high-throughput sequencing technology; In step 1.3, differentially expressed genes were screened out using the criteria of |log2FC| ≥ 1.5 and p-value < 0.05, and then the differentially expressed genes common to each group were screened out using the Venn diagram, and the common differentially expressed genes were taken as specific genes.

3. The method for constructing a microplastic toxicity prediction model based on transcriptomics and QSAR model according to claim 1, characterized in that: In step 2, the physicochemical descriptors selected for the QSAR model include the concentration, particle size, potential, type and shape of microplastics according to the degree of cytotoxicity.

4. The method for constructing a microplastic toxicity prediction model based on transcriptomics and QSAR model according to claim 1, characterized in that: In step 2, the specific steps of determining the gene expression descriptors of specific genes in cells after infection are as follows: Step 2.1, blow the cells that have been poisoned by microplastics off the culture dish, centrifuge at low temperature and transfer the upper aqueous phase; Step 2.2, add an equal volume of isopropanol, place at -20°C for 12 h, and repeat centrifugation until a white precipitate appears; Step 2.3, wash the RNA precipitate, repeat centrifugation, retain the RNA precipitate, and test the purity and concentration of the RNA; Step 2.4, adding primers of specific genes to the extracted RNA for PCR amplification, and then determining the expression level of the specific gene.

5. The method for constructing a microplastic toxicity prediction model based on transcriptomics and QSAR model according to claim 1, characterized in that: In step 4, when using the test set to evaluate the performance of each trained QSAR model: First, the test set is input into each trained QSAR model for prediction, and each trained QSAR model outputs the corresponding model prediction value; Then, the difference evaluation index used to characterize the difference between the model prediction value and the actual observation value of each trained QSAR model is calculated; Finally, each trained QSAR model is evaluated according to the difference evaluation index, and the QSAR model with the best difference evaluation index is selected as the optimal QSAR model.

6. The method for constructing a microplastic toxicity prediction model based on transcriptomics and QSAR model according to claim 5, characterized in that: In step 4, the difference evaluation indicators include mean absolute error, multiple correlation coefficient, root mean square error and standard error. When the difference evaluation indicators are optimal, the mean absolute error between the model prediction value and the actual observation value of the QSAR model is minimized, the multiple correlation coefficient is maximized, the root mean square error is minimized, and the standard error is minimized.

7. The method for constructing a microplastic toxicity prediction model based on transcriptomics and QSAR model according to claim 1, characterized in that: In step 4, the machine learning algorithms used to construct the QSAR model include random forest algorithm, support vector machine algorithm, decision tree algorithm, gradient boosting decision tree algorithm and extreme gradient boosting algorithm.

8. The method for constructing a microplastic toxicity prediction model based on transcriptomics and QSAR model according to claim 1, characterized in that: In step 5, when the performance of the selected optimal QSAR model is verified: First, the training set validation is performed: the physicochemical descriptors and gene expression descriptors of each microplastic sample in the training set are substituted into the selected optimal QSAR model to obtain the model-predicted value of the cell survival rate of each microplastic sample in the training set; then the training set validation index used to characterize the difference between the model-predicted value and the actual observed value of the optimal QSAR model is calculated; if the training set validation index meets the corresponding index validation requirements, the training set validation of the optimal QSAR model is determined to be qualified; Then, the test set validation is performed: the physicochemical descriptors and gene expression descriptors of each microplastic sample in the test set are substituted into the selected optimal QSAR model to obtain the model-predicted value of the cell survival rate of each microplastic sample in the test set; the test set validation index used to characterize the difference between the model-predicted value and the actual observed value of the optimal QSAR model is then calculated; if the test set validation index meets the corresponding index validation requirements, the test set validation of the optimal QSAR model is determined to be qualified; Finally, a judgment is made: if both the training set verification and the test set verification are qualified, the performance verification is judged to be qualified.

9. The method for constructing a microplastic toxicity prediction model based on transcriptomics and QSAR model according to claim 8, characterized in that: In step 5, the training set validation indicators include mean absolute error, multiple correlation coefficient, root mean square error, standard error, and robustness Q 2 LOO , and when the mean absolute error is less than 0.4, the multiple correlation coefficient is greater than 0.6, the root mean square error is less than 0.4, the standard error is less than 0.4, and the robustness Q 2 LOO When it is greater than 0.5, the training set validation of the optimal QSAR model is considered qualified; the validation indicators of the test set include the prediction ability Q 2 ext , and when the prediction ability Q 2 ext When the value is greater than 0.5, the test set of the optimal QSAR model is judged to be qualified.

10. The method for constructing a microplastic toxicity prediction model based on transcriptomics and QSAR model according to claim 1, characterized in that: In step 5, the specific steps for selecting the optimal QSAR model for the prediction range of microplastic toxicity are as follows: First, the mean and standard deviation of the physicochemical descriptors and the mean and standard deviation of the gene expression descriptors of each microplastic sample in the training set were calculated; Then, the optimal QSAR model is used to predict the microplastic samples in the training set and the test set to obtain the model-predicted value of the cell survival rate of each microplastic sample, and then the residual between the model-predicted value and the actual observed value of the cell survival rate of each microplastic sample is calculated, and the residual is normalized to obtain the normalized residual res; Then a Williams diagram is drawn, the horizontal axis of the Williams diagram is the leverage value hi of the microplastic sample, and the vertical axis of the Williams diagram is the normalized residual res of the microplastic sample. Then the warning value h* of the microplastic sample is calculated based on the leverage value hi and the normalized residual res. The leverage value hi is used to evaluate the impact of the sample on the optimal QSAR model, and the normalized residual res is used to evaluate the deviation between the model prediction value and the actual observation value. Then, in the Williams diagram, the area with the horizontal axis between 0 and the warning value h* and the vertical axis between the normalized residual res-3 and the normalized residual res+3 is defined as the application domain of the optimal QSAR model; Finally, it is determined whether the model-predicted value and the actual observed value of the cell survival rate of each microplastic sample fall within the application domain. If they do, it indicates that the prediction of the optimal QSAR model is reliable; otherwise, it indicates that the prediction of the optimal QSAR model is unreliable.

Citation Information

Patent Citations

  • Method for predicting organic matter adsorption of micro-plastics by constructing quantitative structure-activity relationship based on quantum chemical parameters

    CN111613276A

  • In vitro toxicogenomics for toxicity prediction

    US20170270239A1

Cited By

  • Microplastic toxicity prediction method integrating machine learning and meta analysis

    CN120951142A

  • Microplastic toxicity prediction method fusing machine learning and meta-analysis

    CN120951142B