Weak cell toxicity prediction model and prediction device and application thereof

By constructing a weak cytotoxicity prediction model based on gene expression data, and utilizing feature gene sets and various machine learning algorithms, the problem of existing models reflecting differences in substance concentration and target site effects was solved, achieving efficient and accurate prediction of cytotoxicity and improving the comprehensiveness of drug safety evaluation.

CN121148464APending Publication Date: 2025-12-16INST OF MATERIA MEDICA CHINESE ACAD OF MEDICAL SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410758607.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-06-13
Publication Date
2025-12-16

AI Technical Summary

Technical Problem

Existing cytotoxicity prediction models based on compound structure information are unable to accurately reflect substance concentration information and its effects on different target sites. Furthermore, the collinearity problem of high-dimensional gene expression data leads to a decrease in model reliability and accuracy, and differences in training data quality affect the model construction effect.

Method used

We constructed a weak cytotoxicity prediction model based on gene expression data using machine learning algorithms. We used a set of characteristic genes (TOP2A, CCNA2, VAT1, CDK1, POLE2, TSC22D3, HES1, MCM10, TMPO, CENPX) for modeling and combined algorithms such as random forest, support vector machine, XGBoost, and LightGBM to build an efficient and accurate prediction model.

Benefits of technology

It enables accurate prediction of weak cytotoxicity in different cell lines, provides a more comprehensive evaluation of drug safety, and is applicable to the prediction of cytotoxicity of small molecule compounds, proteins, peptides, antibodies, nucleic acids and other substances, thereby improving the efficiency of candidate substance discovery in the new drug development process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure FT_2
    Figure FT_2
  • Figure FT_3
    Figure FT_3
Patent Text Reader

Abstract

The invention belongs to the technical field of medicines, and discloses a weak cell toxicity prediction model, a weak cell toxicity prediction device and application of the weak cell toxicity prediction model. The weak cell toxicity prediction model and the weak cell toxicity prediction device are constructed based on simplified transcriptomics characteristics, and the transcriptomics characteristics of the model comprise TOP2A, CCNA2, VAT1, CDK1, POLE2, TSC22D3, HES1, MCM10, TMPO and CENPX. The invention also discloses application of the machine learning prediction model and the prediction device in prediction of weak cytotoxicity compounds, which is characterized by being applied to prediction of lung-derived, liver-derived, kidney-derived and intestine-derived cell lines, and can be applied to prediction of small-molecular compounds, proteins, nucleic acids and mixtures, such as small-molecular compounds, small-molecular compounds, small-molecular compounds, small-molecular compounds, small-molecular compounds, small-molecular compounds, small-molecular compounds, small-molecular compounds and small-molecular compounds. And an efficient, accurate and rapid prediction method is provided for safety evaluation of candidate substances in a new drug research and development process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of pharmaceutical technology, specifically relating to a weak cytotoxicity prediction model and device and its application. Background Technology

[0002] Drug safety is a key factor in the success or failure of new drug development (Fogel DB. Factors associated with clinical trials that fail and opportunities for improving the likelihood of success: a review. Contemporary clinical trials communications 2018; 11:156-64.). Establishing a reliable toxicity prediction model can rapidly screen high-safety candidate substances in the early stages of drug development, which is of great significance (Badwan BA, et al. Machine learning approaches to predict drug efficacy and toxicity in oncology. Cell Reports Methods 2023; 3.). Cell viability is commonly used to measure the level of cell proliferation, survival, or death in normally growing cells after exposure to an analyte. It can reflect the toxicity of the analyte and is often used as an indicator for in vitro drug safety evaluation (Adan A, et al. Cell Proliferation and Cytotoxicity Assays. Curr Pharm Biotechnol 2016; 17:1213-21.; Johnson S, et al. Assessment of cell viability. Current protocols in cytometry 2013; 64:9.2.1-9.2.26.; Benbow JW, Aubrecht J, Banker MJ, Nettleton D, Aleo MD. Predicting safety toleration of pharmaceutical chemical leads: cytotoxicity correlations to exploratory toxicity studies. Toxicology Letters 2010; 197:175-82.). However, traditional experimental methods often consume a lot of human and material resources. Developing an efficient, accurate and low-cost method for assessing cell viability is of great significance for drug development.

[0003] All existing publicly available cell viability prediction models use the structural information of the analyte for modeling and prediction. For example, invention patent CN114974460 discloses a method for predicting the cytotoxicity of disinfection byproducts based on the molecular structure of compounds using machine learning algorithms; patent CN114171137 discloses a method for predicting the environmental hazards of compounds based on machine learning analysis of compound structural information.

[0004] However, changes in cell viability induced by substance treatment involve complex factors, and predictions based on substance structure information cannot directly reflect the concentration information of the substance and its differences in effects on different target sites. Compared with substance structure information, omics data can reflect changes in the expression levels of various genes in cells after treatment with the analyte, and can be used to systematically quantify the correlation between changes in gene expression and cell viability. It can also accurately reflect the overall effect of treating different cells / tissues at specific treatment concentrations. In addition, prediction models based on omics data can overcome the limitations of existing models in structural characterization and can be applied to predict the cell viability of any analyte with available omics data. This applies not only to small molecule compounds but also to proteins, peptides, antibodies, nucleic acids, and other substances, providing a more comprehensive perspective for drug safety evaluation.

[0005] Building predictive models based on gene expression data shows great promise, but the modeling process still faces certain challenges. Gene expression data is characterized by high dimensionality and collinearity, containing expression values ​​of tens of thousands of genes, and there are correlations between gene expressions. Modeling with all genes may lead to a significant decrease in the reliability, accuracy, and computational efficiency of the model. Therefore, identifying the gene set highly correlated with cell viability and building an efficient predictive model is of great significance. Furthermore, the performance of machine learning models largely depends on the quality of the training data (Gong Y, Liu G, Xue Y, Li R, Meng L. A survey on dataset quality in machine learning. Information and Software Technology 2023:107268.). Publicly available modeling data often comes from different databases with varying experimental conditions, leading to noise within the data and affecting the model's construction effectiveness and generalization ability. Addressing these challenges, identifying the characteristic gene set highly correlated with cell viability, optimizing the model building process, and constructing a highly robust cell viability prediction model have significant practical application value for drug safety assessment. Summary of the Invention

[0006] This invention overcomes the shortcomings of existing technologies by disclosing a weak cytotoxicity prediction model and device, as well as their applications. It overcomes the limitations of structure-based compound toxicity prediction in existing technologies, achieving a systematic identification of weak cytotoxicity of compounds at the gene level. The disclosed characteristic gene combinations and the prediction model constructed based on these combinations can achieve accurate prediction across various cell lines, thus providing an efficient, accurate, and rapid prediction method for discovering high-safety candidate substances in the new drug development process.

[0007] To achieve the above objectives, the present invention provides the following solution:

[0008] In a first aspect, the present invention provides a weak cytotoxicity prediction model based on gene expression data, characterized in that: the method for constructing the model includes the following steps:

[0009] S1: Obtain modeling data: Collect cell expression profile data and cell viability values ​​after treatment with the test substance. Label samples according to cell viability values. The weak toxicity threshold of the test substance is 80% cell viability: cell viability greater than 80% is a positive sample; cell viability less than or equal to 80% is considered to have weak toxicity and is a negative sample. Obtain the modeling dataset.

[0010] S2: Model Construction: Machine learning algorithms are used to fit the modeling dataset with modeling features, and the model hyperparameters are modeled and optimized to obtain a weak cytotoxicity prediction model. The modeling features include genes TOP2A, CCNA2, VAT1, CDK1, POLE2, TSC22D3, HES1, MCM10, TMPO, and CENPX.

[0011] Specifically, the transcriptome data comes from open-source databases GEO, ArrayExpress, CMap, LINCS, and L1000CDS2.

[0012] The machine learning algorithms include random forest, support vector machine, XGBoost, LightGBM, or a voting model of the above four machine learning models.

[0013] Specifically, the characteristic gene combination described in S2 includes one or more of the genes TOP2A, CCNA2, VAT1, CDK1, POLE2, TSC22D3, HES1, MCM10, TMPO, and CENPX.

[0014] Secondly, the present invention provides a weak cytotoxicity prediction device based on gene expression data, characterized in that it includes a data receiving module, a data processing module, and a result generation module.

[0015] The data receiving module is used to receive the expression values ​​of a combination of characteristic genes after cells are treated with the substance to be tested. The combination of characteristic genes includes the genes TOP2A, CCNA2, VAT1, CDK1, POLE2, TSC22D3, HES1, MCM10, TMPO, and CENPX.

[0016] The data processing module refers to the process of calling machine learning algorithms to calculate the input data based on the expression values ​​of characteristic gene combinations, and obtaining a prediction result on whether the test substance has cytotoxicity; the machine learning algorithm includes one or more of random forest, support vector machine, XGBoost and LightGBM;

[0017] The result generation module is used to send the prediction results out.

[0018] Specifically, the characteristic genes in the data receiving module include one or more of the following genes: TOP2A, CCNA2, VAT1, CDK1, POLE2, TSC22D3, HES1, MCM10, TMPO, and CENPX; the cell lines include lung-derived cell lines, kidney-derived cell lines, liver-derived cell lines, and intestinal-derived cell lines.

[0019] Furthermore, the data processing module specifically includes the following steps:

[0020] S1: Collect modeling data: Collect the expression values ​​of characteristic genes of cells treated with the test substance and the cell viability values ​​after treatment with the test substance. The characteristic genes are TOP2A, CCNA2, VAT1, CDK1, POLE2, TSC22D3, HES1, MCM10, TMPO, and CENPX. Cell viability greater than 80% is considered a positive sample; cell viability less than or equal to 80% is considered to have weak toxicity and is a negative sample. Obtain the modeling dataset.

[0021] S2: Model building: Apply machine learning algorithms to fit the modeling data to obtain a predictive model. The machine learning algorithms include one or more of random forest, support vector machine, XGBoost and LightGBM.

[0022] S3: Predict the weak cytotoxicity of the test substance: Call the prediction model described in S2 to calculate the change value of the characteristic gene expression of the test substance and obtain the prediction result of the weak cytotoxicity of the test substance.

[0023] The third aspect of the technical solution of the present invention is to provide the application of the weak cytotoxicity prediction model described in the first aspect and the weak cytotoxicity prediction device described in the second aspect of the present invention in predicting the cell viability value of the test substance in cell lines from different tissue sources.

[0024] The substances to be tested include small molecule compounds, proteins, nucleic acids, carbohydrates, mixtures, and nanomaterials. The cell lines include lung-derived cell lines, kidney-derived cell lines, liver-derived cell lines, and intestinal-derived cell lines.

[0025] Beneficial technical effects

[0026] This invention discloses a machine learning model and prediction device based on transcriptomics features for predicting weak cytotoxicity, and also discloses a set of characteristic gene combinations for weak cytotoxicity prediction. This enables accurate detection of the weak cytotoxicity of analytes and can be applied to predict the weak cytotoxicity of analytes in lung, liver, kidney, and intestinal cell lines. Furthermore, this invention overcomes the limitations of existing technologies in structural characterization and can be applied to predict the weak cytotoxicity of any analyte with available omics data. It can be applied not only to small molecule compounds but also to the cytotoxicity prediction of proteins, peptides, antibodies, nucleic acids, and other analytes, providing a more comprehensive perspective for drug safety evaluation. The implementation of this technology provides an efficient, accurate, and rapid prediction method for the safety assessment of candidate substances in the new drug development process, and offers some reference for the innovation of new technologies and methods in the field of predictive toxicology. Attached Figure Description

[0027] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0028] Figure 1 A weak cytotoxicity prediction model.

[0029] Figure 2 Distribution of cell viability in the external test set. AD represents the distribution of cell viability data for A549, HEK293, HepG2, and HT-29 cell lines treated with the analyte in the external test set, respectively.

[0030] Figure 3 Evaluation results of the weak cytotoxicity prediction model in lung-derived cell lines. AE represents the evaluation results of Support Vector Machine, Random Forest, XGBoost, LightGBM, and Voting model, respectively.

[0031] Figure 4 Evaluation results of the weak cytotoxicity prediction model on kidney-derived cell lines. AE represents the evaluation results of Support Vector Machine, Random Forest, XGBoost, LightGBM, and Voting model, respectively.

[0032] Figure 5Evaluation results of the weak cytotoxicity prediction model on liver-derived cell lines. AE represents the evaluation results of Support Vector Machine, Random Forest, XGBoost, LightGBM, and Voting model, respectively.

[0033] Figure 6 Evaluation results of the weak cytotoxicity prediction model on intestinal cell lines. AE represents the evaluation results of Support Vector Machine, Random Forest, XGBoost, LightGBM, and Voting model, respectively.

[0034] Figure 7 A weak cytotoxicity prediction device. Detailed Implementation

[0035] The objectives, advantages, and features of this invention will be illustrated and explained through the following non-limiting description of preferred embodiments. These embodiments are merely typical examples of applying the technical solutions of this invention, and all technical solutions formed by equivalent substitutions or equivalent transformations fall within the scope of protection claimed by this invention.

[0036] The invention discloses a set of characteristic genes and a cell viability prediction model for cell viability prediction. The method includes the following steps:

[0037] Example 1. A weak cytotoxicity prediction model

[0038] A weak cytotoxicity prediction model based on gene expression data, characterized in that the model construction method includes the following steps:

[0039] S1: Obtain modeling data: Collect cell expression profile data and cell viability values ​​after treatment with the test substance. Label samples according to cell viability values. The weak toxicity threshold of the test substance is 80% cell viability. Cell viability greater than 80% is a positive sample; cell viability less than or equal to 80% is considered to have weak toxicity and is a negative sample. Obtain the modeling dataset.

[0040] S2: Model Construction: Machine learning algorithms are used to model the features and optimize the model hyperparameters to obtain the predictive model. The modeling features include genes TOP2A, CCNA2, VAT1, CDK1, POLE2, TSC22D3, HES1, MCM10, TMPO, and CENPX.

[0041] Specifically, the open-source transcriptome data collection described in step S1 includes the following steps:

[0042] The expression profile data of the test substances were downloaded from Expanded CMap LINCS Resource 2020 (https: / / clue.io / data / CMap2020#LINCS2020). Transcriptome data of A549 cells were extracted after 24 hours of treatment with the test substances (final concentration 10 μM). The parse function of the cmapPy package was used to retrieve level 5 data based on the perturbation ID (Pert id) of the corresponding perturbation conditions of 1,243 compounds. The MODZ (Moderated z score, MODZ) method was applied to merge multiple expression profiles of the same compound: the pairwise Spearman correlation coefficients of multiple expression profiles were calculated as weights, with negative correlation set to a minimum value (0.01). The weighted average of gene expression in multiple expression profiles was calculated to obtain the gene expression profile used for modeling.

[0043] Specifically, step S1, which involves acquiring cell viability data, includes the following steps: ATP plays a crucial role in various physiological processes of cells, directly providing energy to the body. It is an important indicator reflecting cell viability and is positively correlated with the number of viable cells. Therefore, quantitative detection of ATP in cell lysates reflects the number of viable cells in the sample being tested. This invention employs... The Luminescent Cell Viability Assay kit (Promega) evaluates the effect of analytes on cell viability by detecting ATP-quantified cell viability after treatment with the analyte.

[0044] Cells (A549, HEK293, HepG2, and HT-29) were seeded in 96-well plates and cultured for 24 h. The corresponding assay substances were added, and the cells were incubated for 48 h. Cells were then lysed using lysis buffer (Cat. No. E1531; Promega; USA). The relative light units (RLUs) of the cell lysis buffer in each well were measured using the CellTiter-Glo kit, and cell viability was calculated using the following formula:

[0045] Cell viability (%) = RLUs 待测物质 / RLUs 溶剂 ×100%

[0046] Specifically, the machine learning modeling described in step S2 includes the following steps:

[0047] Predictive models were built using Support Vector Machine (svm.SVC from the Python sklearn package), Random Forest (RandomForestClassifier from the Python sklearn package), XGBoost (xgboost package), and LightGBM (lightgbm package). Bayesian optimization (bayes_opt package in Python) was used, employing Gaussian processes iteratively to find the optimal hyperparameter combination. The optimal hyperparameters for each model were obtained through 5x cross-validation. Furthermore, a voting method (votingClassifier(voting="soft") from sklearn.ensemble) was used to fuse the four models to obtain an ensemble model based on soft voting. Example 2. Collection of external test set data.

[0048] ATP plays a crucial role in various physiological processes of cells, directly providing energy to the body. It is an important indicator of cell viability and is positively correlated with the number of viable cells. Therefore, quantitative detection of ATP in cell lysates reflects the number of viable cells in the sample. This invention employs... The Luminescent Cell Viability Assay kit (Promega) evaluates the effect of analytes on cell viability by detecting ATP-quantified cell viability after treatment with the analyte.

[0049] Cells (A549, HEK293, HepG2, and HT-29) were seeded in 96-well plates and cultured for 24 h. The corresponding assay substances were added, and the cells were incubated for 48 h. Cells were then lysed using lysis buffer (Cat. No. E1531; Promega; USA). The relative light units (RLUs) of the cell lysis buffer in each well were measured using the CellTiter-Glo kit, and cell viability was calculated using the following formula:

[0050] Cell viability (%) = RLUs 待测物质 / RLUs 溶剂 ×100%

[0051] The effects of 350 test samples (including small molecule compounds, nucleic acids, cytokines, and antibodies; Table 1) on the cell viability of A549, HEK293, HepG2, and HT-29 cell lines were investigated (results are attached). Figure 2 ).

[0052] Table 1

[0053]

[0054]

[0055]

[0056] Example 3. Evaluation results of a weak cytotoxicity prediction model in lung-derived cell lines.

[0057] We invoked a weak cytotoxicity prediction model to evaluate its predictive ability on the A549 cell test set treated with the compound. The model predictions were evaluated using the following metrics: accuracy, precision, recall, specificity, and the area under the curve (AUC) obtained from the receiver operating characteristic curve (ROC).

[0058] Accuracy = (TP + TN) / (TP + TN + FP + FN)

[0059] Precision = TP / (TP + FP)

[0060] Recall = TP / (TP + FN)

[0061] Specificity = TN / (TN + FP)

[0062] Among them, TP (true positive), TN (true negative), FP (false positive), and FN (false negative) represent true positive, true negative, false positive, and false negative, respectively.

[0063] The predictive capabilities of the five models described above on the test set data of A549 cells treated with the test substance were evaluated respectively (see attached). Figure 3 The results showed that the AUROC scores of Support Vector Machine, Random Forest, XGBoost, LightGBM, and Voting Model on the A549 cell test set were 0.83, 0.82, 0.79, 0.79, and 0.80, respectively; the accuracy scores were 0.78, 0.81, 0.75, 0.74, and 0.78, respectively; the precision scores were 0.79, 0.83, 0.78, 0.79, and 0.81, respectively; the recall scores were 0.88, 0.88, 0.83, 0.79, and 0.84, respectively; and the specificity scores were 0.76, 0.77, 0.68, 0.66, and 0.71, respectively.

[0064] Example 4. Evaluation results of a weak cytotoxicity prediction model in kidney-derived cell lines ( Figure 3 )

[0065] We used a weak cytotoxicity prediction model to evaluate its predictive ability on the HEK293 cell test set treated with the compound. The model predictions were evaluated using the following metrics: accuracy, precision, recall, specificity, and the area under the curve (AUC) obtained from the receiver operating characteristic curve (ROC).

[0066] Accuracy = (TP + TN) / (TP + TN + FP + FN)

[0067] Precision = TP / (TP + FP)

[0068] Recall = TP / (TP + FN)

[0069] Specificity = TN / (TN + FP)

[0070] Among them, TP (true positive), TN (true negative), FP (false positive), and FN (false negative) represent true positive, true negative, false positive, and false negative, respectively.

[0071] The predictive capabilities of the above five models on the test set data of HEK293 cells treated with the test substance were evaluated respectively (see attached). Figure 4 The results showed that the AUROC scores of Support Vector Machine, Random Forest, XGBoost, LightGBM, and Voting Model on the HEK293 cell test set were 0.81, 0.80, 0.77, 0.77, and 0.80, respectively; the accuracy scores were 0.73, 0.72, 0.73, 0.72, and 0.73, respectively; the precision scores were 0.69, 0.70, 0.72, 0.72, and 0.72, respectively; the recall scores were 0.99, 0.93, 0.90, 0.89, and 0.91, respectively; and the specificity scores were 0.97, 0.81, 0.76, 0.74, and 0.78, respectively.

[0072] Example 5. Evaluation results of a weak cytotoxicity prediction model in liver-derived cell lines.

[0073] We used a weak cytotoxicity prediction model to evaluate its predictive ability on the HepG2 cell test set treated with the compound. The model predictions were evaluated using the following metrics: accuracy, precision, recall, specificity, and the area under the curve (AUC) obtained from the receiver operating characteristic curve (ROC).

[0074] Accuracy = (TP + TN) / (TP + TN + FP + FN)

[0075] Precision = TP / (TP + FP)

[0076] Recall = TP / (TP + FN)

[0077] Specificity = TN / (TN + FP)

[0078] Among them, TP (true positive), TN (true negative), FP (false positive), and FN (false negative) represent true positive, true negative, false positive, and false negative, respectively.

[0079] The predictive capabilities of the above five models on the test set data of HepG2 cells treated with the test substance were evaluated respectively (see attached). Figure 5 The results showed that the AUROC scores of Support Vector Machine, Random Forest, XGBoost, LightGBM, and Voting models on the HepG2 cell test set were 0.79, 0.77, 0.75, 0.75, and 0.76, respectively; the accuracy scores were 0.72, 0.73, 0.70, 0.70, and 0.74, respectively; the precision scores were 0.67, 0.71, 0.68, 0.69, and 0.72, respectively; the recall scores were 0.86, 0.78, 0.74, 0.72, and 0.78, respectively; and the specificity scores were 0.80, 0.75, 0.72, 0.71, and 0.76, respectively.

[0080] Example 6. Evaluation results of a weak cytotoxicity prediction model in gut-derived cell lines

[0081] We used a weak cytotoxicity prediction model to evaluate its predictive ability on the HT-29 cell test set treated with the compound. The model predictions were evaluated using the following metrics: accuracy, precision, recall, specificity, and the area under the curve (AUC) obtained from the receiver operating characteristic curve (ROC).

[0082] Accuracy = (TP + TN) / (TP + TN + FP + FN)

[0083] Precision = TP / (TP + FP)

[0084] Recall = TP / (TP + FN)

[0085] Specificity = TN / (TN + FP)

[0086] Among them, TP (true positive), TN (true negative), FP (false positive), and FN (false negative) represent true positive, true negative, false positive, and false negative, respectively.

[0087] The predictive capabilities of the five models described above on test set data of HT-29 cells treated with the test substance were evaluated respectively (see attached). Figure 6 The results showed that the AUROC scores of Support Vector Machine, Random Forest, XGBoost, LightGBM, and Voting models on the HT-29 cell test set were 0.79, 0.78, 0.78, 0.76, and 0.78, respectively; the accuracy scores were 0.75, 0.74, 0.74, 0.72, and 0.74, respectively; the precision scores were 0.77, 0.79, 0.79, 0.77, and 0.79, respectively; the recall scores were 0.82, 0.76, 0.75, 0.74, and 0.77, respectively; and the specificity scores were 0.71, 0.67, 0.66, 0.64, and 0.68, respectively.

[0088] Example 7. A weak cytotoxicity prediction device

[0089] A weak cytotoxicity prediction model based on gene expression data (with appendix) Figure 7 It includes: a data receiving module, a data processing module, and a result generation module.

[0090] (1) Data receiving module, used to receive the expression values ​​of characteristic genes after cells are treated with the test substance, including genes TOP2A, CCNA2, VAT1, CDK1, POLE2, TSC22D3, HES1, MCM10, TMPO, and CENPX;

[0091] (2) A data processing module, comprising a model unit and a prediction unit. The model unit includes a weak cytotoxicity prediction model; the prediction unit is responsible for calling the prediction model, calculating the expression values ​​of characteristic genes in cells treated with the input test substance, and obtaining a prediction result regarding whether the test substance possesses cytotoxicity; wherein, the prediction model preferably employs a random forest algorithm.

[0092] (3) Result generation module, used to send out the prediction results. The weak toxicity threshold of the test substance is 80% cell viability: cell viability greater than 80% is considered a non-toxic sample and a positive sample; cell viability less than or equal to 80% is considered to have weak toxicity and a negative sample.

[0093] This invention has many other embodiments, and all technical solutions formed by equivalent transformation or equivalent transformation fall within the protection scope of this invention.

Claims

1. A weak cytotoxicity prediction model based on gene expression data, characterized in that: The method for constructing this model includes the following steps: S1: Obtain modeling data: Collect cell expression profile data and cell viability values ​​after treatment with the test substance. Label samples according to cell viability values. The weak toxicity threshold of the test substance is 80% cell viability. Cell viability greater than 80% is a positive sample; cell viability less than or equal to 80% is considered to have weak toxicity and is a negative sample. Obtain the modeling dataset. S2: Model Construction: Machine learning algorithms are used to fit the modeling dataset with modeling features, and the model hyperparameters are modeled and optimized to obtain a weak cytotoxicity prediction model. The modeling features include genes TOP2A, CCNA2, VAT1, CDK1, POLE2, TSC22D3, HES1, MCM10, TMPO, and CENPX.

2. The weak cytotoxicity prediction model as described in claim 1, characterized in that: The machine learning algorithms include Random Forest, Support Vector Machine, XGBoost, LightGBM, and the voting model of the above four models.

3. A device for predicting weak cytotoxicity based on gene expression data, characterized in that: It includes a data receiving module, a data processing module, and a result generation module: The data receiving module is used to receive the expression values ​​of a combination of characteristic genes after cells are treated with the substance to be tested. The combination of characteristic genes includes the genes TOP2A, CCNA2, VAT1, CDK1, POLE2, TSC22D3, HES1, MCM10, TMPO, and CENPX. The data processing module refers to the process of calling machine learning algorithms to calculate the input data based on the expression values ​​of characteristic gene combinations, and obtaining a prediction result on whether the test substance has cytotoxicity; the machine learning algorithm includes one or more of random forest, support vector machine, XGBoost and LightGBM; The result generation module is used to send the prediction results out.

4. The attenuated cytotoxicity prediction device as described in claim 3, characterized in that: The cell lines mentioned in the data receiving module include lung-derived cell lines, kidney-derived cell lines, liver-derived cell lines, and intestinal-derived cell lines.

5. The application of the weak cytotoxicity prediction model according to any one of claims 1-2 or the cell viability prediction device according to any one of claims 3-4 in predicting the weak cytotoxicity of the analyte in cell lines from different tissue sources.

6. The application as described in claim 5, characterized in that: The cell lines include lung-derived cell lines, kidney-derived cell lines, liver-derived cell lines, and intestinal-derived cell lines.