A ChemBERTa-FP anti-cancer drug prediction method based on the combination of deep learning and the chemical field
Through the ChemBERTa-FP model combining deep learning and chemistry expertise, the ChemBERTa and BERT_base pre-trained models and molecular fingerprint features are used to solve the problem of insufficient prediction accuracy of anti-cancer drugs, and efficient and accurate drug activity prediction and toxicity identification are achieved.
Patent Information
- Application Number
- CN202411724850.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-28
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2044-11-28
AI Technical Summary
The prior art is difficult to effectively capture the characteristics and interactions of complex molecular structures in drug molecular prediction, resulting in insufficient prediction accuracy, especially in the field of prediction and screening of anti-cancer drugs.
A ChemBERTa-FP model combining deep learning and chemistry expertise was used to predict anti-cancer drug activity by using ChemBERTa and BERT_base as the basic pre-trained model, and combining molecular fingerprint features such as MACCS, PubChem and Pharmacophore ErG.
It significantly improves the accuracy and reliability of anti-cancer drug prediction, shortens the drug discovery cycle, can quickly complete activity prediction under standard computer configuration, and has the ability to identify normal cytotoxicity.
Smart Images

Figure CN119673317B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of medical materials for promoting wound healing, and particularly relates to a ChemBERTa-FP anti-cancer drug prediction method combining deep learning and the chemical field. Background Art
[0002] In the process of drug molecule discovery and design, accurately predicting the biological activity and pharmacological properties of molecules has always been a complex and time-consuming challenge. Traditional methods mainly rely on experimental data and empirical formulas, which are not only costly but also inefficient. In recent years, computer-aided screening technologies have shown great potential in this field. These technologies, including computer prediction models based on molecular descriptors / fingerprints and chemical language (SMILES), cover biological activity prediction, drug similarity prediction, toxicity prediction, etc., and can significantly reduce the over-reliance on time-consuming and labor-intensive experiments.
[0003] Among them, machine learning prediction models based on molecular descriptors / fingerprints to characterize chemical structures, including random forest (RF), support vector machine (SVM), extreme gradient boosting (XGBoost), K-nearest neighbor (KNN), etc., have been widely used in various prediction tasks such as biological activity, drug similarity, and toxicity. Although these methods reduce the dependence on experiments, they still have deficiencies in capturing the complex structural features and interactions of drug molecules, and the prediction accuracy needs to be improved.
[0004] Different from the descriptor / fingerprint-based methods, deep learning models based on natural language processing (NLP), such as BERT (Bidirectional Encoder Representations from Transformers), have been successfully applied to the field of text and language processing. The ChemBERTa model draws on the BERT architecture and is specifically designed for chemical language (SMILES), and is pre-trained on the SMILES sequences of 77 million unlabeled compounds in the PubChem library to fully mine the substructure or fragment information of drug molecules and deeply capture the complex molecular structural features and interaction relationships. By performing targeted fine-tuning on the pre-trained ChemBERTa model, competitive results have been shown in various molecular property prediction tasks.
[0005] However, the ChemBERTa algorithm has not yet been effectively applied to the field of anti-cancer drug prediction and screening, which is a largely underexplored potential area. Therefore, developing an efficient molecular prediction method combining deep learning and chemical expertise is of great significance for accelerating the research and development of anti-cancer drugs. The method of the present invention can not only improve the accuracy and reliability of prediction, but also significantly shorten the drug discovery cycle, providing strong impetus and support for the research and development of anti-cancer drugs. Summary of the Invention
[0006] The object of the present invention is to provide an innovative anti-cancer drug prediction method named ChemBERTa-FP, which is an improved ChemBERTa model that combines deep learning technology and molecular fingerprint features, aiming to predict the activity of anti-cancer drugs to improve the efficiency and accuracy of anti-cancer drug research and development. This method specifically addresses common problems in anti-cancer drug development, such as long R & D cycles and difficult-to-analyze complex molecular structures, and optimizes the prediction accuracy through advanced algorithms. The present invention proposes a model framework that combines deep learning and chemical knowledge, aiming to predict the activity of anti-cancer drugs.
[0007] To achieve the above object and other related objects, the technical solution provided by the present invention is: a ChemBERTa-FP anti-cancer drug prediction method based on the combination of deep learning and the chemical field, characterized in that it includes the following steps:
[0008] Step 1: Obtain data from the CI-60DTP database and data from the ChEMBL database;
[0009] Dataset collection: Use data from the NCI-60DTP database and construct a dataset based on the CDRUG dataset construction method, focusing on bioactive molecules related to anti-cancer properties;
[0010] Data integration: Merge with the NCI-60 dataset constructed by DeepCancerMap to construct a dataset covering 9 types of tumors and 60 tumor cell lines, containing 609,593 active records and 3,764,693 inactive records;
[0011] Step 2: Use ChemBERTa and BERT_base as the basic pre-trained models; BERT base pre-trains a deep bidirectional language representation model and adjusts the pre-trained BERTbase model with an additional output layer to create models for a wide range of tasks without major architectural modifications; at the same time, use ChemBERTa as the basic pre-trained model, which is based on RoBERTa as the basic model and is pre-trained on the PubChem 77M dataset by adding a specific chemical language SMILES during training;
[0012] The dimension of the model vector representation of ChemBERTa is 384, the number of heads is 12, and the number of Transformer layers is 6; the fine-tuning process of ChemBERTa-FP classification is divided into the following steps:
[0013] 1) Data Preparation: Preprocess the collected data and use the AutoTokenizer of ChemBERTa to convert the SMILES representation of chemical molecules into numerical Embeddings
[0014] 2) Load the Pre-trained Model: Load the pre-trained ChemBERTa model, which has learned the context representation of the chemical language SMILES;
[0015] 3) Define the Classification Layer: Add a classification layer on top of the ChemBERTa model and three classification layers for molecular fingerprint features: MACCS, PubChem, and Pharmacophore ErG for chemical formula classification;
[0016] 4) Set the Loss Function and Optimizer: Adopt a composite loss function that combines Focal Loss and Cross Entropy Loss, and select the Adam optimization algorithm to adjust the model parameters;
[0017] 5) Fine-tune the Model: Fine-tune the model using the prepared dataset. During this process, the parameters of the ChemBERTa-FP model will be adjusted according to the objectives of the classification task, achieved through the backpropagation algorithm and gradient descent, so that the model can better adapt to the classification task; and feed the final ChemBERTa-FP vector into the linear classification layer for a binary classification task, for anti-cancer screening and prediction of normal cell toxicity;
[0018] 6) Evaluate the Model: After fine-tuning, use the test set to evaluate the model and calculate the AUC metric to evaluate the model's performance;
[0019] 7) Save and Use the Model: After each epoch, evaluate the model performance using the validation set; when the performance on the validation set exceeds the previous best record, save the current model;
[0020] Step 4: To evaluate the performance of the developed model, model performance evaluation metrics are adopted, and the specific definitions are shown in the following equations; including: accuracy, recall, specificity, precision, and score;
[0021]
[0022] Among them, TN represents the number of true negatives, TP represents the number of true positives, FN represents the number of false negatives, and FP represents the number of false positives; ACC represents accuracy, REC represents recall, SPE represents specificity, PRE represents precision, and F1 represents the score.
[0023] The preferred technical solution is: the data in the ChEMBL database collects molecular data of 19 normal cell lines; a total of 6216 active records and 2312 inactive records; the processed data set of each of the above cell lines is randomly divided into three sub-data sets: training set, validation set and test set in a ratio of 8:1:1.
[0024] The preferred technical solution is: the NCI-60DTP database covers data sets of 9 types of tumors and 60 tumor cell lines, including 609,593 active records and 3,764,693 inactive records.
[0025] The preferred technical solution is: in order to more comprehensively evaluate the performance of the developed model and ensure its advancement and practicality in anti-cancer drug prediction; machine learning is supplemented, including: RF, SVM, KNN, NB and deep learning models; these models are all implemented using the Scikit-learn Python package; all models are trained on GPU RTX 4090 (24GB)*1, CPU12vCPU Intel(R)Xeon(R)Platinum 8352V CPU@2.10GHz, memory 90GB, hard disk 500GB.
[0026] The preferred technical solution is: in order to intuitively display the comprehensive performance of the classification model, ROC curve analysis is used; in the ROC curve analysis, the area value under the ROC curve is particularly emphasized as an important indicator for evaluating the classification ability of the model.
[0027] The preferred technical solution is: 0.5 is used as a threshold for the classification of anticancer drugs; when the probability value predicted by the model exceeds 0.5, the corresponding compound is considered to have potential anticancer effects.
[0028] Due to the application of the above technical solution, the present invention has the following advantages compared with the prior art:
[0029] 1. This invention innovatively uses two deep learning models, ChemBERTa and BERT-base, in the field of anticancer molecule screening. This is the first application of such models. We use six key performance indicators, including AUC, ACC, REC, SPE, PRE and F1 score, to comprehensively evaluate the model performance, ensuring the high efficiency and accuracy of the model in predicting anticancer molecules.
[0030] 2. Based on the ChemBERTa model, this invention first integrates three molecular fingerprint features: MACCS, PubChem, and Pharmacophore ErG to form the innovative ChemBERTa-FP model. These molecular fingerprint features enrich the data input of the model. This integration not only enhances the model's predictive ability for the structures and chemical properties of anti-cancer molecules but also significantly improves the prediction accuracy. Additionally, even under standard computer configurations (GPU RTX 4060 (8GB)*1, CPU (AMD Ryzen 9 7945HX with Radeon Graphics) 2.50GHz, memory 16GB), the model can complete the activity prediction of a single molecule in 60 different tumor cell lines within 10 seconds. This significant improvement in computational efficiency not only accelerates the screening process of anti-tumor drugs but also enhances its universality and practical value in biochemical research.
[0031] 3. A major innovation of this invention is the application of the ChemBERTa-FP model to the normal cell dataset, developing a model specifically for identifying molecules toxic to normal cells. This advancement is of great value for differentiating effective and safe anti-cancer molecules, facilitating the screening of non-toxic and safer anti-tumor drugs, thereby accelerating the drug research and development process and clinical application. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 It is a framework diagram of the ChemBERTa pre-trained model.
[0033] Figure 2 It is a framework diagram of the ChemBERTa-FP model. DETAILED DESCRIPTION OF THE INVENTION
[0034] The following specific embodiments illustrate the implementation manners of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification.
[0035] Please refer to Figure 1-2It should be noted that the structures, proportions, sizes, etc. shown in the drawings of this specification are only used to cooperate with the content disclosed in the specification for those familiar with this technology to understand and read, and are not used to limit the limiting conditions under which the present invention can be implemented. Therefore, they do not have substantial technical significance. Any modification of the structure, change in the proportional relationship, or adjustment of the size, without affecting the effects that the present invention can produce and the purposes that can be achieved, should still fall within the scope that can be covered by the technical content disclosed in the present invention. At the same time, the terms such as "upper", "lower", "left", "right", "middle", and "one" cited in this specification are only for the convenience of clear narration and are not used to limit the scope in which the present invention can be implemented. The change or adjustment of their relative relationships, without substantial change in the technical content, should also be regarded as the scope in which the present invention can be implemented.
[0036] Unless otherwise specified, the reagents or materials described in the following examples are all commercially available.
[0037] Example 1: Study on the influence of the amount of training data on the performance of the ChemBERTa-FP model
[0038] In order to further clarify the influence of different amounts on the performance of the ChemBERTa-FP model, the present invention uses the dataset updated by combining the 2016 NCI-60 database in 2023 (containing 609,593 active records and 3,764,693 inactive records), and at the same time trains the ChemBERTa-FP model with the 2016 NCI-60 dataset (containing 591,597 active records and 2,537,387 inactive records) respectively. The performance of these two models is evaluated through 6 key performance indicators such as AUC, ACC, REC, SPE, PRE, and F1 score.
[0039] The results are shown in Table 2. The performance of the ChemBERTa-FP model trained with the combined dataset (2023 + 2016) is higher than that of the 2016 NCI-60 dataset in terms of AUC, ACC, REC, SPE, and F1 score. In particular, its AUC value can reach more than 90%, even better than the current advanced anti-cancer prediction model. It also shows that the data quality of the combined dataset is relatively high. In addition, the larger the sample proportion, the better the prediction performance of the model. Based on this, subsequent anti-tumor model establishment is based on the NCI-60 (2023 + 2016) dataset.
[0040] Table 2. Performance comparison of ChemBERTa-FP on different NCI-60 datasets (%)
[0041]
[0042] Example 2: Study on the influence of 3 fingerprint features on the performance of the ChemBERTa model
[0043] Previous studies have found that combining three fingerprint features (PubChem, MACCS, and Pharmacophore ErG) with a deep learning model (GNN) and comparing it with the non-fused feature model and other traditional machine learning methods, the model after fusing features performs excellently in the overall prediction performance of drug molecular properties (such as ADMET properties). Therefore, in order to further clarify the impact of fusing three fingerprint features on the anti-cancer drug prediction performance of ChemBERTa and BET-base models, based on the NCI-60 dataset, this invention respectively fuses ChemBERTa and BET-base with three molecular fingerprint features (MACCS, PubChem, and Pharmacophore ErG), and successfully constructs ChemBERTa-FP and BERT-base-FP models.
[0044] The model performance results are shown in Table 3. Compared with ChemBERTa, ChemBERTa-FP has a slight improvement in ACC, REC, and F1 scores, especially the AUC score also increases. AUC is an important indicator because it does not depend on a specific classification threshold and can measure the overall effect of the model's classification ability for positive and negative samples. This indicates that fusing fingerprint features enhances the comprehensive performance of ChemBERT. Similarly, similar to ChemBERTa, BERT-base combined with fingerprint features (BERT-base-FP) also has an improvement in accuracy and AUC scores, although this improvement is very small. These results show that both ChemBERTa and BERT-base perform better in prediction ability after combining molecular fingerprint features. In addition, compared with BERT-base-FP, although ChemBERTa-FP has a slight decrease in ACC, PRE, and SPC, the REC, F1 scores, especially the AUC score, are significantly improved. Generally speaking, the ChemBERTa-FP model performance shows the best in the anti-cancer drug prediction ability.
[0045] Table 3. Performance comparison of ChemBERT and BERT-base before and after fusing fingerprint features (%)
[0046]
[0047] Example 3: Comparison of the prediction performance of ChemBERTa-FP with typical machine learning and deep learning (DNN)
[0048] To comprehensively evaluate the performance of ChemBERTa-FP in anti-cancer drug prediction, the present invention trains typical machine learning models (RF, SVM, KNN, NB) and deep learning model (DNN) as baseline models based on the NCI-60 dataset, and conducts a comparative analysis of the performance with the ChemBERTa-FP model.
[0049] As shown in Table 4, the ChemBERTa-FP model performs best in terms of F1 score and AUC value, reaching 65.987% and 92.429% respectively, indicating that the overall performance of this model is optimal. Although it does not perform best in terms of ACC, REC, PRE, and SPC, these values are far higher than those of most machine learning and deep learning models, further indicating that the ChemBERTa-FP model has good comprehensive prediction ability for anti-cancer drugs.
[0050] Table 4. Performance comparison of ChemBERTa-FP model and other baseline models on the NCI-60 cancer cell line dataset (%)
[0051]
[0052] Example 4: Construction of a normal cell line dataset to evaluate the model performance of ChemBERTa-FP
[0053] As is well known, anti-cancer drugs lack selective toxicity to human normal cells, resulting in serious side effects and limiting their clinical applications. Therefore, predicting the cytotoxicity of compounds to normal cell lines is crucial in the early stage of anti-cancer drug discovery. Accordingly, we established a prediction model for 19 normal cell lines, involving 6,216 active records and 2,312 inactive records. These models based on the ChemBERTa-FP method perform well on 19 normal cell lines (Table 5), and the average AUC value of the test set is 91.443%. Such a model can be used to predict whether anti-cancer drugs are toxic to normal cells. Generally speaking, the cell-based prediction model of ChemBERTa-FP can not only be used to predict active molecules against various cancer cell lines, but also to detect the potential cytotoxicity of compounds to normal cell lines, and thus can be used for cell-based anti-cancer drug discovery.
[0054] Table 5. Performance evaluation of ChemBERTa-FP model on 19 normal cell lines (%)
[0055]
[0056]
[0057] Example 5: Case study of temsirolimus
[0058] Temsirolimus is an mTOR and serine / threonine kinase inhibitor, which is approved by the FDA for renal cell carcinoma (an aggressive form of kidney cancer). It is the first clinically effective mTOR inhibitor discovered for the treatment of mantle cell lymphoma and is also approved in Europe. Current research has found that temsirolimus has an inhibitory effect on the proliferation of various tumor cells, such as lung cancer, prostate cancer, colon cancer, and breast cancer. This invention constructs a ChemBERTa-FP model based on the NCI-60 dataset to predict the anti-cancer activity of temsirolimus, and temsirolimus does not appear in the NCI-60 dataset.
[0059] The prediction results are shown in Table 6. It is predicted that temsirolimus is effective against 9 tumor cell lines. Among them, it is effective against 100% of prostate cancer cell lines, 100% of leukemia cell lines, 100% of central neuroblastoma cell lines, and 100% of non-small cell lung cancer cell lines, and the predicted probability of anti-cancer activity is relatively high, reaching more than 80%. This indicates that the prediction results of the ChemBERTa-FP model are consistent with the experimental results of current research scholars, suggesting that this model can effectively predict the anti-cancer activity of drugs. In addition, this invention further uses a ChemBERTa-FP model constructed based on 19 normal cell lines to predict the effect of temsirolimus on normal cells. The results show that temsirolimus has a toxic effect on 68.42% of normal cell lines. The possible reason for the analysis is that the mTOR pathway plays a key role in cell growth, proliferation, and survival. As an mTOR inhibitor, the main mechanism of action of temsirolimus is to block the mTOR pathway, thereby inhibiting the proliferation and survival of cancer cells. However, this pathway is not only active in cancer cells but also plays a role in normal cells. Therefore, theoretically, temsirolimus may have an impact on the growth and function of normal cells. And currently, clinical practice has confirmed that although temsirolimus has significant anti-cancer effects, its toxic and side effects are obvious. In summary, the ChemBERTa-FP model has important value in identifying anti-cancer drugs that are both effective and safe.
[0060] Table 6. Prediction results of ChemBERTa-FP based on the NCI-60 dataset for temsirolimus (%)
[0061] Tumor type Activity prediction probability Effective proportion of cell line quantity Prostate cancer 95.814 100.000 Leukemia 90.638 100.000 Central neurocytoma 82.728 100.000 Non-small cell lung cancer 80.544 100.000 Melanoma 78.284 88.890 Renal cell carcinoma 77.426 100.000 Breast cancer 76.401 100.000 Ovarian cancer 71.511 85.710 Colorectal cancer 62.772 71.430
[0062] Table 7. Prediction results of ChemBERTa-FP based on the 19 normal cell line dataset for temsirolimus (%)
[0063]
[0064]
[0065] Example 2:
[0066] An innovative anti-cancer drug prediction method, named ChemBERTa-FP, is an improved ChemBERTa model that combines deep learning technology and molecular fingerprint features, aiming to predict the activity of anti-cancer drugs to improve the efficiency and accuracy of anti-cancer drug research and development. This method specifically addresses common challenges in anti-cancer drug development, such as long R & D cycles and difficulty in analyzing complex molecular structures, and optimizes the prediction accuracy through advanced algorithms. The present invention proposes a model framework that combines deep learning and chemical knowledge, aiming to predict the activity of anti-cancer drugs. It is realized by the following technical solutions:
[0067] 1. Dataset collection and processing
[0068] Dataset source: The datasets used in the present invention mainly come from the NCI-60DTP (https: / / wiki.nci.nih.gov / display / NCIDTPdata / NCI-60+Growth+Inhibition+Data) and ChEMBL (https: / / www.ebi.ac.uk / chembl / ) databases.
[0069] Construction of anti-cancer related drug molecule datasets:
[0070] Dataset collection: Use the data of the recent version of the NCI-60DTP database (as of July 2023), and based on the dataset construction method of LI et al.: CDRUG: a web server for predicting anticancer activity of chemical compounds, focus on bioactive molecules related to anti-cancer properties. Data integration: Merge with the NCI-60 dataset constructed by Wu et al.'s DeepCancerMap: A versatile deep learning platform for target- and cell-based anticancer drug discovery; to construct a dataset covering 9 types of tumors and 60 tumor cell lines, including 609,593 active records and 3,764,693 inactive records.
[0071] Construction of normal cell related molecular datasets:
[0072] Data source: Based on the molecular data of 19 normal cell lines collected from the ChEMBL database by Wu et al.
[0073] Data volume: A total of 6,216 active records and 2,312 inactive records.
[0074] The dataset after processing each of the above cell lines was randomly divided into three sub-datasets: a training set, a validation set, and a test set, in a ratio of 8:1:1.
[0075] 2. Selection of BERT series pre-trained models
[0076] In this embodiment, two BERT models were used as the basic pre-trained models for this experiment, including ChemBERTa and BERT_base. BERT pre-trained a deep bidirectional language representation model, and an additional output layer could be used to fine-tune the pre-trained BERT model, so as to create a state-of-the-art model for a wide range of tasks without major architectural modifications. Currently, there are different versions of the BERT pre-trained model. We used BERT_base pre-trained on English Wikipedia and Book Corpus. At the same time, we also used ChemBERTa (https: / / huggingface.co / DeepChem / ChemBERTa-77M-MLM / tree / main) as the basic pre-trained model for this study ( Figure 1 ), which is based on RoBERTa as the basic model and was pre-trained on the PubChem 77M dataset (SMILES of 77 million chemical substances). By adding specific chemical language SMILES during training, ChemBERTa can better understand and process chemical information. It shares a similar architecture with BERT, but shows stronger robustness in classification tasks and provides a relatively advanced pre-trained model for small molecule drug property prediction. Therefore, by training the ChemBERTa model using the labeled cancer association dataset, improving and fine-tuning the architecture or adding more training data, and combining other advanced chemoinformatics techniques, such as molecular fingerprint features, the prediction effect of the model on anti-cancer drug properties may be further improved.
[0077] 3. Construction of ChemBERTa-FP
[0078] The dimension of the model vector representation of ChemBERTa is 384, the number of attention heads is 12, and the number of Transformer layers is 6 ( Figure 1 ). The fine-tuning process of ChemBERTa-FP classification ( Figure 2)It can be divided into the following steps: 1) Data preparation: Preprocess the data collected in this study, such as word segmentation, removing stop words, etc., and use the AutoTokenizer of ChemBERTa (DeepChem / ChemBERTa-77M-MLM) developed by Chithrananda et al. to convert the SMILES representation of chemical molecules into numerical Embeddings, which is a high-dimensional numerical encoding that can be processed by the model. This process not only captures the structural features of the molecules but also retains the richness and complexity of chemical information. Then it is input into the ChemBERTa model for deep learning training and prediction. 2) Load the pre-trained model: Load the pre-trained ChemBERTa model, which has learned the context representation of the chemical language SMILES. 3) Define the classification layer: Add a classification layer on top of the ChemBERTa model and classification layers for three molecular fingerprint features (MACCS, PubChem, and Pharmacophore ErG) for chemical formula classification. These fingerprints, as additional features, contain various structural and chemical properties of the molecules, enhancing the prediction ability of the model. 4) Set the loss function and optimizer: To effectively address the challenge of data imbalance and improve the model's ability to identify minority classes, we adopt a composite loss function that combines Focal Loss and Cross Entropy Loss. In addition, we choose the Adam optimization algorithm to adjust the model parameters. 5) Fine-tune the model: Fine-tune the model using the prepared dataset. During this process, the parameters of the ChemBERTa-FP model are adjusted according to the objectives of the classification task, achieved through the backpropagation algorithm and gradient descent, so that the model can better adapt to the classification task. And the final ChemBERTa-FP vectors are fed into the linear classification layer for a binary classification task, for anti-cancer screening and prediction of cytotoxicity to normal cells. The fine-tuning process adapts the broad chemical understanding obtained during pre-training to the specific context of anti-cancer activity, generating a specialized model dedicated to predicting the potential efficacy of molecular compounds as anti-cancer agents. At the same time, the parameter optimization conditions of the model are set as shown in Table 1. 6) Evaluate the model: After fine-tuning is completed, use the test set to evaluate the model, and evaluate the performance of the model by calculating the AUC metric. 7) Save and use the model: After each epoch, evaluate the model performance using the validation set. When the performance on the validation set exceeds the previous best record, we save the current model. In the test phase, we conduct a final evaluation of the model, focusing on its accuracy and reliability in predicting anti-cancer drug activity for subsequent use.
[0079] Table 1. Different model parameters
[0080]
[0081] 4. Baseline Models
[0082] In order to more comprehensively evaluate the performance of the models we developed and ensure their advancement and practicality in anti-cancer drug prediction, a series of classical machine learning models (RF, SVM, KNN, NB) and deep learning models (DNN) are added as baseline models for comprehensive comparison in this invention. These models are all implemented using the Scikit-learn Python package (https: / / github.com / scikit-learn / scikit-learn, version: 0.24.1). All models are trained on GPU RTX4090 (24GB)*1, CPU 12vCPU Intel(R)Xeon(R)Platinum 8352V CPU@2.10GHz, memory 90GB, and hard disk 500GB.
[0083] 5. Model Performance Evaluation
[0084] To comprehensively evaluate the performance of the models we developed, this invention adopts a series of model performance evaluation metrics, specifically defined by the following equations (1-5). These include: Accuracy (ACC), Recall (REC), Specificity (SPE), Precision (PRE), and F1 Score (F1). In addition, to visually display the comprehensive performance of the classification model, we also adopt Receiver Operating Characteristic (ROC) curve analysis. In the ROC curve analysis, we particularly emphasize the Area Under the Curve (AUC) value as an important indicator for evaluating the model's classification ability. In addition, this invention uses 0.5 as the threshold for classifying anti-cancer drugs. This means that when the probability value predicted by the model exceeds 0.5, the corresponding compound is regarded as having potential anti-cancer effects. This setting is based on in-depth analysis of the data characteristics, aiming to balance the sensitivity and specificity of the model, thus ensuring that the selected candidate compounds have high reliability.
[0085]
[0086] Where TN represents the number of true negatives, TP represents the number of true positives, FN represents the number of false negatives, and FP represents the number of false positives.
[0087] The above are only preferred embodiments for explaining the present invention and are not intended to impose any formal restrictions on the present invention. Therefore, any modifications or changes made to the present invention under the same inventive spirit should still be included within the scope intended to be protected by the present invention.
Claims
1. A ChemBERTa-FP anticancer drug prediction method based on the combination of deep learning and chemistry, characterized by: The following steps are involved: Step 1: Obtain data from the CI-60DTP database and the ChEMBL database; Dataset collection: Data from the NCI-60DTP database were used, based on the CDRUG dataset construction method, focusing on bioactive molecules associated with anticancer properties; Data integration: Merged with the NCI-60 dataset constructed by DeepCancerMap to construct a dataset covering 9 types of tumors and 60 tumor cell lines, including 609,593 active records and 3,764,693 inactive records; Step 2: ChemBERTa and BERT_base were used as base pre-trained models; BERT base pre-trained a deep bidirectional language representation model, and the pre-trained BERTbase model was adjusted with an additional output layer to create models for a wide range of tasks without major architectural modifications; at the same time, ChemBERTa was used as a base pre-trained model, which was based on RoBERTa as a base model and pre-trained on the PubChem 77M dataset, adding a specific chemical language SMILES to the training; Step 3: The dimension of the model vector representation of ChemBERTa is 384, the number of heads is 12, and the number of Transformer layers is 6; the fine-tuning process of ChemBERTa-FP classification is divided into the following steps: 1) Data preparation: preprocess the collected data and use ChemBERTa's AutoTokenizer to convert the SMILES representation of chemical molecules into numerical Embeddings; 2) Load pre-trained model: Load the pre-trained ChemBERTa model, which has learned the contextual representation of the chemical language SMILES; 3) Define the classification layer: Add a classification layer on top of the ChemBERTa model and three molecular fingerprint features MACCS, PubChem, and Pharmacophore ErG for chemical formula classification; 4) Setting the loss function and optimizer: A composite loss function of Focal Loss combined with Cross Entropy Loss was used, and the Adam optimization algorithm was used to adjust the model parameters; 5 Fine-tune the model: Use the prepared dataset to fine-tune the model. In this process, the parameters of the ChemBERTa-FP model are adjusted according to the objectives of the classification task through the back-propagation algorithm and gradient descent to make the model better suited to the classification task; and the final ChemBERTa-FP vector is fed into the linear classification layer for the binary classification task, which is used for anti-cancer screening and prediction of toxicity to normal cells; 6) Evaluate the model: After fine-tuning, use the test set to evaluate the model and calculate the AUC metric to evaluate the model performance. 7) Save and use the model: After each epoch, use the validation set to evaluate the model performance; when the performance on the validation set exceeds the previous best record, save the current model; Step 4: In order to evaluate the performance of the developed model, the model performance evaluation indicators were used, which are specifically defined in the equation shown below; including: accuracy, recall, specificity, precision, and score; Among them, TN represents the number of true negatives, TP represents the number of true positives, FN represents the number of false negatives, FP represents the number of false positives; ACC represents accuracy, REC represents recall, SPE represents specificity, PRE represents precision, and F1 represents score.
2. The ChemBERTa-FP anticancer drug prediction method based on the combination of deep learning and chemistry according to claim 1, characterized in that: The data in the ChEMBL database collected molecular data from 19 normal cell lines; a total of 6216 active records and 2312 inactive records; the processed data sets of each of the above cell lines were randomly divided into three sub-data sets: training set, validation set and test set in a ratio of 8:1:
1.
3. The ChemBERTa-FP anticancer drug prediction method based on the combination of deep learning and chemistry according to claim 1, characterized in that: The NCI-60DTP database covers datasets of 9 types of tumors and 60 tumor cell lines, including 609,593 active records and 3,764,693 inactive records.
4. The ChemBERTa-FP anticancer drug prediction method based on the combination of deep learning and chemistry according to claim 1, characterized in that: A threshold of 0.5 was used for the classification of anticancer drugs; when the probability value predicted by the model exceeded 0.5, the corresponding compound was considered to have potential anticancer effects.
Citation Information
Patent Citations
Anticancer peptide recognition method and system
CN117292742A
Method for analyzing anti-cancer activity difference of turtle back and belly armor based on metabonomics in combination with network pharmacology
CN117517536A