Anti-ampullaria carcinoma active molecule prediction model based on machine learning and construction method thereof

By constructing a prediction model of anti-ampullary cancer active molecules based on machine learning, the problem of lack of effective screening models in the development of ampullary cancer drugs is solved, and the rapid and accurate screening of ampullary cancer active molecules is achieved, reducing the cost and time of drug R&D.

CN120126548APending Publication Date: 2025-06-10LANZHOU UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510175979.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-18
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

The existing technology lacks anti-cancer drug screening models for ampulla cancer, resulting in high cost and low efficiency in the research and development of ampulla cancer drugs.

Method used

Machine learning methods are used to construct an anti-ampula cancer activity molecular prediction model based on structure-activity relationships, and compound activity prediction is achieved by model training of ampulla cancer cytotoxicity data of marketed drugs.

Benefits of technology

It has achieved rapid and accurate screening of ampulla cancer active molecules, reduced the cost and time of drug development, and has important clinical application value.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120126548A_ABST
    Figure CN120126548A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of biological information, and provides an anti-ampulla cancer active molecule prediction model based on machine learning and a construction method thereof. According to the method, anti-cancer activity screening on the ampulla cancer cells is carried out on marketed drugs, compound toxicity data of the ampulla cancer cells are obtained, molecular descriptors of the compounds are calculated, and a machine learning algorithm is used for establishing an anti-ampulla cancer activity prediction model of the compounds. The method has high accuracy and is convenient to use, and the anti-ampulla cancer activity of the molecule can be judged only by providing a compound structure. The model can be used for large-scale virtual screening of ampulla cancer treatment drugs, can reduce the time and fund cost of a traditional drug screening method, and is beneficial to solving the problem that specific drugs still do not exist in current clinical treatment of ampulla cancer.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of bioinformatics technology, and particularly relates to a prediction model of anti-ampullary cancer active molecules based on machine learning and a construction method thereof. Background Art

[0002] Ampullary cancer is a rare digestive system tumor that originates from the ampulla region of the duodenum, which is the junction of the bile duct, pancreatic duct, and duodenum, accounting for approximately 0.2% of all gastrointestinal tumors. Ampullary cancer has the characteristics of low incidence and complex histopathology. Its low incidence rate makes it difficult to recruit a sufficient number of patients for corresponding clinical studies, while the complex pathological features bring certain difficulties in treatment. Ampullary cancer is usually classified into three different histological subtypes: pancreatobiliary type, intestinal type, and mixed type. Patients with different subtypes have different prognoses and are suitable for different treatment methods. Currently, the drug treatment of ampullary cancer usually adopts chemotherapy regimens based on fluorouracil and gemcitabine, and there are no targeted drugs for ampullary cancer. The traditional drug R & D process usually takes a long time and costs a lot, while the reutilization of existing drugs for new indications can reduce the cost of drug R & D to a certain extent. Machine learning is a data mining technology in the field of artificial intelligence, which can discover potential connections from a large amount of data. The quantitative structure-activity relationship (QSAR) model is a machine learning method for studying the relationship between the chemical structure and biological activity of molecules, and plays an important role in the screening and discovery of new drugs. Due to the rarity of ampullary cancer, there is currently no anti-ampullary cancer drug screening model. Therefore, to solve the problem of lack of drugs for ampullary cancer, the present invention uses machine learning methods to construct a structure-activity relationship model, and trains the model with the cytotoxicity data of ampullary cancer cells of marketed drugs, so as to achieve the prediction of active molecules. Summary of the Invention

[0003] In view of the existing technical problems, the present invention provides a prediction model of anti-ampullary cancer active molecules based on machine learning and a construction method thereof, which specifically includes the following contents:

[0004] In the first aspect, the present invention provides a prediction model of anti-ampullary cancer active molecules based on machine learning. The construction method of the prediction model of anti-ampullary cancer active molecules includes the following steps:

[0005] (1) Administer a batch of compounds to ampullary cancer cells, and obtain the cell growth inhibition rate after 48 hours of administration as the activity data; the compounds are shown in Table 1;

[0006] (2) Preprocess the activity data obtained in step (1). Using an inhibition rate of 20% at a concentration of 10 μM as the standard, classify the compound activities. Those with an inhibition rate greater than or equal to 20% are considered active, and those less than 20% are considered inactive; also, standardize the molecular structures, desalt, and remove duplicate molecules of the compound SMILES structures.

[0007] (3) Summarize the compound SMILES structures processed in step (2), and calculate to generate a total of 9 molecular fingerprints: Estatefingerprint (EstFP), Extended fingerprint (ExtFP), Fingerprint (FP), Graph onlyfingerprint (GraFP), Klekota-Roth fingerprints (KRFP), MACCS fingerprint (MACCSFP), Pubchem fingerprint (PubchemFP), Atom Pairs 2D Fingerprinter (AD2D), and Substructure fingerprint (SubFP).

[0008] (4) Divide the data processed in step (3) into a test set and a training set at a ratio of 8:2; optimize the hyperparameters of the model through grid search, and establish anti-ampullary cancer active molecule prediction models using 5 machine learning algorithms, namely logistic regression algorithm, support vector machine algorithm, random forest algorithm, K-nearest neighbor algorithm, and neural network algorithm, with the 9 molecular fingerprints of the compounds respectively.

[0009] (5) Evaluate the prediction capabilities of a total of 45 models trained by the combination of machine learning algorithms and molecular fingerprints through ten-fold cross-validation and test set validation using AUC, accuracy, balanced accuracy, recall rate, specificity, precision, and F1 score for model evaluation, select the top ten models in terms of prediction performance, and obtain an anti-ampullary cancer active molecule prediction model based on machine learning.

[0010] Table 1 430 compounds included in the data used for constructing the anti-ampullary cancer active molecule prediction model

[0011]

[0012]

[0013]

[0014]

[0015]

[0016]

[0017]

[0018]

[0019] Preferably, the ampulla cancer cells in the step (1) include DPC-X1, DPC-X3 and DPC-X4 cells.

[0020] Preferably, the molecular fingerprint calculation in the step (3) is completed by PaDEL-Descriptor software.

[0021] Preferably, the grid search in the step (4) is performed by using the GridSearchCV function in the Sklearn package.

[0022] Preferably, the activity data in the step (1) is measured by the MTT or CCK8 method.

[0023] Preferably, the model construction in the step (4) is performed by using the LogisticRegression, SVC, RandomForestClassifier, KNeighborsClassifier and MLPClassifier functions.

[0024] In a second aspect, the present invention provides an application of the anti-ampulla cancer active molecule prediction model described in the first aspect above in the prediction of anti-ampulla cancer active molecules.

[0025] In a third aspect, the present invention provides a method for predicting anti-ampulla cancer active molecules, the method comprising:

[0026] Data acquisition and processing: performing data processing on a target small molecule drug to obtain the SMILES structural formula of the target small molecule drug, and calculating and generating a total of 9 kinds of molecular fingerprint data including Estate fingerprint (EstFP), Extended fingerprint (ExtFP), Fingerprint (FP), Graph only fingerprint (GraFP), Klekota-Roth fingerprints (KRFP), MACCS fingerprint (MACCSFP), Pubchem fingerprint (PubchemFP), Atom Pairs2D Fingerprinter (AD2D) and Substructure fingerprint (SubFP);

[0027] Molecular activity prediction: Import the molecular fingerprint data of the target small molecule drug into the model described in any one of claims 1-5 to predict the anti-ampullary cancer activity of the molecule; among them, the molecules of the compounds predicted to be active by at least six of the top ten models are regarded as anti-ampullary cancer active molecules.

[0028] Fourthly, the present invention provides an anti-ampullary cancer active molecule obtained by predicting by the method described in the third aspect above.

[0029] Preferably, the active molecule is digoxin.

[0030] Fifthly, the present invention provides the application of digoxin in the preparation of anti-ampullary cancer drugs.

[0031] Sixthly, the present invention provides an electronic device, including a processor, a memory, and a computer program stored in the memory and operable on the processor; when the computer program is executed by the processor, it realizes the anti-ampullary cancer active molecule prediction model described in the first aspect above.

[0032] Seventhly, the present invention provides a computer-readable storage medium, in which computer-readable instructions are stored, and when the computer-readable instructions are executed by a processor, the anti-ampullary cancer active molecule prediction model described in the first aspect above is realized.

[0033] The beneficial effects of the present invention are as follows: The present invention constructs a prediction model of anti-ampullary cancer active molecules for the first time with a machine learning algorithm; the model can predict the activity of molecules by analyzing the chemical structure of compounds, and its operation is convenient and the operation is rapid, and it can quickly and accurately screen the activity of large-scale compounds, which is of great significance for the drug research and development of ampullary cancer. Description of the Drawings

[0034] Figure 1 Technical roadmap of a prediction model of anti-ampullary cancer active molecules based on machine learning and its construction method;

[0035] Figure 2 Schematic diagram of the combination of machine learning algorithm and molecular fingerprint;

[0036] Figure 3 Heat map of ten-fold cross-validation evaluation of the top ten models;

[0037] Figure 4 Heat map of test set validation evaluation of the top ten models;

[0038] Figure 5 Molecular structure of digoxin. Detailed Embodiments

[0039] The following examples further describe and illustrate the present invention, but do not limit the scope of protection required by the present invention. The experimental methods used in the following examples are all conventional methods unless otherwise specified.

[0040] Example 1 Construction of a machine learning-based predictive model for anti-ampullary cancer active molecules and screening of active drugs

[0041] The technical roadmap of the machine learning-based predictive model for anti-ampullary cancer active molecules and its construction method is as Figure 1 shown, and specifically includes:

[0042] 1. Data source and preprocessing

[0043] (1) Data acquisition: Resuscitate and culture ampullary cancer cell lines DPC-X1, DPC-X3 or DPC-X4 cells. Plate the cells with good growth status, and after adherence, add the medium containing the drug (drug information is shown in Table 1 above) and continue to culture for 48 h. The drug concentration in the medium is 10 μM. Detect the cell viability after 48 h of drug administration by the MTT method to obtain drug activity data. Download the SMILES structural formula of the administered compound from the Chemical Abstracts Service (CAS) database of the United States to obtain drug structure data.

[0044] (2) Data preprocessing: Preprocess the obtained data, unify the drug activity data as the cell growth inhibition rate after 48 h of drug administration, and take 20% of the inhibition rate at this concentration as the standard to classify the compound activity; take 20% of the inhibition rate at 10 μM concentration as the standard to classify the compound activity. Compounds with an inhibition rate greater than or equal to 20% are considered active, and those less than 20% are considered inactive; and perform molecular structure standardization, desalting and removal of duplicate molecules on the obtained compound SMILES structural formula.

[0045] 2. Molecular descriptor calculation

[0046] Induce the SMILES structural formula of the preprocessed compound, import it into the PaDEL-Descriptor software to calculate molecular descriptors, and calculate and generate nine molecular fingerprints including Estate fingerprint (EstFP), Extended fingerprint (ExtFP), Fingerprint (FP), Graph only fingerprint (GraFP), Klekota-Roth fingerprints (KRFP), MACCS finger print (MACCSFP), Pubchem fingerprint (PubchemFP), Atom Pairs2D Fingerprinter (AD2D) and Substructure fingerprint (SubFP).

[0047] 3. Construct a prediction model using machine learning algorithms

[0048] (1) Model construction: Divide the data into a training set and a test set with a ratio of 8:2. Optimize the hyperparameters of the model through grid search, and establish a prediction model for anti-ampullary cancer active molecules using five machine learning algorithms, namely logistic regression algorithm, support vector machine algorithm, random forest algorithm, K-nearest neighbor algorithm and neural network algorithm, with 9 molecular fingerprints of the compound, as shown in Figure 2 shown.

[0049] (2) Model evaluation: After completing the model construction, evaluate the prediction ability of a total of 45 models trained by the combination of machine learning algorithms and molecular fingerprints with AUC, accuracy, balanced accuracy, recall rate, specificity, precision and F1 score through ten-fold cross-validation and test set validation, and draw a heat map of the evaluation parameters of the top ten models. The ten-fold cross-validation evaluation heat map of the top ten models is as shown in Figure 3 shown, and the test set validation evaluation heat map of the top ten models is as shown in Figure 4 shown. The top ten models obtained in the test set validation are the prediction models for anti-ampullary cancer active molecules, including KNN_AD2D, LR_EStateFP, NN_GraphFP, RF_AD2D, RF_EStateFP, RF_ExtFP, RF_MACCSFP, SVM_FP, SVM_GraphFP and SVM_PubchemFP.

[0050] 4. Predict anti-ampullary cancer active molecules through the model

[0051] (1) Data acquisition and processing: Two conditions, Small molecule and Approved, were set in the ChEMBL compound database to obtain data on small molecule drugs approved by the FDA; data processing was performed to obtain the SMILES structural formula of the drugs and calculate the nine molecular fingerprints described in 2 above.

[0052] (2) Molecular activity prediction: The molecular fingerprint data of FDA-approved drugs were imported into the top ten models in the test set validation to predict the anti-ampullary cancer activity of the molecules. Molecules predicted as active compounds in at least six models were regarded as active molecules.

[0053] 5. Verification of the inhibitory rate of the active molecules predicted by the model by the MTT method against ampullary cancer cells

[0054] (1) When the cells grew to 80%, trypsin digestion was used, and the cells were resuspended by pipetting with RPMI-1640 medium. The cell suspension was transferred to a 96-well plate, 100 μL per well, with 15,000 cells.

[0055] (2) After culturing for 24 h for the cells to adhere, the drug was diluted with RPMI-1640 medium. The medium in the plate was discarded, and the medium containing the drug was added. After culturing for 48 h, 10 μL of 5 mg / ml MTT solution was added to each well. After incubation for 4 h, the supernatant was discarded, and the formazan crystals were dissolved with DMSO. The absorbance was measured at a wavelength of 492 nm, and the half-maximal inhibitory concentration IC 50 .

[0056] 6. Result analysis

[0057] In the present invention, the drug activity data measured by the MTT method were used as model labels, and the drug structures characterized by molecular fingerprints were used as model features. An anti-ampullary cancer active molecule prediction model was constructed through nine molecular fingerprints and five machine learning algorithms, and combined prediction was performed using the top ten models with the best performance, so as to achieve accurate and rapid molecular activity judgment.

[0058] The machine learning anti-ampullary cancer active molecule prediction model proposed in the present invention has good prediction performance. The AUC values of the top three models in the ten-fold cross-validation exceed 0.8, while the AUC values of the top three models in the test set validation are all greater than 0.7. This model operates quickly and can complete the activity screening of a large number of drugs in a short time, and the predicted drugs have good anti-ampullary cancer activity through experimental verification. The present invention can effectively assist scientific researchers in drug development for ampullary cancer and reduce the human, material and time costs required for drug screening to a certain extent.

[0059] Based on the above prediction method, the present application finds that among the top ten models, seven models, namely LR_EStateFP, NN_GraphFP, RF_AD2D, RF_EStateFP, RF_ExtFP, RF_MACCSFP, SVM_FP, SVM_GraphFP, and SVM_PubchemFP, all predict gemtuzumab as an active molecule, and the molecular structure is as Figure 5 shown. Moreover, through the MTT assay, it is found that the predicted active molecule gemtuzumab has a half-maximal inhibitory concentration IC 50 of 2.48 ± 0.34 μM against the ampulla cancer cell line DPC-X1 after 48 hours of drug administration, showing significant anti-ampulla cancer activity; this indicates that the prediction model described in the present application can effectively predict anti-ampulla cancer active molecules.

[0060] The above description is a detailed description of the preferred feasible embodiments of the present invention, but the embodiments are not intended to limit the scope of the patent application of the present invention. Any equivalent changes or modifications made under the technical spirit disclosed by the present invention should fall within the scope of the patent covered by the present invention.

Claims

1. A machine learning-based molecular prediction model for anti-ampullary carcinoma activity, characterized in that: The method for constructing the anti-ampullary carcinoma activity molecule prediction model comprises the following steps: (1) administering a compound in batches to ampullary cancer cells to obtain activity data on the cell growth inhibition rate 48 hours after administration; the compound is listed in Table 1 of the specification; (2) Preprocessing the activity data obtained in step (1) and classifying the activity of the compounds based on an inhibition rate of 20% at a concentration of 10 μM, where an inhibition rate of 20% or more is considered active and an inhibition rate of less than 20% is considered inactive; and performing molecular structure standardization, desalting, and removal of duplicate molecules on the structural formula of the compound SMILE S; (3) Summarizing the SMILES structural formulas of the compounds processed in step (2), and calculating and generating nine types of molecular fingerprints, including Estate fingerprint (EstFP), Extended fingerprint (ExtFP), Fingerprint (FP), Graph only fingerprint (Gra FP), Klekota-Roth fingerprints (KRFP), MACCS fingerprint (MACCSFP), Pubchem finger print (PubchemFP), Atom Pairs 2D Fingerprinter (AD2D) and Substructure fingerprint (Sub FP); (4) dividing the data processed in step (3) into a test set and a training set in a ratio of 8:2; optimizing the hyperparameters of the model through grid search, and establishing a molecular prediction model for anti-ampullary carcinoma activity using five machine learning algorithms, namely, logistic regression algorithm, support vector machine algorithm, random forest algorithm, K nearest neighbor algorithm, and neural network algorithm, based on the nine molecular fingerprints of the compounds; (5) A total of 45 models trained by the machine learning algorithm and molecular fingerprint combination were evaluated for their predictive ability using ten-fold cross validation and test set validation using AUC, accuracy, balanced accuracy, recall, specificity, precision, and F1 score. The top ten models with the best predictive performance were selected to obtain a molecular prediction model for anti-ampullary carcinoma activity based on machine learning.

2. The molecular prediction model for anti-ampullary carcinoma activity according to claim 1, characterized in that: The ampullary cancer cells of step (1) include DPC-X1, DPC-X3 and DPC-X4 cells.

3. The molecular prediction model for anti-ampullary carcinoma activity according to claim 1, characterized in that: The molecular fingerprint calculation in step (3) is completed by PaDEL-Descriptor software.

4. The molecular prediction model for anti-ampullary carcinoma activity according to claim 1, characterized in that: The grid search in step (4) is performed using the GridSearchCV function in the Sklearn package.

5. The molecular prediction model for anti-ampullary carcinoma activity according to claim 1, characterized in that: The activity data in step (1) is measured by MTT or CCK8 method.

6. Use of the anti-ampullary carcinoma active molecule prediction model according to any one of claims 1 to 5 in the prediction of anti-ampullary carcinoma active molecules.

7. A method for predicting active molecules against ampullary carcinoma, characterized in that: The method comprises: Data acquisition and processing: Data processing is performed on the target small molecule drugs to obtain the SMILES structural formula of the target small molecule drugs, and nine types of molecular fingerprint data are calculated and generated, including Estate fingerprint (EstFP), Extended fingerprint (ExtFP), Fingerprint (FP), Graph only fingerprint (GraFP), Klekota-Roth fingerprints (KRFP), MACCS fingerprint (MACCSFP), Pubchem fingerprint (PubchemFP), Atom Pairs2D Fingerprinter (AD2D) and Substructure fingerprint (SubFP); Molecular activity prediction: The molecular fingerprint data of the target small molecule drug is imported into the model described in any one of claims 1-5 to predict the anti-ampullary carcinoma activity of the molecule; the molecules of the compound predicted to be active by at least six of the top ten models are regarded as active molecules against ampullary carcinoma.

8. The active molecules against ampullary carcinoma predicted by the method of claim 7.

9. The active molecule against ampullary carcinoma according to claim 8, characterized in that The active molecule is digetoin.

10. Application of digetoin in the preparation of anti-ampullary cancer drugs.