Molecular ADMET property prediction algorithm based on multi-modal and multi-scale characteristics

By constructing a multimodal and multi-scale feature fusion framework, using SMILES strings to extract multiple features of drug molecules and combining them with deep learning models, the accuracy and efficiency problems of drug ADMET property prediction were solved, and the success rate of drug research and development was improved.

CN120823910APending Publication Date: 2025-10-21NANJING TECH UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510919537.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-03
Publication Date
2025-10-21

Smart Images

  • Figure CN120823910A_ABST
    Figure CN120823910A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence drug research and development, in particular to a molecular ADMET property prediction algorithm based on multi-mode and multi-scale characteristics. According to the algorithm, multi-scale feature extraction is carried out on a small molecule SMILES character string by constructing a multi-modal feature fusion framework, and training and prediction are completed on 73 ADMET properties covering absorption, distribution, metabolism, excretion, toxicity and general characteristics. The invention provides a multi-modal and multi-scale feature fusion strategy, atomic features, molecular fingerprint features and physicochemical property features are integrated, and a multi-scale feature set covering molecules is formed. A progressive information extraction architecture is combined with a graph neural network (GCN), a gated loop unit (GRU) and an attention mechanism to extract features, so that the accuracy of ADMET property prediction is remarkably improved. In a word, the method marks an important step for predicting the ADMET property of the medicine.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence drug research and development technology, and in particular to a molecular ADMET property prediction algorithm based on multimodal and multiscale features. Background Art

[0002] Drug discovery is a costly, time-consuming process with a high failure rate. It takes an average of 12-15 years from new drug development to market approval, costing over $2.5 billion, and 80%-90% of drug projects are terminated in the preclinical stage. In the later stages of drug development, evaluating pharmacological properties such as ADMET properties (absorption, distribution, metabolism, excretion, and toxicity) becomes key to reducing the failure rate of drug development. Studies have shown that if the ADMET properties of drugs can be reasonably predicted in the preclinical and clinical stages, the development process can be accelerated, costs can be saved, and the success rate of drug development will be significantly improved.

[0003] Currently, the ADMET properties of small molecule drugs are primarily evaluated through in vivo and in vitro experiments. These traditional experimental methods are costly, inefficient, and require significant human and material resources. With the development of computer-aided drug design, computer simulation has become a highly effective experimental approach. While traditional experimental methods require drug synthesis prior to experimentation, computer simulations can provide a priori predictions. Traditional machine learning methods have limited accuracy in predicting ADMET properties. With the development of deep learning technology, some deep learning models have achieved some success in predicting drug ADMET properties. However, some challenges remain, such as the fact that existing ADMET property prediction schemes focus solely on drug structural characteristics and suffer from insufficient predictive performance. Summary of the Invention

[0004] The purpose of this invention is to solve the problem of single feature utilization and insufficient prediction performance in the prior art. We propose a molecular ADMET property prediction algorithm based on multimodal and multiscale features.

[0005] The concept behind this invention is to construct a multimodal, multiscale feature fusion framework to predict the ADMET properties of small molecule drugs. 1) Data preparation: collecting small molecule drug SMILES and corresponding 73 ADMET property data. 2) Drug molecules are used as model input, and the 73 properties are used as model training and prediction targets. 3) Three features are pre-calculated from the drug's SMILES string, such as extracting six atomic attributes, molecular fingerprints, and calculating the physicochemical properties of the molecule. 4) The model is constructed using multimodal feature fusion and a multi-layer network.

[0006] The specific technical solution of the present invention is: a molecular ADMET property prediction algorithm based on multimodal and multiscale features, comprising the following steps:

[0007] Obtain a dataset of small molecule drugs and 73 ADMET properties to be predicted. Divide the collected 73 ADMET properties into training, validation, and test sets in a 6:2:2 ratio. Train a model independently for each property, and ultimately combine all 73 models for prediction.

[0008] When a molecule is input into the model, it extracts three types of features. 1. It extracts atomic features from the SMILES string. 2. It extracts 166 molecular fingerprints based on the MACCS molecular fingerprint. 3. It calculates the physicochemical properties of the molecule and removes invalid physical and chemical properties to obtain 202 valid molecular descriptors. The atomic features are input into a two-layer graph neural network to extract graph-level features. The graph-level features and molecular fingerprint features are combined through a GRU to extract sequence features. The sequence features and physical and chemical property features are combined and an attention mechanism is used to extract cross-modal correlation features. Finally, a fully connected layer is used to output the prediction result. This output can be used for binary classification and regression prediction.

[0009] The present invention has the following beneficial effects:

[0010] 1. The present invention designs a multimodal feature fusion architecture that integrates three types of features: atomic-level features, molecular fingerprint features, and physical and chemical property features to comprehensively characterize molecular properties.

[0011] 2. The present invention designs a multi-scale feature extraction network that integrates GCN graph neural network (structure), GRU (sequence) and attention mechanism (physical and chemical properties) for hierarchical extraction to enhance feature utilization efficiency.

[0012] 3. The present invention uses SMILES as input to predict ADMET properties, and the model shows excellent performance in predicting 73 properties. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] Figure 1 This is the model algorithm framework diagram;

[0014] Figure 2 is the average performance graph of the model on four indicators; DETAILED DESCRIPTION

[0015] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments.

[0016] Reference Figure 1 , a molecular ADMET property prediction algorithm based on multimodal and multiscale features, comprising the following steps:

[0017] S1: The training phase used a training set of 73 ADMET properties from molecular SMILES. Each property was divided into a training set, a validation set, and a test set in a ratio of 6:2:2. These 73 ADMER properties were divided into 7 absorption, 5 distribution, 13 metabolism, 3 excretion, 36 toxicity, and 9 general properties. The specific properties are as follows: the 7 absorption properties are: Caco2 (intestinal permeability), oral bioavailability 20%, MDCK (intestinal permeability), oral bioavailability 50%, drug efflux pump inhibitor (P-gp inhibitor), drug efflux pump substrate (P-gp substrate), and skin permeability. The 5 distribution properties are: blood-brain barrier (central nervous system), blood-brain barrier (logBB permeability), human plasma free fraction, plasma protein binding rate, and apparent volume of distribution. The 13 metabolic properties are: breast cancer resistance protein inhibitor, drug metabolism (CYP1A2 inhibitor), drug metabolism (CYP1A2 substrate), drug metabolism (CYP2C19 inhibitor), drug metabolism (CYP2C19 substrate), drug metabolism (CYP2C9 inhibitor), drug metabolism (CYP2C9 substrate), drug metabolism (CYP2D6 inhibitor), drug metabolism (CYP2D6 substrate), drug metabolism (CYP3A4 inhibitor), drug metabolism (CYP3A4 substrate), hepatic uptake rate (OATP1B1), hepatic uptake rate (OATP1B3). The 3 excretion properties are: clearance, renal excretion - OCT2 transporter, and drug half-life. The 36 toxicity properties are: Ames test, bird toxicity, bee toxicity, bioconcentration factor, biodegradation, carcinogenicity, marine crustacean toxicity, hepatotoxicity (DILI), eye corrosion, eye irritation, maximum recommended daily dose, fathead minnow toxicity, hepatotoxicity (hht), cardiac arrhythmia, intestinal absorption, Daphnia magna toxicity, genotoxicity (micronucleus), aryl hydrocarbon receptor, androgen receptor, androgen receptor ligand binding domain, aromatase, estrogen receptor, estrogen receptor ligand binding domain, glucocorticoid receptor, peroxisome proliferator-activated receptor gamma, thyroid hormone receptor, Tetrahymena, acute oral toxicity in rats, chronic oral toxicity in rats, respiratory toxicity, skin sensitivity, antioxidant response related, ATAD5 related, heat shock related, matrix metalloproteinase related, and p53 related. The 9 general properties are: boiling point, hydration energy, octanol-water distribution coefficient (logD), octanol-water distribution coefficient (logP), water solubility, vapor pressure, melting point, acid dissociation constant, and base dissociation constant.

[0018] S2: The input of the model is the SMILES string of the drug.

[0019] S3: When the SMILES of small molecule drugs are input into the model, preprocessing is required to traverse the small molecule to obtain the graph matrix of atoms and bonds, extract 166 MACCS molecular fingerprint features, and calculate 202 physical and chemical property features.

[0020] S4: The network takes the three extracted features as input. The atomic features are fed into a two-layer GCN graph neural network layer to extract graph-level features. The graph-level features and molecular fingerprint features are concatenated as input to the GRU layer to obtain time series features. The time series features and 202 physical and chemical property features are concatenated as input to the attention mechanism to obtain cross-modal correlation features. Finally, a fully connected layer is used to obtain the corresponding task results.

[0021] S5: The model uses different loss functions for different tasks. The binary classification task uses weighted binary cross entropy loss (BCE loss), and the regression task uses mean square error loss (MSE loss). The formulas for the two losses are as follows:

[0022] The loss formula for the two-class task is:

[0023]

[0024] where p i Represents the model prediction value, y i represents the binary classification label, σ(x) is

[0025] The loss formula for the regression task is:

[0026]

[0027] where x i Represents the predicted value of the model, y i Represents the true value.

[0028] In the network model:

[0029] Two-layer graph neural network: Input the atomic feature matrix (N*6, N is the number of atoms in the SMILES string), and extract the graph-level features of the drug molecules through two layers of GCN (256 dimensions-128 dimensions).

[0030] GRU sequence modeling: concatenate graph-level features with molecular fingerprints, input them into a bidirectional GRU, and take the output of the last time step to obtain sequence features.

[0031] Attention mechanism: The GRU output is concatenated with 202 physical and chemical properties, and feature importance weights are learned through a multi-head attention layer to obtain cross-modal correlation features.

[0032] In this example, a total of 350,695 small molecule drugs were used. 73 prediction tasks were trained. These prediction tasks are categorized into two main types: binary classification and regression. There were 49 binary classification tasks and 24 regression tasks, respectively. Evaluation metrics were: AUC and ACC for the binary classification tasks, and R² and RMSE for the regression tasks. Figure 2 This demonstrates the good predictive performance of the model.

[0033] The above description is only a preferred specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with the technical field, within the technical scope disclosed by the present invention, who makes equivalent replacements or changes based on the technical solution and inventive concept of the present invention, should be covered by the scope of protection of the present invention.

Claims

1. A molecular ADMET property prediction algorithm based on multimodal and multiscale features, characterized by: Including steps: 1) Data screening and preparation: Obtain small molecule SMILES strings and their corresponding 73 ADMET property data. Perform normalization preprocessing on the SMILES strings. 2) Multidimensional feature extraction: Traverse the SMILES string to extract atomic features and features. 166-dimensional molecular fingerprint features were extracted based on MACCS, and 202 physicochemical properties were calculated using RDKit. 3) Using multimodal feature fusion, three types of features (atomic features, molecular fingerprint features, and physical and chemical property features) are input into the model, and feature fusion is achieved through network structure design. 4) A multi-layer network is used to extract graph-level features through a two-layer graph neural network, sequence features through GRU, cross-modal correlation features through the attention mechanism, and finally the classification task and regression task output are completed through the fully connected layer.

2. A molecular ADMET property prediction algorithm based on multimodal and multiscale features according to claim 1, characterized in that: In step 1), the preprocessing process is to use the RDKIT package to normalize the small molecule SMILES string.

3. The molecular ADMET property prediction algorithm based on multimodal and multiscale features according to claim 1, characterized in that: In step 1), 73 ADMET properties cover absorption, distribution, metabolism, excretion and general properties, specifically Caco2 (intestinal permeability), oral bioavailability 20%, MDCK (intestinal permeability), oral bioavailability 50%, drug efflux pump inhibitors (P-gp inhibitors), drug efflux pump substrates (P-gp substrates), skin permeability, blood-brain barrier (central nervous system), blood-brain barrier (logBB permeability), human plasma free fraction, plasma protein binding rate, apparent distribution volume, breast cancer Drug metabolism (CYP1A2 inhibitors), drug metabolism (CYP1A2 substrates), drug metabolism (CYP2C19 inhibitors), drug metabolism (CYP2C19 substrates), drug metabolism (CYP2C9 inhibitors), drug metabolism (CYP2C9 substrates), drug metabolism (CYP2D6 inhibitors), drug metabolism (CYP2D6 substrates), drug metabolism (CYP3A4 inhibitors), drug metabolism (CYP3A4 substrates), hepatic uptake rate (OATP1B1) , liver uptake rate (OATP1B3), clearance, renal excretion - OCT2 transporter, drug half-life, Ames test, bird toxicity, bee toxicity, bioconcentration factor, biodegradation, carcinogenicity, crustacean marine toxicity, hepatotoxicity (DILI), eye corrosion, eye irritation, maximum recommended daily dose, fathead minnow toxicity, hepatotoxicity (h_ht), arrhythmia, intestinal absorption, large Daphnia magna toxicity, genotoxicity (micronucleus), aryl hydrocarbon receptor, androgen receptor, androgen receptor ligand binding domain, aromatase , estrogen receptor, estrogen receptor ligand binding domain, glucocorticoid receptor, peroxisome proliferator-activated receptor γ, thyroid hormone receptor, Tetrahymena, acute oral toxicity in rats, chronic oral toxicity in rats, respiratory toxicity, skin sensitivity, antioxidant response related, ATAD5 related, heat shock related, matrix metalloproteinase related, p53 related, boiling point, hydration energy, octanol-water distribution coefficient (logD), octanol-water distribution coefficient (logP), water solubility, vapor pressure, melting point, acid dissociation constant, base dissociation constant.

4. The molecular ADMET property prediction algorithm based on multimodal and multiscale features according to claim 1, characterized in that: In step 2), the atomic characteristics described include atomic number, atomic charge and distribution state, number of directly connected hydrogen atoms, whether it belongs to an aromatic ring system, whether it is located in a ring structure, and number of adjacent atoms.

5. The molecular ADMET property prediction algorithm based on multimodal and multiscale features according to claim 1, characterized in that: In step 2), the extracted 166-dimensional molecular fingerprint is a MACCS molecular fingerprint extracted using the RDKIT package.

6. The molecular ADMET property prediction algorithm based on multimodal and multiscale features according to claim 1, characterized in that: In step 2), 202 physicochemical properties were calculated using the RDKIT package and obtained after removing invalid molecular descriptors.

7. The molecular ADMET property prediction algorithm based on multimodal and multiscale features according to claim 1, characterized in that: In step 4), a two-layer graph neural network is used to extract graph-level features. A GRU is used to extract molecular fingerprint sequence features. An attention mechanism is used to extract cross-modal correlation features. Reluctant Unit (ReLU) activation functions are used between each network layer. The final output layer is a fully connected layer for binary classification and regression tasks.

Citation Information

Cited By

  • Multi-feature fusion oral bioavailability prediction method, device, equipment and medium

    CN121885236A