ADMET property prediction method for ERα antagonists based on MMS_ResNet_1d model
Through the ADMET property prediction method of ERα antagonist based on the MMS_ResNet_1d deep learning model, the difficulty of predicting the properties of ADMET compounds in early drug development was solved, and rapid and accurate prediction was achieved, reducing the risk of failure in new drug development.
Patent Information
- Application Number
- CN202111388314.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-22
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2041-11-22
AI Technical Summary
The existing technology is difficult to predict the properties of compound ADMET in the early stages of drug research and development quickly and inexpensively, and the main reason for the failure of new drug development is the poor ADMET properties.
Using the ADMET property prediction method of ERα antagonist based on the MMS_ResNet_1d deep learning model, a multi-scale classification model was built for prediction by collecting the molecular structure descriptors and ADMET property data of the compound.
It realizes rapid and effective prediction of ADMET properties in the early stages of drug development, improves the accuracy of prediction, reduces the number of pharmacologically inappropriate compounds, and saves time and costs.
Smart Images

Figure CN114093414B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of ADMET property detection of ERα antagonists, and in particular to a method for predicting the ADMET property of ERα antagonists based on an MMS_ResNet_1d deep learning model. Background Art
[0002] In order for a compound to become a drug candidate, in addition to having good biological activity, it also needs to have good pharmacokinetic properties and safety in the human body, collectively known as ADMET (Absorption, Distribution, Metabolism, Excretion, Toxicity) properties. Among them, ADME mainly refers to the pharmacokinetic properties of the compound, which describes the law of how the concentration of the compound in the body changes over time, and T mainly refers to the toxic side effects that the compound may produce in the human body. No matter how active a compound is, if its ADMET properties are poor, such as being difficult to be absorbed by the human body, or being metabolized too quickly in the body, or having some kind of toxicity, then it is still difficult for it to become a drug, so ADMET property optimization is still needed.
[0003] The main reason for the failure of new drug candidates in the development stage is due to poor ADMET properties. The later the new drug development process fails, the greater the loss. Therefore, the earlier the factors that cause the failure of new drugs are discovered, the better. Many medicinal chemists have proposed that the ADMET properties of compounds should be evaluated in the early stages of drug development, even in the lead compound discovery stage. Selecting compounds with ideal ADMET properties for experimental screening can alleviate the economic pressure of experimental screening to a certain extent; at the same time, the toxicity and metabolic issues that were originally considered in the late stage of the drug development process can be completed in the early stage of lead compound discovery, which can effectively improve the success rate of the later development of candidate drugs.
[0004] Faced with the rapid development of combinatorial chemical proteomics and computer drug screening, conventional biological test methods (including in vitro and in vivo tests) appear to be relatively expensive and lagging behind. How to quickly and cheaply predict the ADMET properties of compounds in the early stages of drug development has become a problem that major pharmaceutical companies and research institutions are very concerned about and need to solve urgently. Summary of the invention
[0005] In order to overcome the shortcomings of the prior art, the present invention proposes a method for predicting the ADMET properties of ERα antagonists based on the MMS_ResNet_1d model, in order to more quickly and effectively perform statistical modeling and prediction of drug properties, improve the accuracy of prediction, thereby reducing the number of pharmacologically inappropriate compounds in the early stages of drug development, saving time and cost.
[0006] In order to achieve the above-mentioned object of the invention, the present invention adopts the following technical scheme:
[0007] The characteristics of the method for predicting the ADMET properties of ERα antagonists based on the MMS_ResNet_1d model of the present invention include the following steps:
[0008] S1: Collect n ADMET properties and m molecular structure descriptors of a series of antagonist compounds acting on the target ERα;
[0009] The m molecular descriptors of the target ERα antagonist compound are used as m independent variables. After data standardization operations are performed on the m independent variables, the characteristic data obtained are recorded as: X = [x1, x2, ..., x i ,…,x m ], x i represents the value of the i-th molecular descriptor; the n ADMET properties of the ERα antagonist compound are calibrated by the binary classification standard, thereby obtaining the total label of the ERα antagonist compound, and represented by the one-hot encoding, denoted as the dependent variable Y = [y1, y2, …, y j ,…,y n ], where y j is the label of the jth property, when the value is 0, it indicates a negative class, and when the value is 1, it indicates a positive class; the feature data X and the dependent variable Y are combined into a data set and divided into a training set D train and validation set D val ;
[0010] S2: Build an MMS_ResNet_1d multi-scale classification model consisting of a data input module, h branch modules and an output fusion module;
[0011] S2.1: The data input module includes a convolution layer Conv1d, a batch normalization layer BatchNorm1d, an activation function layer ReLU and a maximum pooling layer MaxPool1d in sequence. The number of channels of the input data is set to m, and the training set D train Input the data into the data input module as bs according to the size of each batch, and output the intermediate feature X′;
[0012] S2.2: Route of the ath branch module a By g residual blocks After superposition, it is connected to an adaptive pooling layer, and the b-th residual block R b It is composed of the pre-processing unit P_conv connected to the Shortcut unit through a disconnection mechanism; set the b-th residual block R b The built-in parameter is stride b , a∈[1,h];
[0013] S2.2.1: The b-th residual block R b The pre-processing unit P_conv includes a convolutional layer Conv1d b1 , a batch normalization layer BN1d, a ReLu activation function layer, a convolution layer Conv1d b2 , a post-batch normalization layer BN1d, where the convolutional layer Conv1d b1 The convolution kernel size is k ab , step length is s ab , the filling size is Convolutional layer Conv1d b2 The convolution kernel size is k ab , step size is 1, padding size is
[0014] S2.2.2: The shortcut unit of the residual block includes a convolution layer Conv1d with a convolution kernel size of 1 and a stride of 2 and a batch normalization layer BN1d;
[0015] S2.2.3: The intermediate feature X′ is input in parallel into the first residual block R1 of the h branch modules, and passes through the ath branch module Route a After the processing of the pre-processing unit P_conv and the Shortcut unit in the first residual block R1, the convolution block mapping value p_out is output b And the direct block mapping value s_out b , and the disconnection mechanism determines whether the built-in parameter stride1 of the first residual block R1 is "1". If so, the residual mapping value out1 = p_out1 + s_out1 is used as the output of the first residual block R1. Otherwise, the residual mapping value out1 = p_out1 + out0 is used as the output of the first residual block R1. When b = 1, out0 = X';
[0016] When b=2,3,…,g, the b-1th residual block R b-1 Output residual mapping value out b-1 As the b-th residual block R b The input is passed through the b-th residual block R b The processed output residual mapping value out b , so that the g-th residual block R b The output residual mapping value out g ;
[0017] S2.2.4: The last residual block R g Output residual mapping value out gAfter processing by the adaptive pooling layer, the single-scale mapping value Out is obtained a And as the ath branch module Route a The output of h branch modules is obtained.
[0018] S2.3: The output fusion module includes a fusion layer Cat, a flattening layer Flatten and a fully connected layer Fc in sequence, wherein the fusion layer Cat converts After concatenation according to the second dimension, it is processed by the flatten layer and the fully connected layer Fc, and the final output neuron mapping value is recorded as l = [l1,…l j ,…l n ], where l j Represents the mapping value output by the jth neuron in the fully connected layer;
[0019] S3: Training and selecting models:
[0020] S3.1: Initialized learning rate is lr, current iteration number is epoch, and optimal classification accuracy is ACC max , the learning rate adjustment iteration value t = 0, set the adjustment cycle threshold to t max ;
[0021] S3.2: Use formula (1) to construct the binary cross entropy loss L:
[0022]
[0023] In formula (1): σ(l j ) represents the jth neuron mapping value l j Input the sigmoid function to calculate the probability that the j-th property is predicted to be a positive class;
[0024] S3.3: In the epoch iteration, the training set D train After layer normalization processing according to the size of each batch bs, it is sent to the MMS_ResNet_1d model for training, and the gradients of m channels are solved after calculating the cross entropy loss L, and then the weight parameters in the gradient are optimized using the Adam optimizer based on the learning rate lr, so as to obtain the model of the epoch training;
[0025] S3.3: After the epoch iteration training, on the validation set D val The model trained in the epoch is verified according to each batch size of bs, and the determination coefficient ACC of the model trained in the current epoch is calculated. epoch And as an evaluation indicator, if ACC epoch >ACCmax , then ACC epoch Assign to ACC max , and save the parameters of the model trained in the current epoch. If ACC epoch ≤ACC max , then assign t+1 to t, and judge t=t max Is it true? If so, adjust the learning rate lr to 0.5lr; otherwise, keep the learning rate lr;
[0026] S3.4: After assigning epoch+1 to epoch, return to step S3.3 until the determination coefficient no longer increases, stop training and use the last trained model as the optimal classification model F;
[0027] S4: Inputting n kinds of ADMET properties to be tested into the optimal classification model F; and outputting the attributes of the corresponding properties preset by the labels.
[0028] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0029] 1. The present invention uses m molecular descriptor information of ERα antagonists as independent variables and predicts multiple ADMET properties thereof at the same time, which greatly saves the required research and development time compared with single experimental detection, and can reduce the number of pharmacologically inappropriate compounds in the early stage of drug development;
[0030] 2. The present invention uses an improved deep learning algorithm to systematically model and predict the ADMET properties of ERα antagonist compounds to assist in drug design and screening. Compared with other methods, the method of the present invention can perform statistical modeling and prediction more quickly and effectively, and improves the accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] Figure 1 It is a flow chart of the method for predicting the ADMET properties of ERα antagonists based on the MMS_ResNet_1d deep learning model of the present invention;
[0032] Figure 2 It is a schematic diagram of the MMS_ResNet_1d deep learning network architecture of the present invention;
[0033] Figure 3 It is a schematic diagram of the X data structure of the present invention. DETAILED DESCRIPTION
[0034] In this embodiment, in order to save time and improve the accuracy of ADMET property detection of ERα antagonists, a method for predicting ADMET properties of ERα antagonists based on the MMS_ResNet_1d deep learning model is designed. According to the ADMET property data of known molecules, a computational prediction model is established using the MMS_ResNet_1d deep learning network to predict various ADMET properties of unknown molecules, and effectively apply it to the ADMET property pre-detection task of ERα antagonists. Specifically, Figure 1 As shown, the steps are as follows:
[0035] S1: Collect n ADMET properties and m molecular structure descriptors of a series of antagonist compounds acting on the target ERα;
[0036] The m molecular descriptors of the target ERα antagonist compound are used as m independent variables. After data standardization operations are performed on the m independent variables, the characteristic data obtained are recorded as: X = [x1, x2, ..., x i ,…,x m ], x i represents the value of the i-th molecular descriptor; the n ADMET properties of the ERα antagonist compound are calibrated by the binary classification standard, thereby obtaining the total label of the ERα antagonist compound, and represented by the one-hot encoding, denoted as the dependent variable Y = [y1, y2, …, y j ,…,y n ], where y j is the label of the jth property, when the value is 0, it indicates a negative class, and when the value is 1, it indicates a positive class; the feature data X and the dependent variable Y are combined into a data set and divided into a training set D train and validation set D val ;
[0037] In this embodiment, m=729, n=5, that is, 729 molecular descriptors of the ERα antagonist compound are used as X values, and z-score standardization operation is performed, and five ADMET properties of the ERα antagonist compound are used as Y values, respectively, and the meanings of the five properties are: small intestinal epithelial cell permeability (Caco-2), cytochrome P450 enzyme (Cytochrome P450, CYP) 3A4 subtype (CYP3A4), compound cardiac safety evaluation (human Ether-a-go-go Related Gene, hERG), human oral bioavailability (Human Oral Bioavailability, HOB), micronucleus test (Micronucleus, MN); the training set and the validation set are divided into a ratio of 9:1;
[0038] S2: Build an MMS_ResNet_1d multi-scale classification model consisting of a data input module, h branch modules and an output fusion module; in this embodiment, h=3, and the model architecture is as follows Figure 2 As shown;
[0039] S2.1: The data input module includes a convolution layer Conv1d, a batch normalization layer BatchNorm1d, an activation function layer ReLU and a maximum pooling layer MaxPool1d. The number of channels of the input data is set to m, and the training set D train According to the size of each batch, bs is input into the data input module, and the intermediate feature X′ is output; in this embodiment, m=729;
[0040] S2.2: Route of the ath branch module a By g residual blocks After superposition, it is connected to an adaptive pooling layer, and the b-th residual block R b It is composed of the pre-processing unit P_conv connected to the Shortcut unit through a disconnection mechanism; set the b-th residual block R b The built-in parameter is stride b , a∈[1,h]; in this embodiment, g=3;
[0041] S2.2.1: The bth residual block R b The pre-processing unit P_conv includes a convolutional layer Conv1d b1 , a batch normalization layer BN1d, a ReLu activation function layer, a convolution layer Conv1d b2 , a post-batch normalization layer BN1d, where the convolutional layer Conv1d b1 The convolution kernel size is k ab , step length is s ab , the filling size is Convolutional layer Conv1d b2 The convolution kernel size is k ab , step size is 1, padding size is In this embodiment, the convolution kernel sizes in the three branch modules are k 1b =3, k 2b =5, k 3b =7, step length s a1 =1,s a2 =2, s a3 =2;
[0042] S2.2.2: The shortcut unit of the residual block contains a convolution layer Conv1d with a convolution kernel size of 1 and a stride of 2 and a batch normalization layer BN1d;
[0043] S2.2.3: The intermediate feature X′ is input in parallel into the first residual block R1 of the h branch modules, and passes through the ath branch module Route a After the processing of the pre-processing unit P_conv and the Shortcut unit in the first residual block R1, the convolution block mapping value p_out is output b And the direct block mapping value s_out b , and the disconnection mechanism determines whether the built-in parameter stride1 of the first residual block R1 is "1". If so, the residual mapping value out1 = p_out1 + s_out1 is used as the output of the first residual block R1. Otherwise, the residual mapping value out1 = p_out1 + out0 is used as the output of the first residual block R1. When b = 1, out0 = X';
[0044] When b=2,3,…,g, the b-1th residual block R b-1 Output residual mapping value out b-1 As the b-th residual block R b The input is passed through the b-th residual block R b The processed output residual mapping value out b , so that the g-th residual block R b The output residual mapping value out g ; In this embodiment, S1 = 1, S2 = 2, S3 = 2;
[0045] S2.2.4: The last residual block R g Output g After being processed by the adaptive pooling layer, it is used as the branch module Route shown in the figure. a The output is denoted as Out a ;
[0046] S2.2.4: The last residual block R g Output residual mapping value out g After processing by the adaptive pooling layer, the single-scale mapping value Out is obtained a And as the ath branch module Route a The output of h branch modules is obtained.
[0047] S2.3: The output fusion module includes a fusion layer Cat, a flattening layer Flatten and a fully connected layer Fc. The fusion layer Cat converts After concatenation according to the second dimension, it is processed by the flatten layer and the fully connected layer Fc, and the final output neuron mapping value is recorded as l = [l1,…l j ,…ln ], where l j Represents the mapping value output by the jth neuron in the fully connected layer;
[0048] S3: Training and selecting models:
[0049] S3.1: Initialized learning rate is lr, current iteration number is epoch, and optimal classification accuracy is ACC max , the learning rate adjustment iteration value t = 0, set the adjustment cycle threshold to t max ; In this embodiment, the initial learning rate lr = 0.01, epoch = 1, ACC max =0,t max =2, that is, when the cumulative accuracy does not improve twice, the learning rate is reduced for adjustment, and t is reset to 0;
[0050] S3.2: Use formula (1) to construct the binary cross entropy loss L:
[0051]
[0052] In formula (1): σ(l j ) represents the jth neuron mapping value l j Input the sigmoid function to calculate the probability that the j-th property is predicted to be a positive class;
[0053] In this embodiment, in order to speed up the convergence of the model, the present invention designs to perform layer normalization on the X data before inputting the model, which is implemented using the torch.nn.LayerNorm command in pytorch; the input data structure is as follows Figure 3 As shown, the size is [32,729,1]; in multi-label classification tasks, the problem can be converted into multiple binary classification problems, the loss of each label of a sample is calculated, and then the average is taken to get the final loss. Sigmoid is used as the activation function of the output layer. At this time, the output of the last layer cannot be regarded as a distribution because the sum is not 1. Now each neuron in the output layer is regarded as a binomial distribution, which is equivalent to converting a multi-label problem into a binary classification problem on each label. Use binary_crossentropy (binary cross entropy loss function) as the loss function. The cross entropy is expressed as the difference between the true probability distribution and the predicted probability distribution. The smaller the value of the cross entropy, the better the model prediction effect.
[0054] S3.3: After the epoch iteration training, on the validation set D val The model trained in the epoch is verified according to each batch size of bs, and the determination coefficient ACC of the model trained in the current epoch is calculated. epochAnd as an evaluation indicator, if ACC epoch >ACC max , then ACC epoch Assign to ACC max , and save the parameters of the model trained in the current epoch. If ACC epoch ≤ACC max , then assign t+1 to t, and judge t=t max Is it true? If so, adjust the learning rate lr to 0.5lr; otherwise, keep the learning rate lr;
[0055] S3.4: After assigning epoch+1 to epoch, return to step S3.3 until the determination coefficient no longer increases, stop training and use the last trained model as the optimal classification model F;
[0056] S4: Input n kinds of ADMET properties to be tested into the optimal classification model F; and output the attributes of the corresponding properties preset by the label. In this embodiment, the corresponding property attributes preset by each property classification label are: 1) Caco-2: '1' represents that the compound has good intestinal epithelial cell permeability, and '0' represents that the compound has poor intestinal epithelial cell permeability; 2) CYP3A4: '1' represents that the compound can be metabolized by CYP3A4, and '0' represents that the compound cannot be metabolized by CYP3A4; 3) hERG: '1' represents that the compound has cardiotoxicity, and '0' represents that the compound does not have cardiotoxicity; 4) HOB: '0' represents that the compound has good oral bioavailability, and '0' represents that the compound has poor oral bioavailability; 5) MN: '1' represents that the compound has genotoxicity, and '0' represents that the compound does not have genotoxicity. In order to further illustrate the advantages of this method, other deep learning models were also built for training comparison, and the results are shown in Table 1, where ACC represents the accuracy rate, which is a percentage.
[0057] Table 1 Training results of five data validation sets corresponding to different models
[0058]
[0059] It can be seen from Table 1 that the improved MMS_Resnet_1d deep learning model designed by the present invention has an increased accuracy compared to the original model, and compared to other models, it can achieve good accuracy in the classification prediction tasks of five properties.
Claims
1. A method for predicting the ADMET properties of ERα antagonists based on the MMS_ResNet_1d model, characterized in that: The following steps are involved: S1: Collect n ADMET properties and m molecular structure descriptors of a series of antagonist compounds acting on the target ERα; The m molecular descriptors of the target ERα antagonist compound are used as m independent variables. After data standardization operations are performed on the m independent variables, the characteristic data obtained are recorded as: , represents the value of the i-th molecular descriptor; the n ADMET properties of the ERα antagonist compounds are calibrated by the binary classification standard, thereby obtaining the total label of the ERα antagonist compound, which is represented by a unique hot encoding and recorded as the dependent variable ,in, is the label of the jth property, when the value is 0, it indicates a negative class, and when the value is 1, it indicates a positive class; the feature data X and the dependent variable Y are combined into a data set and divided into a training set D train and validation set D val ; S2: Build an MMS_ResNet_1d multi-scale classification model consisting of a data input module, h branch modules and an output fusion module; S2.1: The data input module includes a convolution layer Conv1d, a batch normalization layer BatchNorm1d, an activation function layer ReLU and a maximum pooling layer MaxPool1d in sequence. The number of channels of the input data is set to m, and the training set D train According to the size of each batch, bs is input into the data input module and the intermediate features are output ; S2.2: The ath branch module Route a By g residual blocks After superposition, it is connected to an adaptive pooling layer, and the bth residual block It is composed of the pre-processing unit P_conv connected to the Shortcut unit through a disconnection mechanism; set the bth residual block The built-in parameter is stride b , ; S2.2.1: The b-th residual block R b The pre-processing unit P_conv includes a convolutional layer Conv1d b1 , a batch normalization layer BN1d, a ReLu activation function layer, a convolution layer Conv1d b2 , a post-batch normalization layer BN1d, where the convolutional layer Conv1d b1 The convolution kernel size is , the step length is , the filling size is , convolutional layer Conv1d b2 The convolution kernel size is , step size is 1, padding size is ; S2.2.2: The shortcut unit of the residual block includes a convolution layer Conv1d with a convolution kernel size of 1 and a stride of 2 and a batch normalization layer BN1d; S2.2.3: The intermediate characteristics Input the first residual block of h branch modules in parallel In the a-th branch module Route a The first residual block After processing by the pre-processing unit P_conv and the Shortcut unit, the convolution block mapping value is output and direct block mapping value , and the disconnection mechanism determines the first residual block The built-in parameter of stride1 is whether it is "1". If so, the residual mapping value As the first residual block The output of , otherwise, the residual mapping value As the first residual block The output of, when b=1, ; When b=2,3,…,g, the b-1th residual block Output residual map value As the b-th residual block The input is passed through the bth residual block. The processed output residual map value , so that the g-th residual block The output residual map value of ; S2.2.4: The last residual block R g Output residual map value After processing by the adaptive pooling layer, the single-scale mapping value is obtained And as the a-th branch module Route a The output of h branch modules is obtained. ; S2.3: The output fusion module includes a fusion layer Cat, a flattening layer Flatten and a fully connected layer Fc in sequence, wherein the fusion layer Cat converts After concatenation according to the second dimension, it is processed by the flatten layer Flatten and the fully connected layer Fc, and the final output neuron mapping value is recorded as ,in, Represents the mapping value output by the jth neuron in the fully connected layer; S3: Training and selecting models: S3.1: Initialized learning rate is lr, current iteration number is epoch, and optimal classification accuracy is ACC max , the learning rate adjustment iteration value t=0, set the adjustment cycle threshold to t max ; S3.2: Use formula (1) to construct the binary cross entropy loss L: (1) In formula (1): Represents mapping the jth neuron to a value Input the sigmoid function to calculate the probability that the j-th property is predicted to be a positive class; S3.3: In the epoch iteration, the training set D train After layer normalization processing according to the size of each batch bs, it is sent to the MMS_ResNet_1d model for training, and the gradients of m channels are solved after calculating the cross entropy loss L, and then the weight parameters in the gradient are optimized using the Adam optimizer based on the learning rate lr, so as to obtain the model of the epoch training; S3.4: After the epoch of iterative training, on the validation set D val The model trained in the epoch is verified with a batch size of bs, and the determination coefficient of the model trained in the current epoch is calculated. And as an evaluation indicator, if >ACC max , then Assign to ACC max , and save the parameters of the model trained in the current epoch. If ≤ACC max , then assign t+1 to t, and judge t=t max Is it true? If so, adjust the learning rate lr to 0.5lr; otherwise, keep the learning rate lr; S3.5: After assigning epoch+1 to epoch, return to step S3.3 until the determination coefficient no longer increases, stop training and use the last trained model as the optimal classification model F; S4: Inputting n kinds of ADMET properties to be tested into the optimal classification model F, and outputting the attributes of the corresponding properties preset by the labels.
Citation Information
Patent Citations
Organic ER affinity quick screening and forecast method based on receptor binding mode
CN101059520A
Medicament module pharmacokinetic property and toxicity predicting method based on capsule network
CN109979541A