Compound ADMET property prediction method based on combined molecular characterization and electronic equipment

By converting the molecular map, SMILES sequence and molecular fingerprint of the compound into joint molecular characterization, and using deep learning algorithms to train the ADMET property prediction model, the problem of insufficient molecular characterization during training of the ADMET prediction model in the prior art is solved, and the accuracy and stability of the prediction results are improved.

CN120072102APending Publication Date: 2025-05-30HANGZHOU BIO SINCERITY PHARMA TECH CO LTD

Patent Information

Application Number
CN202510049700.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-13
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The drug molecule characterization used in the prior art when training drug ADMET prediction model is not comprehensive enough, resulting in poor accuracy of ADMET prediction results.

Method used

The compound ADMET property prediction method based on joint molecular characterization is used. By obtaining the molecular graph and SMILES sequence of the compound to be predicted, it is processed into graph feature vectors, sequence global feature vectors and molecular fingerprint vectors, and these vectors are spliced ​​into joint molecular characterization, and the ADMET property prediction model is obtained by inputting deep learning algorithms to train.

Benefits of technology

By integrating the characteristics of molecular maps, SMILES sequences and molecular fingerprints, the feature loss is reduced and the accuracy and stability of ADMET properties prediction are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120072102A_ABST
    Figure CN120072102A_ABST
Patent Text Reader

Abstract

The invention provides a combined molecular characterization-based compound ADMET property prediction method and electronic equipment. The method comprises the following steps: acquiring a molecular map and an SMILES sequence of a to-be-predicted compound molecule; processing the molecular image into an image feature vector, and processing the SMILES sequence into a sequence global feature vector and a molecular fingerprint vector; splicing the graph feature vector, the sequence global feature vector and the molecular fingerprint vector to obtain a joint molecular representation; inputting the combined molecular characterization into an ADMET property prediction model, and outputting the ADMET property corresponding to the predicted compound molecule to be predicted by the ADMET property prediction model; wherein the ADMET property prediction model is obtained by training a deep learning algorithm. According to the method, the three different types of molecular features of the to-be-predicted compound molecules are fused, so that feature deletion is effectively reduced, and the accuracy and stability of ADMET property prediction results are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of artificial intelligence drug R & D, and specifically relates to a method for predicting ADMET properties of compounds based on combined molecular characterization and an electronic device. Background Art

[0002] In the process of drug discovery and development, ADMET (absorption, distribution, metabolism, excretion, and toxicity of drugs) properties are crucial for the success of new drugs. The absorption and distribution of drugs affect their bioavailability in the body, while metabolism and excretion determine the duration and safety of drugs in the body. Toxicity assessment is an important part of ensuring the safety of new drugs. Therefore, effective ADMET property prediction can assist in screening potential compounds, saving time and resources. Through accurate ADMET prediction, researchers can increase the success rate of candidate drugs and reduce the risk of failure.

[0003] In recent years, with the rapid development of deep learning technology, ADMET prediction models can be constructed by using deep neural networks. These models can be trained using large-scale public datasets to capture the potential features of compounds. For example, the existing patent document CN 118522372 A discloses a method and model for predicting molecular ADMET properties based on fused fingerprints. The property information in small molecules is extracted, the semantic information in small molecules is extracted and fused to obtain a molecular representation through a three-part architecture, and then the ECFP (extended connectivity fingerprint) is fused to enrich the features. Then, a molecular ADMET property prediction model is trained through the above features. In the above literature, the SMILES (Simplified Molecular Input Line Entry Specification) string and InChi (International Chemical Identifier) string of the molecule are used for characterization. The above string representation methods cannot fully capture the three-dimensional structure and spatial relationship of the molecule, which may lead to the loss of important information, and molecules with similar structures but different string representations may be misclassified. Although the above solution fuses ECFP, it is only a binary vector generated by feature encoding the molecular structure, which can capture the chemical properties of the molecule to a certain extent, cannot fully capture the complex structure information of the molecule, and cannot adapt to newly emerging chemical properties. Therefore, the drug molecular characterization used in the prior art for training the ADMET prediction model may lose important features of the molecule, thus affecting the model training result and further affecting the accuracy of ADMET prediction. Summary of the Invention

[0004] The technical problem to be solved by this application is that the drug molecular representation used in the training of the drug ADMET prediction model in the prior art is not comprehensive enough, and the results obtained by the ADMET prediction model trained thereby for predicting the ADMET of drugs are poor in accuracy. Furthermore, a method for predicting the ADMET properties of compounds based on joint molecular representation and an electronic device are provided.

[0005] In a first aspect, the technical solution of this application provides a method for predicting the ADMET properties of compounds based on joint molecular representation, including:

[0006] Obtain the molecular graph and SMILES sequence of the compound molecule to be predicted;

[0007] Process the molecular graph into a graph feature vector, and process the SMILES sequence into a sequence global feature vector and a molecular fingerprint vector;

[0008] Concatenate the graph feature vector, the sequence global feature vector, and the molecular fingerprint vector to obtain a joint molecular representation;

[0009] Input the joint molecular representation into the ADMET property prediction model, and the ADMET property prediction model outputs the predicted ADMET properties corresponding to the compound molecule to be predicted; wherein, the ADMET property prediction model is trained by a deep learning algorithm.

[0010] In some embodiments of the method for predicting the ADMET properties of compounds based on joint molecular representation, the process of processing the molecular graph into a graph feature vector includes:

[0011] Perform graph feature encoding on the molecular graph to obtain an initial feature vector; the graph feature encoding includes: taking the atoms in the molecular graph as point features and the chemical bonds in the molecular graph as edge features;

[0012] Input the initial feature vector into a graph convolutional neural network model, and perform transfer, update, and iterative operations on the point features and the edge features respectively until the number of iterations reaches the set number of iterations;

[0013] Recombine the updated point features and updated edge features output by the graph convolutional neural network model to obtain the graph feature vector.

[0014] In some embodiments of the method for predicting the ADMET properties of compounds based on joint molecular representation, the process of processing the SMILES sequence into a sequence global feature vector includes:

[0015] Input the SMILES sequence into a language learning model, and perform word segmentation, encoding, vector mapping, and iterative operations on the SMILES sequence until the number of iterations reaches the set number of iterations;

[0016] Use the output vector obtained by the vector mapping described in the last iteration process as the sequence global feature vector.

[0017] In some solutions, the compound ADMET property prediction method based on joint molecular characterization processes the SMILES sequence into a molecular fingerprint vector, including:

[0018] Input the SMILES sequence into the RDKit tool to extract the molecular fingerprint corresponding to the SMILES sequence; among them, the RDKit tool is an open-source toolkit for chemoinformatics. Based on two-dimensional and three-dimensional molecular operations of compounds, it can use machine learning methods to generate compound descriptors, generate molecular fingerprints, calculate compound structure similarities, and display two-dimensional and three-dimensional molecules, etc.

[0019] Perform normalization processing on the molecular fingerprint to obtain the molecular fingerprint vector.

[0020] In some solutions, the compound ADMET property prediction method based on joint molecular characterization, after splicing the graph feature vector, the sequence global feature vector, and the molecular fingerprint vector to obtain a joint molecular characterization, includes:

[0021] Use the vector obtained by arranging and connecting the graph feature vector, the sequence global feature vector, and the molecular fingerprint vector corresponding to the same compound molecule to be predicted in sequence as the joint molecular characterization of the same compound molecule to be predicted.

[0022] In some solutions, the compound ADMET property prediction method based on joint molecular characterization, the training of the ADMET property prediction model includes:

[0023] Obtain a compound dataset, and after standardizing and information annotating the compounds in the compound dataset, obtain a sample compound dataset. The information annotation includes compound attributes and the attribute values of the compound attributes;

[0024] Obtain the joint molecular characterization and the corresponding ADMET properties of each sample compound in the sample compound dataset; the joint molecular characterization includes the graph feature vector of the sample compound molecule, the sequence global feature vector of the sample compound molecule, and the molecular fingerprint vector of the sample compound molecule;

[0025] Use the joint molecular characterization and ADMET properties of the compound samples to train a deep learning algorithm to obtain an ADMET property prediction model; the deep learning algorithm includes a feedforward neural network.

[0026] The compound ADMET property prediction method based on joint molecular characterization described in some solutions, which uses the joint molecular characterization and ADMET properties of the compound samples to train a deep learning algorithm to obtain an ADMET property prediction model, includes:

[0027] Determine the task type of the sample compound according to the information annotation of the compound sample;

[0028] Determine the loss function and evaluation index in the training process of the feedforward neural network according to the task type of the sample compound.

[0029] In a second aspect, the technical solution of the present application provides a computer-readable storage medium, in which program information is stored. After the computer reads the program information, it executes the steps of the compound ADMET property prediction method based on joint molecular characterization described in any one of the technical solutions in the first aspect.

[0030] In a third aspect, the technical solution of the present application provides a computer program product, including computer programs / instructions, and when the computer programs / instructions are executed by a processor, they implement the steps of the compound ADMET property prediction method based on joint molecular characterization described in any one of the technical solutions in the first aspect.

[0031] In a fourth aspect, the technical solution of the present application provides an electronic device, including a memory, a processor, and a computer program stored on the memory. The processor executes the computer program to implement the steps of the compound ADMET property prediction method based on joint molecular characterization described in any one of the technical solutions in the first aspect.

[0032] The above technical solutions provided by the present application, compared with the prior art, have the following technical effects:

[0033] The compound ADMET property prediction method and electronic device based on joint molecular characterization provided by this application obtain the molecular graph and SMILES sequence of the compound to be predicted, expand the molecular graph and SMILES sequence into graph feature vectors, sequence global feature vectors, and molecular fingerprint vectors, and then splice the three vectors to obtain a joint molecular characterization. The ADMET property prediction model is obtained in advance using the joint molecular characterization for a deep learning algorithm. Therefore, by inputting the joint molecular characterization of the compound molecule to be predicted into the ADMET property prediction model, the ADMET properties corresponding to the compound molecule to be predicted can be obtained. Through the above solution of this application, the prediction of ADMET properties is realized using the graph feature vectors corresponding to the molecular graph, the sequence global feature vectors obtained from the SMILES sequence, and the molecular fingerprint vectors. Through the graph feature vectors, the relationships and features between nodes in the molecular graph can be fully captured. Through the sequence global feature vectors, various relationships between different parts of the SMILES sequence are concerned, avoiding the loss of important information when processing long sequences, and thus capturing the rich context information of the SMILES sequence. Through the molecular fingerprint vectors, the features of the molecule can be covered from simple physical and chemical properties to complex molecular descriptors, making up for the structural information that is difficult to learn in the first two types of vector features. That is, this solution effectively reduces feature loss and improves the accuracy and stability of the ADMET property prediction results through the fusion of three different types of molecular features of the compound molecule to be predicted. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] Figure 1 is a flowchart of the compound ADMET property prediction method based on joint molecular characterization according to an embodiment of this application;

[0035] Figure 2 is a schematic block diagram of the principle of the graph convolutional neural network model used to obtain graph feature vectors from the molecular graph according to an embodiment of this application;

[0036] Figure 3 is a schematic block diagram of the principle of the language learning model used to obtain global feature vectors from the SMILES sequence according to an embodiment of this application;

[0037] Figure 4 is a schematic diagram of the ADMET property prediction model training process according to an embodiment of this application;

[0038] Figure 5 is a schematic diagram of the hardware connection relationship of the electronic device used to execute the compound ADMET property prediction method based on joint molecular characterization according to an embodiment of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0039] The following further describes the specific embodiments of this application with reference to the accompanying drawings.

[0040] It is easy to understand that, according to the technical solution of the present application, under the premise of not changing the essence of the present application, various structural forms and implementation methods that can be mutually replaced by those of ordinary skill in the art. Therefore, the following specific embodiments and the accompanying drawings are only exemplary illustrations of the technical solution of the present application, and should not be regarded as all of the present application or as a limitation or restriction on the technical solution of the application.

[0041] The orientation terms such as up, down, left, right, front, back, front side, back side, top, bottom, etc. mentioned or possibly mentioned in this specification are defined relative to the structures shown in the respective drawings. They are relative concepts, and thus may change accordingly depending on their different positions and different usage states. Therefore, these or other orientation terms should not be construed as restrictive terms.

[0042] This embodiment provides a method for predicting the ADMET properties of compounds based on joint molecular representation, which is applied to a computer system for predicting the ADMET properties of drugs, such as Figure 1 as shown, the method includes the following steps:

[0043] S100: Obtain the molecular graph and SMILES sequence of the compound molecule to be predicted.

[0044] When it is determined to predict the ADMET properties of a certain compound molecule, the structure of the compound molecule is determined. Define this compound molecule as the compound molecule to be predicted, and obtain the molecular graph and SMILES sequence according to its determined molecular structure, where the molecular graph can also be obtained by converting the SMILES sequence through the RDKit tool.

[0045] S200: Process the molecular graph into a graph feature vector, and process the SMILES sequence into a sequence global feature vector and a molecular fingerprint vector.

[0046] In this step, the molecular graph can be processed into a graph feature vector by means of graph encoding; use a model for processing language to map the processing result into a vector after processing the SMILES sequence according to language processing methods such as word segmentation, keyword expansion, semantic expansion, etc. to obtain the sequence global feature vector. The molecular fingerprint vector can be obtained by directly processing the SMILES sequence through the RDKit tool.

[0047] S300: Concatenate the graph feature vector, the sequence global feature vector, and the molecular fingerprint vector to obtain a joint molecular representation.

[0048] The specific concatenation method can directly combine different vectors.

[0049] S400: Input the combined molecular representation into an ADMET property prediction model, and the ADMET property prediction model outputs the predicted ADMET properties corresponding to the compound molecule to be predicted; wherein, the ADMET property prediction model is obtained by training with a deep learning algorithm.

[0050] The deep learning algorithm can select existing deep neural networks in the prior art, such as deep convolutional networks, recurrent neural networks, deep residual networks, etc. Use existing open-source data to obtain training samples in advance, and use the training samples to train the selected deep learning algorithm. After the training is completed, the parameters in the deep learning algorithm are solidified to obtain the ADMET property prediction model. The ADMET property prediction model can be directly used to predict the ADMET properties of the compound molecule to be predicted. It can be understood that the data form of the input of the training samples for training the deep learning algorithm is consistent with the combined molecular representation obtained by splicing the graph feature vector, the sequence global feature vector, and the molecular fingerprint vector.

[0051] In the solution of the above embodiment, by obtaining the molecular graph and SMILES sequence of the compound molecule to be predicted, the molecular graph and SMILES sequence are expanded into a graph feature vector, a sequence global feature vector, and a molecular fingerprint vector, and then the three vectors are spliced to obtain a combined molecular representation. Use the combined molecular representation to obtain the ADMET property prediction model by a deep learning algorithm in advance. Therefore, inputting the combined molecular representation of the compound molecule to be predicted into the ADMET property prediction model can obtain the ADMET properties corresponding to the compound molecule to be predicted. Through the above solution of the present application, the prediction of ADMET properties is realized by using the graph feature vector corresponding to the molecular graph, the sequence global feature vector obtained from the SMILES sequence, and the molecular fingerprint vector. The relationship and features between nodes in the molecular graph can be fully captured through the graph feature vector. The various relationships between different parts in the SMILES sequence are concerned through the sequence global feature vector, avoiding the loss of important information when processing long sequences, and then capturing the rich context information of the SMILES sequence. The molecular fingerprint vector can cover the features of the molecule from simple physicochemical properties to complex molecular descriptors, making up for the structural information that is difficult to learn in the features of the first two vectors. That is, this solution effectively reduces feature loss and improves the accuracy and stability of the ADMET property prediction result by fusing three different types of molecular features of the compound molecule to be predicted.

[0052] In the above solution, in step S200, the processing of the molecular graph into a graph feature vector includes:

[0053] S201: Perform graph feature encoding on the molecular graph to obtain an initial feature vector; the graph feature encoding includes: taking the atoms in the molecular graph as point features and the chemical bonds in the molecular graph as edge features.

[0054] Specifically, after taking the atoms in the molecular graph as point features and the chemical bonds in the molecular graph as edge features, the adjacency matrix corresponding to the molecular graph is obtained. The initial feature vector of the molecular graph is constructed according to the existing information annotation and topological structure of each atom and chemical bond in the molecular graph. The initial feature vector of each point includes the atomic number, the number of bonds connected to each atom, the formal charge, chirality, the number of hydrogens of the atom, hybridization and aromaticity, and the one-hot encoding of the atomic mass (divided by 100), etc. The initial feature vector of each edge contains the chemical bond type (whether the chemical bond is a conjugated bond or a cyclic bond), and whether the chemical bond contains stereochemical information.

[0055] S202: Input the initial feature vector into the graph convolutional neural network model, and perform transfer, update, and iterative operations on the point features and the edge features respectively until the number of iterations reaches the set number of iterations.

[0056] The graph convolutional neural network model can be a D-MPNN model. In this step, the initial feature vector is input into the D-MPNN model. Since the initial feature vector contains the relevant information of points and chemical bonds, the initial feature vector also contains the information that should be in the adjacency matrix, that is, it contains the adjacent relationship between different points. As Figure 2 shown, the D-MPNN model includes a message passing module, a node update module, an edge update module, and a readout module. The message passing module first receives the initial feature vector. According to the adjacent relationship of points, it generates messages from its neighbor points for each point and performs multiple rounds of message passing. After each round of message passing, the node update module can update the features of each point according to the messages of each point and its neighbor points. The edge update module performs weighted sum or applies a non-linear activation function to the messages of each passed point to integrate the messages from different neighbor points and adjacent edges, so as to update the edge features. The updated point features and edge features enter the message passing module again to iterate to the next round, affecting the message passing of the next round. Repeat the above processes of message passing, edge update, and point update, and iterate multiple times to fully capture the relationships and features between each point in the molecular graph until the set number of iterations is reached. In this solution, the set number of iterations can be determined according to the system data processing ability and the required prediction duration. For example, if it is required to obtain the prediction results of ADMET properties in a short time, and / or the system's data processing ability is average, the number of iterations can be selected as a smaller value. If it is required to obtain the prediction results of ADMET properties in a long time, and / or the system's data processing ability is strong, the number of iterations can be selected as a larger value. In this embodiment, 3 times is selected as an example.

[0057] S203: Recombine the updated point features and updated edge features output by the graph convolutional neural network model to obtain the graph feature vector.

[0058] After reaching the set number of iterations, the readout module is used to aggregate the features of all points and edges to obtain the global molecular feature vector A. The global molecular feature vector A has multiple dimensions. In this embodiment, 300 dimensions are taken as an example.

[0059] The commonly used MPNN (Message Passing Neural Networks) model is a graph convolutional neural network model, and D-MPNN is an improved graph convolutional neural network model. D-MPNN, like MPNN, has two stages: message passing and state update. The information transmitted by D-MPNN adopted in this solution is not only the correlation between atoms, but also includes the correlation between chemical bonds, which can reduce the repetition in information transmission compared with MPNN.

[0060] In the above solution, in step S200, processing the SMILES sequence into a sequence global feature vector includes:

[0061] S211: Input the SMILES sequence into the language learning model, perform word segmentation, encoding, vector mapping and iterative operations on the SMILES sequence until the number of iterations reaches the set number of iterations.

[0062] Among them, the language learning model can be selected as BERT (Bidirectional Encoder Representations from Transformers), that is, the pre-trained language representation model. Combined with Figure 3 the shown framework, the language learning model includes a word segmentation module, an embedding module, an encoding module and a recurrent iteration network.

[0063] In this step, the tokenization module first obtains the pre-trained model data of the BERT model. For the SMILES sequence of the compound molecule to be predicted in this step, its string is determined, and the string is tokenized using the vocabulary of the pre-trained model data used by BERT. The embedding module performs one-hot encoding on the tokenized SMILES sequence, and then uses the embedding module to map the corresponding data tokens and their positions in the encoding result to a continuous vector space, where the data tokens are directly mapped to the corresponding vectors, and the position encoding is implemented through the embedding vectors, thereby retaining the position information of each encoding result in the sequence. After the above embedding and position encoding, a continuous vector is obtained. The continuous vector is input into the encoding module, and the encoding module uses the multi-head attention mechanism to allow the module to calculate self-attention in parallel, so as to simultaneously focus on various relationships between different parts of the input continuous vector, avoid losing important information when processing longer vectors, and thus capture richer context information. After the multi-head attention layer, the encoding module adds the input and the output to form a residual connection, which helps to alleviate the problem of gradient disappearance in deep networks. The result after the residual connection is layer-normalized to keep the mean and variance of the output of each layer relatively stable. The output after the residual connection and layer normalization will be passed into the feed-forward neural network, and the input and output of the feed-forward neural network will be passed into the residual connection and layer normalization again. The above attention layer, residual connection, and feed-forward neural network are iterated N times through the recurrent network. As mentioned above, the value of N can be selected according to the data processing ability of the system and the requirements for prediction time. In this solution, it can be selected as 3 times.

[0064] S212: Use the output vector obtained by mapping the vector in the last iteration process as the sequence global feature vector.

[0065] In each iteration, the language learning model will learn a higher-level feature representation. In the feed-forward neural network, the output of each layer will be used as the input of the next layer, gradually deepening the understanding of the input SMILES sequence. After being processed for the set number of iterations, the result output by the vector mapping will represent the sequence global feature vector of the entire SMILES sequence, which is represented by the 384-dimensional feature vector B in this solution.

[0066] In the above solution, processing the SMILES sequence into a molecular fingerprint vector in step S200 includes:

[0067] S221: Input the SMILES sequence into the RDKit tool to extract the molecular fingerprint corresponding to the SMILES sequence.

[0068] In specific implementation, the RDKit tool can be directly used to convert the SMILES sequence of the compound molecule to be predicted into an RDKit2DNormalized fingerprint, and this fingerprint is used as the molecular fingerprint. The molecular fingerprint may include a total of 200 global molecular features that can be calculated by the RDKit tool, covering the simple physical and chemical properties to complex molecular descriptors of the compound molecule to be predicted, and is used to make up for the structural information that is difficult to learn in the graph convolutional neural network model and the language learning model.

[0069] S222: Normalize the molecular fingerprint to obtain the molecular fingerprint vector.

[0070] The molecular fingerprint is normalized by a normalization method to obtain a molecular fingerprint vector. In this solution, taking the 200-dimensional feature vector C as an example, this solution can improve the stability and performance of the ADMET property prediction model.

[0071] In the above solution, step S300 includes: using the vector obtained by sequentially arranging and connecting the graph feature vector, the sequence global feature vector, and the molecular fingerprint vector corresponding to the same compound molecule to be predicted as the joint molecular representation of the same compound molecule to be predicted. As mentioned above, assume that a compound molecule to be predicted obtains a 300-dimensional graph feature vector A, a 384-dimensional sequence global feature vector B, and a 200-dimensional molecular fingerprint vector C after being processed in step S200. Then, the above vectors are sequentially spliced in the order of A, B, and C, and finally an 884-dimensional joint molecular representation is obtained. In this solution, the representation of the compound molecule to be predicted corresponds to three categories: SMILES sequence, molecular graph, and molecular fingerprint. The SMILES sequence is concise, does not require complex structure conversion, is easy to encode and process, and is applicable to a variety of compounds. The molecular graph directly uses the molecular structure for modeling, regards atoms as points and chemical bonds as edges, retains the topological information and spatial structure of the molecule, can more intuitively capture the internal relationships of the molecule, and better represents the interactions between molecules. The molecular fingerprint is a binary vector generated by feature encoding of the molecular structure, and can capture the chemical characteristics of the molecule to a certain extent. Compared with the molecular representation methods of the prior art, the above joint molecular representation more fully and comprehensively considers the potential characteristics of the compound molecule to be predicted.

[0072] The ADMET property prediction model adopted in the above solution is trained in the following way:

[0073] S110: Obtain a compound data set, and perform standardization processing and information annotation on the compounds in the compound data set to obtain a sample compound data set. The information annotation includes compound attributes and the attribute values of the compound attributes.

[0074] Specifically, collect open-source ADMET data information, which can be collected from Therapeutics Data Commons, ChEMBL, OCHEM, DeepChem, and various literatures. Classify and merge the ADMET data information by attribute. After merging, the data only retains the SMILES column, the corresponding attribute column, and the data source column. For example, for solubility as an attribute, its corresponding value is 3.1 for solubility. Standardize the SMILES through the RDKit tool, including removing atomic mapping numbers, deleting redundant hydrogen atom labels, unifying isomers, standardizing the representation of all rings to ensure consistent ring numbering, normalizing atoms and bonds according to atomic sequence or molecular structure and rearranging, removing metal atoms and complexes, etc., to perform relevant processing to make the representation of molecules more consistent and accurate. Filter out the missing values where the standardized SMILES is empty, which can be achieved using Pandas. For classification tasks, if the same SMILES from different sources has different labels, discard it; for regression tasks, after unifying the units, if the same SMILES from different sources has different values, and the difference between the maximum value and the minimum value is less than 10% of the average value, take the average value, otherwise discard the SMILES.

[0075] After the above processing, the labeled molecular ADMET dataset is obtained as the compound dataset in this step.

[0076] S120: Obtain the combined molecular characterization of each sample compound in the sample compound dataset and the corresponding ADMET properties; the combined molecular characterization includes the graph feature vector of the sample compound molecule, the sequence global feature vector of the sample compound molecule, and the molecular fingerprint vector of the sample compound molecule.

[0077] The processing method for obtaining the combined molecular characterization in this step is similar to the methods described in the previous steps S200 and S300. As Figure 4As shown, the SMILES sequence passes through the embedding layer, multi-head attention layer, residual and normalization layer, feed-forward neural network, and N-loop iterations of residual and normalization of the language learning model to obtain the sequence global feature vector; the SMILES sequence is processed by the RDKit tool to obtain the molecular fingerprint; the molecular graph undergoes graph feature encoding, message passing, node update, edge update, and readout by the graph convolutional neural network model to obtain the graph feature vector, and then the three vectors are concatenated to obtain the joint molecular representation. The graph feature vector extracts the global relationship between atoms and chemical bonds in the molecular graph and is a high-level abstraction of the molecular geometry and topological structure. The sequence global feature vector extracts the sequence pattern and structural features under the SMILES sequence and contains deep sequence information related to chemical semantics. The molecular fingerprint vector extracts various properties and structural features of the molecule and is a concise expression of the molecular characteristics and specific structural patterns. By concatenating the three vectors, a comprehensive molecular feature representation can be obtained.

[0078] S130: Train the deep learning algorithm using the joint molecular representation and ADMET properties of the compound samples to obtain the ADMET property prediction model; the deep learning algorithm includes a feed-forward neural network.

[0079] Specifically, ten different random seeds can be set, and different random seeds can ensure the randomness of different sub-grouping results. The compounds under each ADMET property are randomly divided into a training set, a validation set, and a test set at a ratio of 8:1:1 under a specific random seed, resulting in a total of ten different data random division groups. Select a deep feed-forward neural network and use the above data division groups for training to finally obtain the ADMET property prediction model.

[0080] Preferably, in the above training process: determine the task type of the sample compound according to the information annotation of the compound sample; determine the loss function and evaluation index in the training process of the feed-forward neural network according to the task type of the sample compound. Specifically, the task types include classification tasks and regression tasks, and the classification tasks include binary classification tasks and multi-class classification tasks. During the training process, cross-entropy is used as the loss function for classification tasks, and accuracy and AUROC (Area Under the ROC Curve; the ROC curve represents the relationship between the true positive rate and the false positive rate at different thresholds) are used as evaluation indicators. For binary classification tasks, the final model output is passed through the Sigmoid function to limit the output value within the range of (0,1). For multi-class classification, the final model output is transformed using the Softmax function so that the sum of the classification scores across classes is 1. The regression task uses the mean squared error as the loss function and RMSE (root mean squared error) and R 2 (coefficient of determination) as the evaluation indicator.

[0081] During the training process, by default, the deep learning algorithm trains with random samples for at least 100 epochs. In the training process of the deep learning algorithm, the Adam optimizer is used, which has an adaptive learning rate and can automatically adjust the learning rate at different stages. Specifically, in the first two training epochs, the Adam optimizer is used as a warm-up period to linearly increase the learning rate from 10 -4 to 10 -3 , and then in the remaining training epochs, the learning rate is exponentially decreased from 10 -3 to 10 -4 . The batch size of the Adam optimizer can be selected as 50 data. The model with the minimum AUROC or RMSE for each property prediction task is saved as the best model to determine the optimal weights and parameters. Table 1 shows the AUROC and accuracy evaluation metrics corresponding to different classification tasks, and Table 2 shows specific examples of using RMSE (root mean square error) and R 2 (coefficient of determination) as evaluation metrics for different regression tasks. The results shown in Table 1 and Table 2 are the averages of ten groups of randomly divided different data.

[0082] Table 1 Classification tasks use AUROC and accuracy as evaluation metrics

[0083] Task AUROC Accuracy bioavailability_20 0.839257 0.8322 bioavailability_50 0.851387 0.790423 PAMPA 0.76233 0.856931 HIA_90 0.951118 0.894676 HIA_30 0.954186 0.951988 Pgp_inhibition 0.94675 0.900078 Pgp_substrate 0.906663 0.84096 BBB 0.9552 0.903008 Cyp_inhibition 0.912479 0.862658 Cyp_substrate 0.866687 0.792574 HIV 0.797691 0.720546 AMES 0.897164 0.825686 Carcinogen 0.777581 0.721111 DILI 0.835896 0.760987 Eye_corrosion 0.989649 0.956925 Eye_irritation 0.978734 0.944636 Hematotoxicity 0.821287 0.763813 Hemolytic_toxicity 0.964278 0.889583 hERG_10uM 0.917624 0.845128 hERG_1uM 0.938649 0.877983 LD50_mouse 0.866021 0.805667 LD50_rat 0.858545 0.799556 Mitochondrial_toxicity 0.8893 0.839931 Nephrotoxic 0.806564 0.735762 Ototoxicants 0.739457 0.688208 Respiratory 0.855271 0.795566 Skin_reaction 0.841329 0.774073 Tox21 0.875933 0.934313

[0084] Table 2 Regression tasks use RMSE and R 2 as evaluation metrics

[0085]

[0086]

[0087] As shown in Table 1, for the tasks bioavailability_20 and bioavailability_50, they respectively represent using 20% and 50% as thresholds. Labels with bioavailability above the threshold are set to 1, and otherwise to 0. In their training process, cross-entropy is used as the loss function, and the AUROC reaches 0.839257 and the Accuracy reaches 0.8322, meeting the requirements of the evaluation metrics. In Table 2, the tasks CL_microsome, CL_microsome(mouse), and CL_microsome(rat) respectively refer to the clearance rates in humans, mice, and rats. The mean squared error is used as the loss function. During the training process of CL_microsome, the RMSE reaches 0.4732 and R 2 reaches 0.676858, meeting the requirements of the evaluation metrics.

[0088] Using the ADMET property prediction model provided by the above solution in combination with the joint molecular representation to predict the ADMET properties of compound molecules, and predicting the ADMET properties of the same compound molecules using the model obtained by combining the conventional molecular representation method in the prior art with conventional training, and judging its performance according to specific evaluation indexes. The performance comparison results are shown in Table 3:

[0089] Table 3 Performance Comparison between the Solution of the Present Application and the Existing Solution

[0090]

[0091] In the examples shown in the above table, the results of predicting the ADMET properties of compound molecules using the ADMET property prediction model of the present application in combination with the joint molecular representation are all better than the prediction results of the existing solutions. On the one hand, the solution of the present application adopts the method of joint molecular representation, which can avoid the loss of characteristics of compound molecules and fully mine the data characteristics. On the other hand, the solution of the present application uses different models to process different molecular characteristics. For example, the graph feature vector is processed by the D-MPNN model, and the sequence global feature vector is processed by the BERT model, which helps to improve the robustness of ADMET property prediction and the accuracy of prediction results.

[0092] The embodiment of the present application also provides a computer-readable storage medium, in which program information is stored. After the computer reads the program information, it executes the steps of the method for predicting the ADMET properties of compounds based on joint molecular representation according to any one of the above method embodiments.

[0093] The embodiment of the present application also provides a computer program product, including a computer program / instructions, and when the computer program / instructions are executed by a processor, the steps of the method for predicting the ADMET properties of compounds based on joint molecular representation according to any one of the above method embodiments are implemented.

[0094] The embodiment of the present application also provides an electronic device, such as Figure 5As shown, the electronic device includes at least one processor 501 and at least one memory 502. Program information is stored in at least one of the memories 502. After reading the program information, at least one of the processors 501 executes the method for predicting the ADMET properties of a compound based on a combined molecular representation according to any of the above method embodiments. The device may further include: an input device 503 and an output device 504. The processor 501, the memory 502, the input device 503, and the output device 504 may be communicatively connected. The memory 502, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules. By running the non-volatile software programs, instructions, and modules stored in the memory 502, the processor 501 executes various functional applications and data processing, that is, implements the method for predicting the ADMET properties of a compound based on a combined molecular representation provided by any of the above solutions. The memory 502 may include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created according to the use of the method for predicting the ADMET properties of a compound based on a combined molecular representation, etc. In addition, the memory 502 may include a high-speed random access memory and may also include non-volatile memories, such as at least one disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some embodiments, the memory 502 may optionally include a memory remotely provided with respect to the processor 501, and these remote memories can be connected to the device for executing the method for predicting the ADMET properties of a compound based on a combined molecular representation through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof. The input device 503 can receive input user clicks and generate signal inputs related to user settings and function controls of the method for predicting the ADMET properties of a compound based on a combined molecular representation. The output device 504 may include a display device such as a display screen. When the one or more modules are stored in the memory 502 and are run by the one or more processors 501, they execute the method for predicting the ADMET properties of a compound based on a combined molecular representation in any of the above method embodiments.

[0095] According to needs, the above technical solutions can be combined to achieve the best technical effect.

[0096] The above are only the principles and preferred embodiments of this application. It should be noted that for those of ordinary skill in the art, based on the principles of this application, several other variations can also be made, which should also be regarded as the protection scope of this application.

Claims

1. A method for predicting the ADMET properties of compounds based on combined molecular characterization, characterized in that: include: Obtain the molecular graph and SMILES sequence of the compound to be predicted; Processing the molecular graph into a graph feature vector, and processing the SMILES sequence into a sequence global feature vector and a molecular fingerprint vector; The graph feature vector, the sequence global feature vector and the molecular fingerprint vector are concatenated to obtain a joint molecular representation; The combined molecular characterization is input into an ADMET property prediction model, and the ADMET property prediction model outputs the predicted ADMET property corresponding to the compound molecule to be predicted; wherein the ADMET property prediction model is trained by a deep learning algorithm.

2. The method for predicting compound ADMET properties based on combined molecular characterization according to claim 1, characterized in that: The processing of the molecular graph into a graph feature vector comprises: Performing graph feature encoding on the molecular graph to obtain an initial feature vector; the graph feature encoding includes: taking atoms in the molecular graph as point features and taking chemical bonds in the molecular graph as edge features; Inputting the initial feature vector into the graph convolutional neural network model, transferring, updating and iterating the point features and the edge features respectively until the number of iterations reaches the set number of iterations; The updated point features and updated edge features output by the graph convolutional neural network model are recombined to obtain the graph feature vector.

3. The method for predicting compound ADMET properties based on combined molecular characterization according to claim 1, characterized in that: Processing the SMILES sequence into a sequence global feature vector includes: Input the SMILES sequence into the language learning model, perform word segmentation, encoding, vector mapping and iterative operation on the SMILES sequence until the number of iterations reaches the set number of iterations; The output vector obtained by the vector mapping in the last iteration process is used as the global feature vector of the sequence.

4. The method for predicting compound ADMET properties based on combined molecular characterization according to claim 1, characterized in that: Processing the SMILES sequence into a molecular fingerprint vector includes: Input the SMILES sequence into the RDKit tool to extract the molecular fingerprint corresponding to the SMILES sequence; The molecular fingerprint is normalized to obtain the molecular fingerprint vector.

5. The method for predicting compound ADMET properties based on combined molecular characterization according to claim 1, characterized in that: The step of concatenating the graph feature vector, the sequence global feature vector and the molecular fingerprint vector to obtain a joint molecular representation comprises: The graph feature vector, the sequence global feature vector and the molecular fingerprint vector corresponding to the same compound molecule to be predicted are sequentially arranged and connected to obtain a vector as the joint molecular representation of the same compound molecule to be predicted.

6. The method for predicting compound ADMET properties based on combined molecular characterization according to any one of claims 1 to 5, characterized in that: The training of the ADMET property prediction model includes: Acquire a compound data set, and obtain a sample compound data set after standardization and information annotation of the compounds in the compound data set, wherein the information annotation includes compound attributes and attribute values ​​of the compound attributes; Obtaining a joint molecular representation of each sample compound in a sample compound data set and a corresponding ADMET property; the joint molecular representation includes a graph feature vector of the sample compound molecule, a sequence global feature vector of the sample compound molecule, and a molecular fingerprint vector of the sample compound molecule; The deep learning algorithm is trained using the combined molecular characterization and ADMET properties of the compound sample to obtain an ADMET property prediction model; the deep learning algorithm includes a feedforward neural network.

7. The method for predicting compound ADMET properties based on combined molecular characterization according to claim 6, characterized in that: The deep learning algorithm is trained using the combined molecular characterization and ADMET properties of the compound sample to obtain an ADMET property prediction model, including: Determining the task type of the sample compound according to the information annotation of the sample compound; The loss function and evaluation index in the feedforward neural network training process are determined according to the task type of the sample compound.

8. A computer-readable storage medium, characterized in that: The storage medium stores program information, and after the computer reads the program information, the computer executes the steps of the method for predicting the ADMET properties of compounds based on combined molecular characterization according to any one of claims 1 to 7.

9. A computer program product, characterized in that The invention comprises a computer program / instruction, characterized in that when the computer program / instruction is executed by a processor, the steps of the method for predicting the ADMET properties of compounds based on combined molecular characterization according to any one of claims 1 to 7 are implemented.

10. An electronic device comprising a memory, a processor and a computer program stored in the memory, characterized in that: The processor executes the computer program to implement the steps of the method for predicting ADMET properties of compounds based on combined molecular characterization according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Molecular ADMET property prediction method and model based on fusion fingerprints

    CN118522372A

Cited By

  • Method and system for predicting admet properties of drug-like small molecules

    CN122822137A