A small molecule drug design method, device, electronic equipment and storage medium

By using the TransGEM model to generate small molecule drugs based on gene expression differences, the inefficiency problem in drug development for diseases with unknown targets has been solved. This approach enables the efficient and low-cost generation of new compounds and is applicable to the development of drugs for novel diseases and diseases with unknown targets.

CN117059196BActive Publication Date: 2025-11-07HUAZHONG AGRI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310941750.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-07-28
Publication Date
2025-11-07
Estimated Expiration
2043-07-28

AI Technical Summary

Technical Problem

Existing technologies are inefficient in drug development for diseases with unknown targets, rely on large amounts of computing resources and lack flexibility, making it difficult to effectively generate new compounds with biological activity.

Method used

Using the TransGEM model, a deep learning approach involving gene expression encoders, molecular embedding layers, decoders, and generators is employed to generate small molecule drugs by leveraging the gene expression differences between diseased cells and normal tissues, thus avoiding dependence on target information.

Benefits of technology

It improves drug development efficiency, reduces costs, and enables the generation of new compounds with potential biological activity, making it suitable for drug design targeting diseases with unknown objectives.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117059196B_ABST
    Figure CN117059196B_ABST
Patent Text Reader

Abstract

The application discloses a small-molecule drug design method and device, electronic equipment and storage medium. The method comprises the following steps: acquiring a compound related to a research disease and gene expression data before and after the compound treating cells, calculating the difference between the gene expression data before and after the compound treating cells, and constructing a data set with the compound and the gene expression data; inputting the data set into a pre-constructed TransGEM model to complete the training of the TransGEM model, wherein the input vector of the TransGEM model is the compound and the corresponding difference, and the output vector of the TransGEM model is a reconstructed compound; acquiring gene expression data of a specific disease cell and a corresponding normal tissue cell, and calculating the gene expression difference; inputting the gene expression difference into the trained TransGEM model to generate a small-molecule drug corresponding to the specific disease. The application is easy to use and efficient, and can be used for small-molecule drug design for diseases with unknown targets.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of biological medicine, and in particular to a small molecule drug design method and device, electronic equipment and storage medium. BACKGROUND

[0002] Drug development is an expensive, complex, lengthy, and low success rate process (Zhu H. (2020). Big Data and Artificial Intelligence Modeling for Drug Discovery. Annual Review of Pharmacology and Toxicology. 60:573-589; Chan HCS., et al. (2019). Advancing Drug Discovery via Artificial Intelligence. Trends in Pharmacological Sciences. 40:592-604). The cost of developing a new drug and bringing it to market is as high as $3 billion, and it takes an average of 10 years (Chan HCS., et al. (2019). Advancing Drug Discovery via Artificial Intelligence. Trends in Pharmacological Sciences. 40:592-604; Cavasotto CN., and Di Filippo JI. (2021). Artificial intelligence in the early stages of drug discovery. Archives of Biochemistry and Biophysics. 698:108730). Reducing costs and speeding up the development of new drugs is a great challenge facing the pharmaceutical industry. Computer-aided drug development is a prominent method of modern preclinical drug discovery, which applies machine learning / deep learning methods to multiple stages of drug development (Sabe VT., et al. (2021). Current trends in computer aided drug design and a highlight of drugs discovered via computational techniques: A review. European Journal of Medicinal Chemistry. 224:113705).Computer-aided drug development has the potential to improve the existing drug development industry by providing rational guidance for drug discovery (Wang L., et al. (2019). Artificial intelligence facilitates drug design in the big data era. Chemometrics and Intelligent Laboratory Systems. 194:103850).

[0003] Traditionally, computer-aided drug development relies on two strategies: drug repurposing and de novo drug design. Related studies show that the synthetic chemical space can contain about 10 60 - 10 100 molecules, of which 10 23 - 10 60 molecules can be potential drug-like compounds, but the number of synthesized molecules is only 10 8 - 10 10(Xia X., et al. (2019). Graph-based generative models for de Novo drug design. Drug Discovery Today: Technologies. 32-33: 45-53; Polishchuk PG., Madzhidov TI., and Varnek A. (2013). Estimation of the size of drug-like chemical space based on GDB-17 data. J Comput Aided Mol Des. 27:675-679). This indicates that the available chemical space and molecular diversity for drug repurposing are significantly limited. De novo drug design can further explore the chemical space in an automated manner to generate new compounds (Pereira T., et al. (2021). Diversity oriented Deep Reinforcement Learning for targeted molecule generation. J Cheminform. 13: 1-17). De novo drug design is usually divided into ligand-based or receptor-based methods (Xu M., Ran T., and Chen H. (2021). De Novo Molecule Design Through the Molecular Generative Model Conditioned by 3D Information of Protein Binding Sites. J. Chem. Inf. Model. 61:3240-3254). Ligand-based de novo drug design generates new molecules based only on prior knowledge of a given structural ligand (Wang Z., et al. (2020). Combined strategies in structure-based virtual screening. Physical Chemistry Chemical Physics. 22:3149-3159). However, most of the generated molecules deviate from the expected properties and lack biological activity.Receptor-based de novo drug design largely depends on target protein information, especially three-dimensional structure information of the target protein (Robson B. (2022). De novo protein folding on computers. Benefits and challenges. Computers in Biology and Medicine. 143: 105292). However, the target protein of a new disease is still not determined, and the target protein of some complex diseases has not been determined or the structure has not been resolved. These diseases pose a challenge to receptor-based de novo drug design methods.

[0004] Early drug discovery mainly focused on phenotypic drug discovery (Vincent F., et al. (2022). Phenotypic drug discovery: recent successes, lessons learned and new directions. Nat Rev Drug Discov. 1-16). From the beginning of the molecular biology revolution to the completion of the human genome sequencing, the focus of drug discovery gradually shifted to specific disease targets (Vincent F., et al. (2022). Phenotypic drug discovery: recent successes, lessons learned and new directions. Nat Rev Drug Discov. 1-16). However, with the continuous progress of computer performance, research on phenotypic drug discovery has gradually recovered since 2011 (Swinney DC., and Lee JA. (2020). Recent advances in phenotypic drug discovery. F1000Res. 9: F1000 Faculty Rev-944). Phenotypic drug discovery combined with artificial intelligence has become increasingly mature, and it has become an accepted drug discovery model by the academic and pharmaceutical communities (Vincent F., et al. (2022). Phenotypic drug discovery: recent successes, lessons learned and new directions. Nat Rev Drug Discov. 1-16).Gene expression profiles can be used to characterize cellular and biological phenotypes

[18] and have been successfully applied to drug repurposing (Donner Y., Kazmierczak S., Fortney K. (2018). Drug Repurposing Using Deep Embeddings of Gene Expression Profiles. Mol. Pharmaceutics. 15:4314-4325), drug mechanism analysis (Rickardson L., et al. (2005). Identification of molecular mechanisms for cellular drug resistance by combining drug activity and gene expression profiles. Br J Cancer. 93:483-492), lead compound discovery (Hassane DC., et al. (2008). Discovery of agents that eradicate leukemia stem cells using an in silico screen of public gene expression data. Blood. 111:5654-5662) and side effect prediction (Wood JR., et al. (2005). Valproate-induced alterations in human theca cell gene expression: clues to the association between valproate use and metabolic side effects. Physiological Genomics. 20:233-243). Some studies have shown that there is some intrinsic connection between gene expression and drug structure (Zhu J., et al. (2021). Prediction of drug efficacy from transcriptional profiles with deep learning. Nat Biotechnol. 1-9). Gene expression data has also been used for de novo drug design.Born et al. (Born J., et al. (2021). PaccMannRL: De novo generation of hit-like anticancer molecules from transcriptomic data via reinforcement learning. iScience. 24: 102269) used two variational autoencoder models to learn the embedding and generation rules of gene expression profiles and compound molecules, respectively. The encoder of the gene expression variational autoencoder model is connected to the decoder of the molecule variational autoencoder model, so that new compound molecules can be generated based on gene expression profiles. However, this study needs to retrain the model with sufficient data when generating molecules for a specific disease, and the application flexibility is insufficient. Moreover, the model also relies on reinforcement learning to improve the quality of generated molecules, which consumes a lot of computing resources. Mendez-Lucio et al. (Mendez-Lucio O., et al. (2020). De novo generation of hit-like molecules from gene expression signatures using artificial intelligence. Nat Commun. 11: 10) designed a conditional generative adversarial network framework with gene expression profiles and Gaussian noise as input. Two discriminators are used to evaluate the correlation between the generated molecules and the gene expression profiles and their authenticity. Through adversarial training and conditional constraints, the model framework generates new compounds that meet the given conditions. However, this study generates molecules using gene expression data from disease target gene knockouts, which also relies on target information of the disease. SUMMARY

[0005] The purpose of the present application is to overcome the above technical deficiencies, provide a small molecule drug design method, device, electronic equipment and storage medium, which is easy to use, high efficiency, can be used for small molecule drug design for diseases with unknown targets, provides a new idea for disease research and development, and has a broad application prospect in the field of new disease drug research and development and unknown target disease drug research and development.

[0006] To achieve the above technical purpose, the present application adopts the following technical solutions:

[0007] In a first aspect, the present application provides a small molecule drug design method, comprising the following steps:

[0008] Obtaining compounds related to the research disease and gene expression data before and after the compounds treat cells, and calculating the difference between the gene expression data before and after the compounds treat cells, then constructing a data set with the compounds and the gene expression data;

[0009] inputting the dataset into a pre-constructed TransGEM model to complete training of the TransGEM model, wherein an input vector of the TransGEM model is the compound and the corresponding difference value, and an output vector of the TransGEM model is a reconstructed compound;

[0010] obtaining gene expression data of a specific disease cell and a corresponding normal tissue cell, and calculating a gene expression difference value;

[0011] inputting the gene expression difference value into the trained TransGEM model to generate a small molecule drug corresponding to the specific disease.

[0012] In some embodiments, the TransGEM model at least includes a gene expression encoder, a molecular embedding layer, a decoder, and a generator connected in sequence.

[0013] In some embodiments, the decoder is composed of stacked decoder layers, and the decoder layers are composed of a masked multi-head self-attention layer, a multi-head attention layer, and a feedforward neural network layer, the multi-head attention layer is used to extract an attention matrix between genes and molecules, and the attention matrix is used to calculate an attention score of each gene with respect to a molecule to identify a potential disease target.

[0014] In some embodiments, the attention matrix is specifically:

[0015]

[0016] wherein, is an attention matrix, Q represents a query matrix, K represents a key matrix, V represents a value matrix, d K represents a dimension of a row vector of the key matrix, and T indicates matrix transposition.

[0017] In some embodiments, the gene expression encoder is used to encode the input difference value into x ∈ R (n×d) , wherein d = d_g + d_c, n represents a number of genes in each gene expression, x represents encoded data, R represents a real number, d_g represents a dimension of difference value embedding, and d_c represents a dimension of disease corresponding cell line embedding.

[0018] In some embodiments, the TransGEM model further includes a loss function used to optimize parameters of the model, wherein the loss function is a Kullback-Leibler divergence, and the mathematical formula is:

[0019]

[0020] wherein n represents the number of samples, y represents the molecular representation corresponding to the samples, represents the reconstructed molecular representation corresponding to the samples.

[0021] In some embodiments, the result of the gene expression difference value is kept to one decimal place.

[0022] In a second aspect, the present application provides a small molecule drug design device, comprising:

[0023] a training data construction module, configured to obtain compounds related to a research disease and gene expression data before and after the compounds treat cells, and calculate the difference between the gene expression data before and after the compounds treat the cells, and construct a data set with the compounds and the gene expression data;

[0024] a training module, configured to input the data set into a pre-constructed TransGEM model to complete the training of the TransGEM model, wherein the input vector of the TransGEM model is the compound and the corresponding difference value, and the output vector of the TransGEM model is a reconstructed compound;

[0025] a data acquisition module, configured to obtain gene expression data of a specific disease cell and corresponding normal tissue cells, and calculate a gene expression difference value;

[0026] an output module, configured to input the gene expression difference value into the trained TransGEM model to obtain a small molecule drug corresponding to the specific disease.

[0027] In a third aspect, the present application provides an electronic device, comprising a processor and a memory;

[0028] The memory has stored a computer program which can be executed by the processor;

[0029] The processor executes the computer program to implement the steps in the small molecule drug design method described above.

[0030] In a fourth aspect, the present application provides a computer readable storage medium, which stores one or more programs which can be executed by one or more processors to implement the steps in the small molecule drug design method described above.

[0031] Compared with the prior art, the small molecule drug design method, device, electronic equipment and storage medium provided by the application are based on the association between gene expression data and molecular drugs, adopt a deep learning method, and generate small molecule compounds with disease treatment potential for a specific disease through a TransGEM model, so that the efficiency of drug research and development can be effectively improved. In addition, the small molecule drug design method has the advantages of low cost, easy use and high efficiency, and has wide application prospects in the fields of new disease drug research and development and unknown target disease drug research. BRIEF DESCRIPTION OF DRAWINGS

[0032] Figure 1 is a flowchart of the small molecule drug design method provided by the embodiment of the application;

[0033] Figure 2 is a framework diagram of the TransGEM model in the small molecule drug design method provided by the embodiment of the application;

[0034] Figure 3 is a framework diagram of the gene expression encoder and the Transformer decoder in the small molecule drug design method provided by the embodiment of the application;

[0035] Figure 4 is a property distribution diagram of a molecule generated by the TransGEM test result in the small molecule drug design method provided by the embodiment of the application;

[0036] Figure 5 is a property distribution diagram of a molecule generated by two specific embodiments of the application;

[0037] Figure 6 is a structure diagram of a molecule generated in the first specific embodiment of the application and a docking diagram of the molecule and a corresponding target;

[0038] Figure 7 is a structure diagram of a molecule generated in the second specific embodiment of the application and a docking diagram of the molecule and a corresponding target;

[0039] Figure 8 is a functional module schematic diagram of the small molecule drug design device provided by the embodiment of the application;

[0040] Figure 9 is a hardware structure schematic diagram of the electronic equipment provided by the embodiment of the application. DETAILED DESCRIPTION

[0041] In order to make the purpose, technical scheme and advantages of the application more clear, the application will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the application and do not limit the application.

[0042] Referring to Figure 1 The small molecule drug design method provided by the present application comprises the following steps:

[0043] S100, obtaining the gene expression data of a compound related to a research disease and the gene expression data of the compound before and after treating cells, and calculating the difference between the gene expression data before and after treating the cells, and then constructing a data set with the compound and the gene expression data;

[0044] S200, inputting the data set into a pre-constructed TransGEM model to complete the training of the TransGEM model, wherein the input vector of the TransGEM model is the compound and the corresponding difference, and the output vector of the TransGEM model is a reconstructed compound;

[0045] S300, obtaining the gene expression data of a specific disease cell and a corresponding normal tissue cell, and calculating the gene expression difference;

[0046] S400, inputting the gene expression difference into the trained TransGEM model to obtain a small molecule drug corresponding to the specific disease.

[0047] In this embodiment, based on the correlation between gene expression data and small molecule drugs, a deep learning method is used to generate small molecule compounds with disease treatment potential for a specific disease through a TransGEM model, which can effectively improve the efficiency of drug research and development. In addition, the small molecule drug design method has the advantages of low cost, easy use and high efficiency, and has broad application prospects in the fields of new disease drug research and development and unknown target disease drug research and development.

[0048] In some embodiments, the step S100, the compound related to the disease and the gene expression data before and after the compound treatment of the cell are from databases including but not limited to LINCS (https: / / lincsproject.org / LINCS) database. The data from different databases can be processed into corresponding forms. For example, the subLINC dataset is constructed based on the LINCS 1000 database. The data in the dataset is from the 3-level data of the small molecule compound in the LINCS 1000 database. Considering that in addition to the type of compound, different compound doses and perturbation times also affect the final gene expression state of the cell line, only the gene expression profile of the small molecule compound with a dose of 10 μM and a perturbation time of 24 hours is retained. In the LINCS 1000 database, only the expression values of 978 landmark genes are actually measured, and the expression values of the other 11350 genes are inferred from the 978 landmark genes. If the expression values of all 12350 genes are used, too much noise may be introduced, affecting the performance of the model. Therefore, the gene expression profile of each sample in the dataset only retains the expression values of the 978 landmark genes. Since the model needs the gene expression change value of the compound molecule after perturbing the cell line, the final data obtained is the expression difference value of the 978 landmark genes of each sample and its control sample. In order to ensure the relative accuracy of the gene expression difference value, the "compound molecule + cell line" is taken as the label, and the data containing only 1 sample under the same label is removed. For repeated samples under the same label, the data is combined and de-duplicated by taking the median. Finally, the expression difference value of the 978 landmark genes of each sample is retained to 1 decimal place. Finally, only the data related to the 14 cell lines (A375, A549, HA1E, HEK293, HELA, HEPG2, HT29, JURKAT, MCF10A, MCF7, MDAMB231, PC3, THP1 and YAPC) with the largest number of samples are retained and used to construct the subLINCS dataset.

[0049] In some embodiments, the step S200, the TransGEM model is a phenotype-based small molecule de novo design model. The model can generate small molecule compounds with disease treatment potential only using the gene expression difference between the diseased cell line and the normal tissue cell, and does not depend on the target information of the disease.

[0050] Optionally, please refer to Figures 2-3 , the TransGEM model at least includes a gene expression encoder, a molecular embedding layer, a decoder and a generator connected in sequence.

[0051] In this embodiment, the cell line names and gene expression difference values of 978 landmark genes before and after molecular perturbation are encoded by a gene expression encoder. The corresponding self-referencing embedded strings (SELFIES) of these molecules are embedded by a molecular embedding layer. The above two pieces of information are input into the decoder of the Transformer model. The purpose is to understand the relationship between the self-action of the molecule and the cell line-specific gene expression changes disturbed by the molecule through the decoder. Finally, the self of these molecules is reconstructed by a generator. In the process of generating molecules, the attention matrix between the 978 landmark genes and the molecule is extracted from the Transformer model decoder, and based on this attention matrix, the attention score of each gene relative to the molecule is calculated, and the genes with high attention scores may become potential disease targets.

[0052] In some embodiments, since the gene expression difference value is the size of the degree of gene expression change, it is not suitable for one-hot encoding form. Secondly, if the gene expression difference value is directly input into the model, it will cause the embedding matrix to be too sparse and lose a lot of difference information. Therefore, this study proposes a unique gene expression encoder, as shown in Figure 3 The cell line names corresponding to a specific disease and the expression difference values of 978 landmark genes of the cell line and normal tissue are encoded by a gene expression encoder into x∈R (n×d) n=978 represents the number of genes in each gene expression; d=d_g+d_c, where d_g=9 is the dimension of the gene expression difference embedding, and d_c=15 is the dimension of the disease corresponding cell line embedding. Specifically, the expression difference value of each gene is kept to one decimal place and then encoded into a 9-dimensional vector, where the first dimension represents whether the gene is up-regulated or down-regulated relative to the control sample. Gene expression up-regulation (expression difference value greater than 0) is 1, and gene expression down-regulation (expression difference value less than 0) is 0. The 2nd to 9th dimensions represent the degree of change in gene expression relative to the control sample. Take the absolute value of the gene expression difference value, then perform a 10-fold transformation, and then convert it to a binary representation. If the binary representation is less than 8 bits, fill the front with 0. The cell line corresponding to a specific disease is encoded by one-hot encoding form into a 15-dimensional vector, where the first 14 dimensions correspond to the 14 cell lines in the training data set, and the last 1 dimension represents an unknown cell line, mainly for model expansion.

[0053] In some embodiments, for the molecular embedding layer, the molecule is represented as SELFIES, which can be easily decomposed into tokens of the form "[*]". All non-redundant SELFIES in the subLINCS dataset are split into tokens. These tokens are used to build a dictionary, where tokens with a frequency less than 5 are discarded. The start, end, and unknown tokens are defined as<sos> 、 <eos>and <unk>Tag <sos>Located before a molecule SELFIES, indicates the start of each molecule SELFIES. Tag <eos>Located after the molecule SELFIES, represents the end of each molecule SELFIES. Token <unk>all tokens not in the dictionary are replaced. The three tokens mentioned earlier are also added to the dictionary, resulting in a final dictionary size of 52. Finally, based on this dictionary, each molecule represented by a SELFIES is embedded into a where m denotes the number of tokens of SELFIES, d m denotes the embedding dimension (d m = 64).

[0054] In some embodiments, please refer to Figures 2-3 For the decoder, the decoder is composed of stacked decoder layers. The decoder layer is composed of a masked multi-head self-attention layer, a multi-head attention layer and a feed-forward neural network layer. Each layer is subject to residual connection and layer normalization. The masked multi-head self-attention layer is used to learn the context information of Input1 itself. The multi-head attention layer is used to learn the internal association information between Input1 and Input2, and the attention matrix is obtained from this layer. The calculation of attention involves three groups of matrices, namely the query (Q) matrix, the key (K) matrix and the value (V) matrix. Attention can be mathematically represented by the following formula:

[0055]

[0056] where is the attention matrix; d K denotes the dimension of the row vector of the K matrix; and T indicates matrix transposition.

[0057] In some embodiments, the TransGEM model further comprises a loss function for optimizing the parameters of the model, wherein the loss function is Kullback-Leibler divergence, and the mathematical formula is:

[0058]

[0059] where n denotes the number of samples, y denotes the molecule representation corresponding to the sample, denotes the reconstructed molecule representation corresponding to the sample.

[0060] In order to better understand the present application, the following will be combined Figures 4-7 The technical solutions of the application are described in detail, and the related terms are explained as follows: LINCS 1000: gene expression database; effectiveness: the proportion of effective molecules in all newly generated molecules; uniqueness: the proportion of non-repeated molecules in effective molecules; novelty: the proportion of molecules not appearing in the training data set in effective molecules; internal diversity: a quantitative indicator of the diversity of generated unique effective molecules; QED: Quantitative Estimation of Drug-likeness, used to evaluate the drug-likeness of compounds; SA: Synthetic Accessibility, used to evaluate the chemical synthesis difficulty and cost of compounds.

[0061] Specifically, first, taking the LINCS 1000 database as an example, the subLINC data set is constructed, and the TransGEM model is trained and preliminarily evaluated using the subLINCS data set. 500 compound molecules are generated for each sample of the test set, and the performance of each evaluation index of the generated molecules is shown in Table 1 and Figure 4 As shown in the table. The results show that for all samples, the TransGEM model can generate 100% effective new compound molecules. Excellent novelty indicates that the TransGEM model does not have overfitting phenomenon and has the ability to generate new molecules outside the training data set. The average uniqueness of the newly generated molecules is 87.4%, and the uniqueness of the molecules generated for other cell lines is above 80% except for the HEPG2 cell line. This shows that the TransGEM model can pay attention to the differences in different gene expression changes, so it does not tend to generate repeated molecules. The average internal diversity of the newly generated molecules also basically reaches 80%. Moreover, the newly generated molecules are basically consistent with the distribution of the octanol / water partition coefficient (LogP), molecular weight, QED value and SA score of the molecules in the subLINCS data set. The above results show that the TransGEM model has excellent overall performance.

[0062] Table 1 Evaluation of TransGEM model on subLINCS data set

[0063]

[0064] Based on the above argument, in the first embodiment, the method of the application is used for small molecule drug design for prostate cancer, and the specific steps are as follows:

[0065] I. Collect gene expression data

[0066] The prostate cancer corresponding cell line and the prostate corresponding cell line related gene expression data were collected in the LINCS1000 database. Among them, the prostate cancer corresponding cell line selects PC3, and the prostate corresponding cell line selects LHSAR. Then, the gene expression difference of 978 landmark genes between LHSAR cell line and PC3 cell line is calculated, and 1 decimal place is retained.

[0067] II. Small molecule drug design for prostate cancer using TransGEM model

[0068] The above obtained gene expression difference is used as the input of the TransGEM model, and the cell line is set to PC3. 1000 small molecule compounds are generated using the TransGEM model.

[0069] III. Evaluation of the generated results

[0070] The effectiveness, novelty, uniqueness, internal diversity, QED and SA of the above generated 1000 small molecule compounds are evaluated, and the results are shown in Table 2 and Figure 5 The results show that 100% effective new compound molecules can also be generated for prostate cancer, and the uniqueness and diversity are more than 80%, and the properties of the generated molecules are comparable to FDA approved drugs.

[0071] Table 2 Evaluation of molecules generated for prostate cancer

[0072]

[0073] IV. Molecular docking verification

[0074] For prostate cancer, a known drug target is poly ADP ribose polymerase 1 (PARP1) (Deshmukh D., and Qiu Y. (2015). Role of PARP-1 in prostate cancer. Am J Clin Exp Urol. 3: 1-12). The PARP1 protein with its ligand olaparib is a drug approved by the US Food and Drug Administration for the treatment of PC, and its complex crystal structure has been determined (Deeks ED. (2015). Olaparib: first global approval. Drugs. 75: 231-240). The crystal structure of the PARP1 protein was downloaded from the PDB database (https: / / www.rcsb.org / ), with PDB ID 7kk4 (Ryan K., et al. (2021). Dissecting the molecular determinants of clinical PARP1 inhibitor selectivity for tankyrase 1. J Biol Chem. 296: 100251). The attention matrix of the TransGEM model in generating molecules targeting PC was extracted, and then the attention scores of the 978 landmark genes corresponding to each generated molecule were calculated. The molecules corresponding to the top 500 PARP1 genes were retained, and a total of 1000 molecules were obtained. These retained molecules were constructed into a molecular library for docking with the PARP1 protein. In addition, two molecules with good docking results, P1 and P2, were analyzed. As shown in FIGS. Figure 6 A-C, both molecules P1 and P2 exhibit ideal drug similarity and lower synthetic complexity. In addition, similarity analysis shows that the structures of P1 and P2 have high novelty, which are different from the structure of olaparib. The docking scores of molecules P1 and P2 with the PARP1 protein are comparable to the docking score of olaparib. Previous studies have shown that olaparib forms stable hydrogen bonds with the His862, Gly863, Tyr896, and Ser904 residues of the PARP1 protein, among which Gly863 and Ser904 are key active amino acids (Ryan K., et al. (2021). Dissecting the molecular determinants of clinical PARP1 inhibitor selectivity for tankyrase 1. J Biol Chem. 296: 100251). Analysis of the docking results of molecules P1 and P2 with the PARP1 protein shows that these two molecules can also form stable hydrogen bonds with the Gly863 and Ser904 residues of the PARP1 protein ( Figure 6 D and 6E).

[0075] V. Prostate cancer potential target discovery

[0076] In the process of generating molecules targeting prostate cancer, the top 10 genes with the highest attention scores were collected from the TransGEM model (Table 3). Surprisingly, all 10 genes have been reported to be associated with the onset of prostate cancer, with PARP1 being a target protein that has been approved for PC treatment

[53] . This indicates that when generating molecules for a specific disease, the TransGEM model does indeed pay more attention to genes that are more closely related to the onset of the disease.

[0077] Table 3 Top 10 genes in attention when generating molecules for prostate cancer

[0078]

[0079]

[0080] In the second example, the method of the present application is used for small molecule drug design for non-small cell lung cancer, with the following specific steps:

[0081] I. Collect gene expression data

[0082] Gene expression data related to the corresponding cell lines of non-small cell lung cancer and the corresponding cell lines of lung tissue are collected in the LINCS 1000 database. Among them, the A549 cell line is selected for non-small cell lung cancer, and the NL20 cell line is selected for lung tissue. Then, the gene expression difference of 978 landmark genes between the NL20 cell line and the A549 cell line is calculated, and 1 decimal place is retained.

[0083] II. Small molecule drug design for prostate cancer using the TransGEM model

[0084] The gene expression difference obtained above is used as the input of the TransGEM model, and the cell line is set to A549. 1000 small molecule compounds are generated using the TransGEM model.

[0085] III. Evaluation of the generated results

[0086] For the 1000 small molecule compounds generated above, their effectiveness, novelty, uniqueness, internal diversity, QED and SA properties are evaluated, and the results are shown in Tables 4 and Figure 5 The results show that 100% effective new compound molecules can be generated for non-small cell lung cancer, and the uniqueness and internal diversity of the generated molecules can reach more than 80%. The distribution of the 4 main properties of the generated molecules is basically consistent with the FDA-approved drugs.

[0087] Table 4 Evaluation of molecules generated in non-small cell lung cancer

[0088]

[0089]

[0090] IV. Molecular docking verification

[0091] For non-small cell lung cancer (NSCLC), the known drug target is epidermal growth factor receptor (EGFR) (Stewart EL., et al. (2015). Known and putative mechanisms of resistance to EGFR targeted therapies in NSCLC patients with EGFR mutations—a review. Transl Lung Cancer Res. 4:67–81). The ligand that binds to the EGFR protein crystal structure is dacomitinib, a drug approved by the U.S. Food and Drug Administration (FDA) for the treatment of NSCLC (Brzezniak C., Carter CA., and Giaccone G. (2013). Dacomitinib, a new therapy for the treatment of non-small cell lung cancer. Expert Opin Pharmacother. 14:247–253). The crystal structure of the EGFR protein was downloaded from the PDB database (https: / / www.rcsb.org / ), PDB ID 4I24 (Gajiwala KS., et al. (2013). Insights into the Aberrant Activity of Mutant EGFR Kinase Domain and Drug Recognition. Structure. 21:209–219). The attention matrix was extracted from the TransGEM model when generating molecules targeting non-small cell lung cancer. Then, the attention scores for 978 marker genes corresponding to each generated molecule were calculated. Molecules corresponding to the top 500 EGFR genes in terms of attention were retained, resulting in a total of 379 molecules. These 379 molecules were used to construct a molecular library for docking with the EGFR protein. Furthermore, two molecules with good docking results, N1 and N2, were analyzed. Figure 7 As shown in A-C, both molecules N1 and N2 show good QED scores and SA scores. Compared with dacomitinib, molecules N1 and N2 show higher structural novelty. Related studies show that dacomitinib forms stable hydrogen bonds with the Gln791, Met793, Cys797, Asp800 and Asp855 residues of the EGFR protein, among which Cys797 is a key active amino acid (Gajiwala KS., et al. (2013). Insights into the Aberrant Activity of Mutant EGFR Kinase Domain and Drug Recognition. Structure. 21: 209-219). In this study, although the docking scores of molecules N1 and N2 with the EGFR protein are lower than those of dacomitinib, the docking results show that these two molecules can also form stable hydrogen bonds with the Cys797 residue of the EGFR protein Figure 7 D and 7E).

[0092] V. Discovery of potential targets for prostate cancer

[0093] In the process of generating molecules targeting prostate cancer, the top 10 genes with the highest attention scores were collected from the TransGEM model (Table 5), of which 8 of the 10 genes were reported to be associated with the onset of the disease. This shows that when generating molecules for a specific disease, the TransGEM model does indeed pay more attention to genes more closely related to the onset of the disease.

[0094] Table 5 Top 10 genes in attention when generating molecules for non-small cell lung cancer

[0095]

[0096] From the above analysis, it can be seen that for the examples of prostate cancer and non-small cell lung cancer, the TransGEM model can generate small molecule compounds with structural novelty and potential biological activity, which demonstrates the feasibility and effectiveness of the present application in generating small molecule drugs for diseases.

[0097] Another embodiment of the present application provides a small molecule drug design device, please refer to Figure 8 The small molecule drug design device comprises a training data construction module 11, a training module 12, a data acquisition module 13 and an output module 14.

[0098] The training data construction module 11 is used to acquire compounds related to the research disease and gene expression data before and after the compounds treat cells, and calculate the difference between the gene expression data before and after the compounds treat cells, and then construct a data set with the compounds and the gene expression data.

[0099] The training module 12 is configured to input the data set into a pre-constructed TransGEM model to complete training of the TransGEM model, wherein an input vector of the TransGEM model is the compound and the corresponding difference value, and an output vector of the TransGEM model is a reconstructed compound.

[0100] The data acquisition module 13 is configured to acquire gene expression data of a specific disease cell and a corresponding normal tissue cell, and calculate a gene expression difference value.

[0101] The output module 14 is configured to input the gene expression difference value into the trained TransGEM model to obtain a small molecule drug corresponding to the specific disease.

[0102] In the embodiment, based on the association between the gene expression data and the small molecule drug, a deep learning method is used to generate a small molecule compound with disease treatment potential for a specific disease through the TransGEM model, which can effectively improve the efficiency of drug research and development. In addition, the small molecule drug design method has the advantages of low cost, easy use and high efficiency, and has a wide application prospect in the field of new disease drug research and development and unknown target disease drug research and development.

[0103] It should be noted that the module referred to in the present application refers to a series of computer program instruction segments capable of completing a specific function, and is more suitable for describing the execution process of small molecule drug design than the program. The specific implementation of each module is described in the above method embodiment, which will not be described here.

[0104] In some embodiments, the TransGEM model at least includes a gene expression encoder, a molecular embedding layer, a decoder and a generator connected in sequence.

[0105] In some embodiments, the decoder is composed of stacked decoder layers, and the decoder layers are composed of a masked multi-head self-attention layer, a multi-head attention layer and a feedforward neural network layer. The multi-head attention layer is configured to extract an attention matrix between genes and molecules, and the attention matrix is configured to calculate an attention score of each gene relative to the molecule to identify a potential disease target.

[0106] In some embodiments, the attention matrix is specifically:

[0107]

[0108] wherein, is the attention matrix, Q represents a query matrix, K represents a key matrix, V represents a value matrix, d K represents the dimension of the row vector of the key matrix, and T indicates matrix transposition.

[0109] In some embodiments, the gene expression encoder is configured to encode the input difference value into x e R (n×d) where d = d_g + d_c, n represents the number of genes in each gene expression, x represents the encoded data, R represents a real number, d_g represents the dimension of the difference value embedding, and d_c represents the dimension of the disease corresponding cell line embedding.

[0110] In some embodiments, the TransGEM model further comprises a loss function configured to optimize the parameters of the model, wherein the loss function is Kullback-Leibler divergence, and the mathematical formula is:

[0111]

[0112] where n represents the number of samples, y represents the molecular representation corresponding to the sample, represents the reconstructed molecular representation corresponding to the sample.

[0113] In some embodiments, the result of the gene expression difference value retains 1 decimal place.

[0114] Another embodiment of the present application provides an electronic device, such as Figure 9 As shown in the figure, the electronic device 10 comprises:

[0115] one or more processors 110 and memories 120, Figure 9 In the embodiment, the processor 110 and the memory 120 can be connected through a bus or other means, Figure 9 In the embodiment, the connection through the bus is taken as an example.

[0116] The processor 110 is configured to complete various control logics of the electronic device 10, which can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), a single-chip microcomputer, an ARM (Acorn RISC Machine), or other programmable logic devices, discrete gates or transistor logic, discrete hardware components, or any combination of these components. In addition, the processor 110 can also be any conventional processor, microprocessor or state machine. The processor 110 can also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors in combination with a DSP and / or any other such configuration.

[0117] The memory 120, as a non-volatile computer readable storage medium, can be used to store non-volatile software programs, non-volatile computer executable programs and modules, such as program instructions corresponding to the small molecule drug design method in the embodiments of the present application. The processor 110 executes various function applications and data processing of the electronic device 10 by running the non-volatile software programs, instructions and units stored in the memory 120, that is, implements the small molecule drug design method in the above method embodiments.

[0118] The memory 120 can include a program storage area and a data storage area, wherein the program storage area can store an operating platform and application programs required by at least one function; the data storage area can store data created according to the use of the electronic device 10, etc. In addition, the memory 120 can include a high-speed random access memory, and can also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state memory device. In some embodiments, the memory 120 can optionally include a memory remotely arranged with respect to the processor 110, and these remote memories can be connected to the electronic device 10 through a network. Examples of the above network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof.

[0119] One or more units are stored in the memory 120, and when executed by the one or more processors 110, the small molecule drug design method in any of the above method embodiments is executed, for example, the method steps S100 to S400 in the above described Figure 1 are executed.

[0120] Another embodiment of the present application provides a computer readable storage medium, which stores computer executable instructions, and the computer executable instructions are executed by one or more processors, for example, the method steps S100 to S400 in the above described Figure 1 are executed.

[0121] By way of example, computer readable storage media can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM), which acts as external cache memory. By way of example, and not limitation, RAM can be provided in numerous forms such as synchronous RAM (SRAM), dynamic RAM (DRAM), RAM, synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), Synchlink DRAM (SLDRAM), and direct Rambus RAM (DRRAM). Combinations of the above can be used. The disclosed memory components or memory of operating environments described herein are intended to include one or any such combinations of memory components.

[0122] To sum up, the small molecule drug design method, device, electronic equipment and storage medium provided by the present application are based on the association between gene expression data and molecular drugs, and use a deep learning method to generate small molecule compounds with disease treatment potential for a specific disease through a TransGEM model, which can effectively improve the efficiency of drug research and development. In addition, the small molecule drug design method has the advantages of low cost, easy use and high efficiency, and has a wide application prospect in the field of new disease drug research and development and unknown target disease drug research and development.

[0123] The specific embodiments of the present application described above do not constitute a limitation on the scope of protection of the present application. Any various other corresponding changes and modifications made according to the technical concept of the present application shall be included in the scope of protection of the claims of the present application.< / unk> < / eos> < / sos> < / unk> < / eos> < / sos>

Claims

1. A method of small molecule drug design, characterized by, The method comprises the following steps: obtaining gene expression data of a compound related to a research disease and cells before and after the compound is used to treat the cells, and calculating the difference between the gene expression data before and after the cells are treated by the compound, and then constructing a data set with the compound and the gene expression data; inputting the data set into a pre-constructed TransGEM model to complete the training of the TransGEM model, wherein the input vector of the TransGEM model is the compound and the corresponding difference, and the output vector of the TransGEM model is a reconstructed compound; obtaining gene expression data of a specific disease cell and a corresponding normal tissue cell, and calculating the gene expression difference; inputting the gene expression difference into the trained TransGEM model to generate a small molecule drug corresponding to the specific disease; the TransGEM model at least comprises a gene expression encoder, a molecular embedding layer, a decoder and a generator connected in sequence; the decoder is composed of stacked decoder layers, the decoder layers are composed of a masked multi-head self-attention layer, a multi-head attention layer and a feedforward neural network layer, the multi-head attention layer is used to extract an attention matrix between genes and molecules, and the attention matrix is used to calculate an attention score of each gene relative to the molecule to identify potential disease targets; The gene expression encoder is used to encode the input difference value as x e R n×d wherein d = d_g + d_c, n represents the number of genes in each gene expression, x represents the encoded data, R represents a real number, d_g represents the dimension of the difference value embedding, and d_c represents the dimension of the disease corresponding cell line embedding.

2. The method of small molecule drug design according to claim 1, wherein, the attention matrix is specifically: wherein, is an attention matrix, Q denotes a query matrix, K denotes a key matrix, and V denotes a value matrix, denotes a dimension of a row vector of the key matrix, and T indicates a matrix transpose.

3. The method of small molecule drug design of claim 1, wherein, the TransGEM model further comprises a loss function, and the loss function is used to optimize the parameters of the model, wherein the loss function is Kullback-Leibler divergence, and the mathematical formula is: , where n denotes the number of samples, denotes a molecular representation corresponding to the sample, denotes a reconstructed molecular representation corresponding to the sample.

4. The method of small molecule drug design of claim 1, wherein, the result of the gene expression difference is kept to one decimal place.

5. A small molecule drug design apparatus, characterized by, comprise: a training data construction module for obtaining gene expression data of a compound related to a research disease and cells before and after the compound is used to treat the cells, and calculating the difference between the gene expression data before and after the cells are treated by the compound, and then constructing a data set with the compound and the gene expression data; a training module for inputting the data set into a pre-constructed TransGEM model to complete the training of the TransGEM model, wherein the input vector of the TransGEM model is the compound and the corresponding difference, and the output vector of the TransGEM model is a reconstructed compound; a data acquisition module for obtaining gene expression data of a specific disease cell and a corresponding normal tissue cell, and calculating the gene expression difference; an output module for inputting the gene expression difference into the trained TransGEM model to obtain a small molecule drug corresponding to the specific disease; the TransGEM model at least comprises a gene expression encoder, a molecular embedding layer, a decoder and a generator connected in sequence; the decoder is composed of stacked decoder layers, the decoder layers are composed of a masked multi-head self-attention layer, a multi-head attention layer and a feedforward neural network layer, the multi-head attention layer is used to extract an attention matrix between genes and molecules, and the attention matrix is used to calculate an attention score of each gene relative to the molecule to identify potential disease targets; The gene expression encoder is used to encode the input difference value as x e R n×d wherein d = d_g + d_c, n represents the number of genes in each gene expression, x represents the encoded data, R represents a real number, d_g represents the dimension of the difference value embedding, and d_c represents the dimension of the disease corresponding cell line embedding.

6. An electronic device, comprising: A computer program product comprising: a processor and a memory; a computer program stored on said memory and executable by said processor; said processor implementing the steps of the method of designing small molecule drugs according to any one of claims 1-4 when executing said computer program.

7. A computer-readable storage medium, characterized in that, A computer program product comprising: a processor and a memory; a computer program stored on said memory and executable by said processor; said processor implementing the steps of the method of designing small molecule drugs according to any one of claims 1-4 when executing said computer program.