Workflows for generating compounds having biological activity against specific biological targets
By combining generative adversarial networks and reinforcement learning in the GENTRL model, compound structures are generated and prioritized, solving the problem of low compound generation efficiency in traditional drug discovery and achieving efficient and low-cost generation of bioactive compounds.
Patent Information
- Application Number
- CN202080073784.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-08-23
- Filing Date
- 2020-08-22
- Publication Date
- 2026-02-10
- Estimated Expiration
- 2040-08-22
AI Technical Summary
In current drug discovery processes, traditional methods are time-consuming, costly, and have low success rates. There is a lack of effective methods for de novo molecular design and biovalidation, especially in generating small molecule compounds with biological activity.
A generative tensor reinforcement learning (GENTRL) model combining generative adversarial networks (GANs) and reinforcement learning (RL) is adopted. By receiving biological target input, the model is trained to generate compound structures. Sammon mapping and Kohonen self-organizing map (SOM) are used for priority ranking, combined with pharmacophore modeling and synthesis verification, to achieve efficient generation and biological evaluation of compounds.
It significantly reduces the time required to generate bioactive compounds from scratch, improves the novelty and synthetic feasibility of compounds, ensures that the generated compounds have high efficiency of activity against specific biological targets, and reduces the cost and time of drug discovery.
Smart Images

Figure CN114667498B_ABST
Abstract
Description
Cross-references to related applications
[0001] This patent application claims priority to U.S. Provisional Application No. 62 / 891,050, filed August 23, 2019, which is incorporated herein by reference in its entirety. Technical Field
[0002] This invention relates to a workflow protocol for identifying therapeutic targets, acquiring background data, training models, generating structures, prioritizing compounds, synthesizing compounds, and biologically evaluating compounds to identify compounds that regulate therapeutic targets. Background Technology
[0003] Related fields description
[0004] Drug discovery is a multidisciplinary field that requires the seamless integration of multiple computational and experimental disciplines. The process typically involves designing, synthesizing, and testing a large number of compounds in an iterative, trial-and-error manner, with most being discarded due to lack of bioactivity, selectivity, physicochemical properties, unfavorable pharmacokinetic (PK) characteristics, unacceptable toxicity, or other issues. Recent advances in artificial intelligence (AI), particularly in generative adversarial networks (GANs) and reinforcement learning (RL), can help design new compounds that have been placed within an acceptable safety and therapeutic window, thereby increasing the likelihood of successful clinical trials.
[0005] The traditional discovery pathway for new molecular entities (NMEs) or first-in-class drugs is a complex and resource-intensive activity that typically takes 10-20 years, with each NME costing between USD 500 million and USD 2.6 billion. 1-3 This uncertainty stems from reliance on tests with unknown clinical relevance and from a lack of rigorous critical evaluation of available information. Failures during the preclinical or clinical evaluation phases are a major contributor to increased drug discovery costs due to lack of clinical efficacy, unacceptable pharmacokinetic characteristics, unacceptable toxicity and selectivity, and evolving regulatory or commercial strategic positioning. 1 Many drug candidates fail to reach the clinical evaluation stage or obtain regulatory approval. Smaller organizations strive to innovate drug discovery projects whose potential therapeutic importance extends beyond early-stage development.
[0006] Many organizations involved in drug discovery are exploring the integration and deployment of AI systems to enhance internal research and development (R&D) for projects involving new compounds, improved target bioactivity and selectivity, pharmacokinetic characteristics, toxicity, solubility, metabolic stability, overall accessibility, and other properties. 4 Despite skepticism about the potential impact of computer models, AI systems have made tremendous progress in many disciplines and technological fields in recent years. 5-7The applications of AI range from computer vision, speech and text analysis, route selection, and autonomous driving to medical diagnostics, drug discovery, and biotechnology. 8 AI technology can identify patterns hidden in massive amounts of data and perform reliable and complex generalizations that traditional data analysis techniques cannot achieve. Many experts in the field of drug discovery predict that AI-driven instruments will become increasingly important and prominent in the near future. 9-11 .
[0007] One innovative AI approach is generative adversarial network technology. 12 GANs have been used to "imagine" new, realistic photographic images, videos, and music with desired sets of properties, and have the potential to generate chemical structures. 13 In drug discovery, while the theoretical basis for GAN applications has been studied, there are no experimental examples demonstrating the ability to generate novel, diverse, and effective small molecule compounds. Early experiments with GANs used binary fingerprint representations of molecular structures. 14,15 It needs to match a known chemical space. String-based representations allow for the generation of new structures for synthesis. 16-20 Much research has focused on graphics and 3D structures. 21 This allows for new applications. Segler et al. extended generative models through transfer and reinforcement learning, using predicted binding affinity as the target. 22 Merk et al. demonstrated the application of SMILES-based recurrent neural networks for generating active molecules using transfer learning. 23,24 Although recent AI developments have focused on de novo molecular design, one of its main drawbacks is the lack of a complete experimental validation cycle: from computer design to chemical synthesis and biological evaluation.
[0008] Therefore, having advanced, experimentally validated methods to generate new, synthetically feasible compounds that are active against biomolecular targets would be advantageous. Summary of the Invention
[0009] In some embodiments, the computer-implemented method may include: receiving input of a biological target; receiving a generative model (e.g., a tensor reinforcement learning (GENTRL) model) trained with a reference compound; generating a structure of the generative compound using the generative model; prioritizing the structures of the generative compounds based on at least one criterion; processing the preferred chemical structures of the generative compounds via a Sammon mapping protocol to obtain the hit structure; and providing the chemical structure of the hit structure. In some aspects, the reference compound includes: a general compound, a compound that modulates a biological target, and a compound that modulates biomolecules other than a biological target.
[0010] In some embodiments, the computer-implemented method may include: receiving input of a biological target or ligand for any biological target (e.g., a biological target or other biological target); receiving input of the properties of the generated compound; receiving at least one generative model trained with a reference compound; generating a structure of the generated compound having each generative model, wherein the generated compound is designed to interact with the biological target and / or be associated with structural features of the ligand; prioritizing the structures of the generated compounds for each generative model based on at least one reward criterion; processing the preferred chemical structures of the generated compounds via a Sammon mapping protocol to obtain the hit structures; and providing the chemical structures of the hit structures.
[0011] In some embodiments, one or more non-transitory computer-readable media are provided storing instructions that, in response to being executed by one or more processors, cause a computer system to perform operations including performing the computer methods described herein to provide the chemical structure of the generated hit structure through a generative model.
[0012] In some embodiments, a computer system may include: one or more processors; and one or more non-transitory computer-readable media storing instructions that, in response to being executed by one or more processors, cause the computer system to perform operations, including performing the methods described herein to provide the chemical structure of the generated hit structure via a generative model.
[0013] The foregoing is merely illustrative and is not intended to be limiting in any way. Other aspects, embodiments, and features will become apparent from the accompanying drawings and the following detailed description, in addition to the illustrative aspects, embodiments, and features described above. Attached Figure Description
[0014] The above and following information, as well as other features of the invention, will become more apparent from the accompanying drawings, the following description, and the appended claims. It should be understood that these drawings merely illustrate a few embodiments of the invention and are therefore not intended to limit its scope; the invention is described with additional specificity and detail using the drawings.
[0015] Figure 1A Examples of workflows for generating novel compounds with biological activity against specific biological targets are shown.
[0016] Figure 1B Examples of workflows for generating novel compounds that match or are associated with selected ligands (e.g., ligands of known biological targets) are shown, such that the generated compounds are associated with the 3D shape, space, hydrogen bonds, and other characteristics of the selected ligands.
[0017] Figure 2A-2C This shows an example of the pharmacophore hypothesis being implemented.
[0018] Figure 3 This includes a nonlinear Sammon map of 40 selected molecules labeled with triangles.
[0019] Figure 4A The dose-response curves include those for compounds 1-6.
[0020] Figure 4B Including ICs displaying compounds 2 and 4 50 The image.
[0021] Figure 4C The structures of compounds 1-6 and their ICs for DDr1 and DDR2 are shown. 50 value.
[0022] Figures 5A-5C Quantum mechanical calculations show the best-fit pharmacophore hypothesis.
[0023] Figure 6 Representative examples of the generated structures are shown compared to the parent DDR1 inhibitor.
[0024] Figure 7 Selectivity curves for compound 1 are included.
[0025] Figure 8A-8I This includes data showing that compounds 1 and 2 significantly block DDR1 autophosphorylation in a dose-dependent manner.
[0026] Figures 9A-9F Data include those showing the effects of compounds 1 and 2 on cellular fibrosis markers α-actin and CTGF (normalized to GAPDH) in MRC-5 cells.
[0027] Figure 10A A schematic example of a chemical generation system platform according to a workflow protocol is shown.
[0028] Figure 10B An example of the workflow for operating the generative model is shown.
[0029] Figure 10C An example of the workflow for operating the generative model is shown.
[0030] Figure 11 An example of a chemical generation system platform based on a workflow and generation model is shown.
[0031] Figure 12 An image showing the implementation of the reward function based on Kohonen SOM is displayed.
[0032] Figure 13 Examples of rejected molecules are shown.
[0033] Figure 14 An example of a molecule generated using the GENTRL model is shown.
[0034] Figure 15 An example of a computing system for performing the computational methods described herein is shown.
[0035] The elements and components in the accompanying drawings may be arranged according to at least one embodiment described herein, and such arrangement may be modified according to the invention provided herein by those skilled in the art. Detailed Implementation
[0036] In the following detailed description, reference is made to the accompanying drawings, which form part of the invention. In the drawings, similar reference numerals generally identify similar components unless the context otherwise requires. The exemplary embodiments described in the detailed description, drawings, and claims are not intended to be limiting. Other embodiments may be used, and other changes may be made, without departing from the spirit or scope of the subject matter presented in the invention. It will be readily understood that aspects of the invention as generally described herein and illustrated in the drawings can be arranged, replaced, combined, separated, and designed in a variety of different configurations, all of which are expressly contemplated within the invention.
[0037] Generally, this technology relates to a workflow for generating compounds that have specific activity against biological targets (e.g., DDR1 or others). The workflow may include four connected components: an input and configuration component; a generation component; a reward component; and an output component that outputs generated compounds, which have been prioritized by the reward component.
[0038] In some embodiments, the input and configuration components may be configured to receive and process input data for a biological target, as well as input data for specific properties of the generated compounds that have specific activities on the biological target. These specific properties may include an acceptable range of rewards for the reward component. For example, specific properties may allow the input data to specify acceptable physicochemical properties, the overall accessibility of the compound, and other properties evaluated using a set of modules in the reward and scoring protocol.
[0039] In some embodiments, the generative component may include at least one model capable of generating compounds, such as compounds with specific activity against biological targets. For example, the generative component may include up to 30 or more AI models, including GENTRL. GENTRL is one example of a generative model. However, the generative component of the workflow is applicable to many other different generative models with different architectures (e.g., GANs, genetic algorithms, RNNs, etc.).
[0040] In some embodiments, the reward component is configured to feed back rewards linked to a single generated compound to the generation component to enable active learning. Similarly, the reward component includes all modules and combinations thereof for performing reward priority ranking. Reward priority ranking allows filtering out suboptimal compounds. In some aspects, this individual structure is evaluated in the reward component as each model (e.g., from 30+) of the generation component generates a chemical structure. Depending on the generation time, the workflow protocol can process millions of generated and rewarded compounds.
[0041] In some embodiments, the output component is configured to select newly generated and scored chemical structures for the hit compound. For example, the Sammon mapping protocol can be used to select newly generated and scored chemical structures for the hit compound.
[0042] In some embodiments, this document illustrates experimental validation of small molecules designed using generative adversarial networks (GANs). This invention provides generative tensor reinforcement learning (GENTRL) models or other generative models or combinations thereof for de novo drug design, which have been shown to be used to generate novel inhibitors of DDR1 kinase. Within 21 days, the GENTRL model generated six compounds. Two of these compounds were novel, low-number nanomolar lead inhibitors with high activity. Further biological and animal studies confirmed antifibrotic activity, microsomal stability, and pharmacokinetic (PK) properties. Therefore, this invention can be used to generate compounds with specific activities (e.g., targeting specific proteins), which can be synthesized and biologically validated.
[0043] In some embodiments, the workflow for generating the protocol can be achieved by using reinforcement learning. 25-27 Variational reasoning 28、29 and tensor decomposition 30-32 The algorithm is formulated using a combination of two generative machine learning steps. The first step generates a large number of potential manifolds for reported compounds. The second step extends the first step to discover new chemical scaffolds for compounds. The protocol can be parameterized using tensor training formats. 32The algorithm learns the structure of manifolds to utilize partially known properties (see below). By combining multiple methods, the algorithm is highly robust: it avoids mode collapse, it uses partially labeled data, and it infers the chemical space of molecules, thereby proposing new and efficient structures not present in the training library.
[0044] In some embodiments, generative protocols including GENTRL models avoid the drawbacks of previously reported generative models. GENTRL avoids the problems of GANS and mode collapse. This overcomes the mode collapse problem that GANs may suffer from, namely their inability to propose an entire chemical class of compounds attractive for drug discovery. Autoencoder-based generative models, such as variational autoencoders (VAEs), are also relevant. 33 And adversarial autoencoders (AAEs). 34 Mode collapse is avoided by compressing the structure space to a latent distribution of a simple organization that typically follows a standard Gaussian distribution. While simple latent distributions are helpful for training machine learning (ML) models, they may not be the optimal mapping of the topology of the latent space of chemical structures. The proposed prior distribution used in the GENTRL model flexibly parameterizes the latent space in a high-dimensional lattice, with a large number of exponentially increasing multidimensional Gaussian numbers at its nodes. This parameterization better links latent codes and properties and handles missing values without explicit input (unlike other semi-supervised models). 35 compared to).
[0045] Previous reinforcement learning-based methods used reward functions that were not very relevant to the generation of new structures because they used trivial rewards, such as the predicted octanol / water partition coefficient. 36 Traditional drug similarity rules 37 and comprehensive accessibility 18,38,39 The current generation protocol uses self-organizing maps (SOM). 40 To estimate how chemical structures relate to reported bioactive compounds (e.g., small molecule kinase inhibitors), we combine time-based structural trends observed in the intellectual property of drug discovery. The method for generating protocols employs three features: 1) SOM trends (e.g., intellectual property trends); 2) general biological pathway SOMs (e.g., general kinase pathways); and 3) specific biological target SOMs (e.g., specific kinases, such as DDR1).
[0046] In some embodiments, the generative protocol uses a trend SOM, a Kohonen-based reward function that distinguishes between “new” and “old” compounds by taking into account the application priority date of the chemical structure disclosed in relevant patents and published patent applications from any country, region, or PCT. Compounds claimed at different times are located in different regions of the Kohonen map. Neurons filled with new chemical entities have been used to actively reward the generative model, making it closer to clusters containing new molecules. Trend SOMs can be used in protocols for any biological target.
[0047] In some embodiments, the generation protocol uses a universal kinase SOM, a Kohonen map mapping that reliably distinguishes kinase inhibitors from other classes of molecules, directing the ML engine toward structures with high kinase levels. These two groups (e.g., kinase inhibitors versus other classes of molecules) are evenly distributed along different regions of the map, providing statistically significant separation. Structures predicted to be kinase inhibitors with a cellular coefficient exceeding 1.3 are then subjected to a Kohonen-based specific classifier, as described below. For other biological targets, this universal SOM may be a universal biological target family SOM.
[0048] In some embodiments, the generation protocol uses a specific kinase SOM that has been trained to isolate DDR1 inhibitors from a total pool of kinase target molecules. DDR1 inhibitors were observed to be distributed within a closely related group of neurons at the edge of the mapping. For other biological targets, this SOM can be a specific biological target SOM.
[0049] In some embodiments, the generation protocol prioritizes the generated structures using a combination of the three SOMs described above, thus placing more emphasis in the final step on nodes occupied by inhibitors of biological targets (e.g., DDR1 inhibitors). This results in obtaining chemical structures that are more likely to be specific inhibitors of a particular biological target.
[0050] In some embodiments, the present invention provides an accelerated workflow protocol 100 for generating the chemical structure of a compound, which in Figure 1A The diagram illustrates that the compound has specific activity against biological targets. Figure 1AThis is a biotarget-based protocol that uses information about a biotarget to design and generate compounds that interact with that biotarget. Workflow protocol 100 includes selecting a biotarget associated with a disease or symptom, or a biotarget for any other biological cause, where modulating the biotarget can produce some benefit (step 102). In this example, DDR1 kinase is selected as the biotarget. Step 102a includes obtaining property inputs for the generated structure, which can be physicochemical, medicinal, reward-responsive, or other properties that can be associated with the desired generated structure. Once the biotarget is selected and the properties are input, protocol 100 can identify reference compounds for training the generative model and for compound generation (step 104). Reference compounds may include different groups of compounds, such as known inhibitors of biological targets (e.g., DDR1 kinase) that have the ability to modulate the biological target above an upper regulatory threshold; known inhibitors of other targets associated with the biological target (e.g., general kinase inhibitors) that do not modulate the biological target or modulate the biological target below a lower regulatory threshold; compounds that are not inhibitors of the biological target or related targets (e.g., non-kinase inhibitory compounds); compounds identified by retrosynthetic analysis (RSA), which facilitates the synthesis of any ultimately identified compound; and compounds identified by analyzing intellectual property (IP) databases, thereby allowing known compounds to be avoided and new compounds to be generated. The obtained reference compounds can be narrowed down by pretreating the reference compounds to remove unfavorable compounds (step 104a), which may be unfavorable for various reasons, such as having a structure unfavorable to biological use or synthesis, poor solubility, or gross errors in structure. In some embodiments, the reference compound may include a general compound, which may be a compound of a known inhibitor of an unrelated target (e.g., a receptor not in the same family or class as the biological target receptor), a compound of an unknown inhibitor of any target, a compound of good medicinal chemistry, or a compound that can be synthesized by a reasonable synthetic protocol.
[0051] Once the database includes reference compounds, a generation protocol (step 106) can be executed, such as a GENTRL model or other models using workflow protocol 100. The generation protocol (step 106) may include model training (step 106a), structure generation (step 106b), reward function refinement (step 106c), and pharmacophore modeling (step 106d). Model training can also be performed in an AI, where the generative model is trained using inputs from the reference compounds in the database. Therefore, embodiments include generative models trained for each type of model. Each trained generative model can generate structures based on the model training. The generated structures of all models can be refined using Kohonen SOMs, such as trend SOMs, general biological target family SOMs, and specific biological target SOMs. In some cases, the generated structure can be refined by excluding known or patented compounds, which is an assessment of the structural novelty of the generated compound. The pharmacophore hypothesis can be realized through pharmacophore modeling to analyze the structure of the scaffold or side group responsible for biological or pharmacological interactions (e.g., modeling the structure with biological targets (e.g., modeling the interaction between the structure and DDR1 kinase), such as docking, binding, dissociation, or other studies). Then, the generation protocol (step 106) can reduce the number of generated structures and refine the structures towards novel and biological target-specific structures.
[0052] Once the generation protocol is established, the generated compounds are refined using a priority protocol through a priority module (step 108). Each priority protocol may include modules for performing rewards and scoring, as described herein. The priority protocol can implement molecular filtering to remove compounds that do not meet certain reward criteria (step 108a), which can be defined and modified for different biological targets (see, for example, step 106c, for refinement of the reward function). Furthermore, molecular filtering can be performed using a medicinal chemistry filter (MCF), which filters out compounds that do not meet certain medicinal chemistry criteria. The MCF filter may include many (e.g., more than 100) different substructure queries and analyses, where compounds that do not meet the criteria can be filtered out. The generated structures can also be processed using cluster analysis and diversity sorting procedures (step 108b). In some cases, a supplier database can be compared with the generated structures so that structures identical to or undesirably similar to known marketable compounds can be filtered from the generated structure group (step 108c). The generated compounds can also be processed again by KohonenSOM to prioritize them (step 108d), such as general kinase SO and specific kinase SOM, which provide additional prioritization of structures that potentially inhibit biological targets (e.g., DDR1 kinase). Models, such as pharmacophore modeling, crystallographic modeling, 3D pharmacophore modeling, and other models that can be used to obtain root mean square deviation (RMSD) values, reflect the fit to the biological target (step 108e).
[0053] The optimized structure can then be processed using the same set of descriptors and the RMSD values output from pharmacophore modeling via the Sammon mapping method (see steps 106d and 108e). Sammon mappings can be constructed, and some (e.g., defined or arbitrary) specific compound structures can be selected.
[0054] In some embodiments, an optional additional intellectual property filter may be implemented to remove compounds that may be covered by patents or protected by other intellectual property rights (step 112). Therefore, the remaining compounds may be novel and patentable.
[0055] The remaining selected compounds are considered "hit" on the biological target, and the hit compounds can then be synthesized (step 114). The synthesis protocol may include analyzing possible synthetic routes to obtain the compounds. Compounds that are not easily synthesized may be removed (e.g., see step 108 for priority sorting).
[0056] The synthesized hit compounds are then analyzed to verify their bioactivity against biological targets, such as the regulation of biological target activity (step 116). The type of analysis varies depending on the biological target. For example, the synthesized hit compounds may be analyzed for binding to and / or regulation of biological targets.
[0057] Then, validated hit compounds with biological activity against the biological target can be provided (step 118). These compounds can be provided for further evaluation, clinical trials, and potential use in patient treatment.
[0058] As an example, the origin of AI-driven novel molecules with nanomolar activity targeting DDR1 kinases can be as follows: Figure 1A As shown, workflow protocol 100 can significantly reduce the timeline required to obtain compounds that have been validated to be biologically active against biological targets. Workflow protocol 100 uses a set of rigorously tuned models to produce efficient structures of molecules that medicinal chemists consider processable. The starting point includes selecting a biological target of interest.
[0059] Figure 1B This includes a workflow protocol 100 for designing other ligands that match or are associated with a provided / selected ligand. In this workflow, the protocol creates the chemical structure of potential ligands based on the chemical structure of the provided / selected ligand. Otherwise, although the focus is on the structure of the provided / selected ligand, Figure 1B The workflow in follows Figure 1A The protocol is to generate new potential ligands that will function similarly to the provided / selected ligands, thus possessing similar biological activity, for example, similar biological activity to the same biological target, although the actual biological target is not explicitly defined. However, the biological target may also be included in the step of receiving input on the properties of the generated compound.
[0060] therefore, Figure 1A This includes protocols for analyzing biological targets, such as binding pockets or other binding features for generating chemical structures that bind to ligands. Figure 1B This includes protocols for analyzing ligands for biological targets, such as analyzing 3D structures, hydrogen bonds, etc., and using ligand information to design additional potential ligands with similar shapes or 3D presence and similar binding characteristics for binding to the same biological targets.
[0061] Biological targets can be selected as examples of suitable biological targets that meet the initial criteria, which are: a) the target must be validated or at least closely associated with a disease or disorder or other condition (e.g., fibrosis) and should be classified as relatively novel in this application; b) the reference dataset should contain a low to moderate (but sufficient) number of molecules to assess the “cognitive” skills of the generative model. For example, five datasets could be used: 1) small molecule compounds targeting biological targets (e.g., inhibitory activity against DDR1 kinase); 2) public databases of reported inhibitors (e.g., kinase inhibitors) as a positive control set; 3) molecules acting on other unrelated biological targets (e.g., non-kinase targets) as negative control sets; 4) patent data on bioactive molecules claimed by top pharmaceutical companies, arranged by priority date; and 5) structural data (e.g., X-ray diffraction data) of known published biological target inhibitors, such as the DDR1 inhibitor in the example. A 1µM IC can be selected. 50 The value serves as a threshold for classifying reference molecules as active or inactive. Databases 2) through 4) can then be preprocessed to eliminate gross outliers and normalize the input chemical space by reducing the number of compounds with similar structures in each cluster.
[0062] Generative protocols offer modeling techniques that can produce and analyze large amounts of data based on available information, pinpointing inferences and objective relationships hidden within research phenomena. In drug discovery, deep learning algorithms in GENTRL or other models can place chemical data in a precise context and reveal deep-seated patterns that are otherwise invisible but potentially valuable to medicinal chemists. Therefore, workflow protocols offer a partial solution to the industry's widely acknowledged problem of declining productivity and cost balance. This workflow demonstrates that generative models are rapidly maturing and, in addition to addressing many other challenges in drug research and development, may become the mainstream tool for de novo development hits and prospect development.
[0063] In some embodiments, generative models can be used in de novo design approaches for compounds to create new chemical entities with desired properties, such as pharmacological activity. Workflow protocols can begin by identifying the most suitable biological targets based on knowledge of disease mechanisms. The first step in the design process is to identify a set of lead compounds to be optimized. The most promising hits (e.g., compounds) are selected, and their property characterization undergoes further optimization. Finally, an efficient method for synthesizing the final drug compound can be established. While standard design methods can create molecular structures that are typically synthesizable within a few reaction steps, these methods often rely on the fact that they depend on well-established chemical knowledge accumulated in the form of synthetic rules or fundamental physical models. Generative models can automatically identify patterns in complex, nonlinear data without requiring manual feature engineering.
[0064] In the generative model example, the GENTRL model operates on an AI-based platform that combines ML and deep learning (DL) models (GANs, autoencoders, RNN-based language models, genetic algorithms, combinatorial methods, ensembles). These methods are combined with RL optimizations and integrated into an end-to-end pipeline. Figure 1A-1B In practice, once the user has entered all the necessary information, the generation protocol will begin with up to 30 models running in parallel. The average duration of a standard generation experiment is approximately 72 hours. This interface allows for real-time tracking of the generation process's progress, as well as the performance and convergence speed of each model. All molecules being generated incrementally can be compared against various metrics using the interactive features of the user interface.
[0065] Further reference Figure 1A-1B The workflow protocol can use a 2D model to filter the generated compounds using a first-line filter, and to filter rewards and scores using a SOM (Systematic Object Model) of the generated compounds. The first-line filter can include modules such as: MCE-18; MCF; ROG5, T-index, similarity; diversity; physicochemical (PC) features; drug similarity, privileged fragments; comprehensive accessibility (SA) score; and clustering. The SOM filter can include: hierarchical active molecule (HAM) dataset; parental SOM; and scaling mapping.
[0066] Furthermore, rewards and scores for 3D models can include: conformation generation (3D conformation, minimization, flexscore); descriptors (similarity to template molecules); pharmacophore modeling (pharmacophore hypothesis, search); shape analysis (shape similarity); and pocket analysis (combining affinity, pocket identification, and scoring). Further structural modifications can be achieved, such as enhanced metabolic stability and bioisosteres (chemical substituents or groups with similar physical or chemical properties that produce biological properties broadly similar to another compound).
[0067] In some embodiments, once all the necessary information has been input, generation in the generation component will begin with up to 30 models running in parallel. Inputs may include 2D or 3D structures (SDF or MOL), X-ray data (target, binding bag), and target names. Reference databases, such as those used for training, can be input. Different AI models can also be identified and input into the system. The average duration of a standard generation experiment is approximately 72 hours. The interface allows for real-time tracking of the generation process's progress, as well as the performance and convergence rate of each model. For example, progress can be displayed on a computer screen. All molecules being generated step-by-step can be compared against various metrics using the interactive features of the user interface.
[0068] for Figure 1A-1B The workflow protocol operates on a platform comprising over 30 high-level deep learning models with multiple generative paradigms based on various molecular representations, such as 3D, graphical, and string representations. Training, validation, and model selection for the reward component, as well as benchmarking and evaluation procedures, are automated and integrated with a comprehensive set of metrics for evaluating generative quality. Molecular novelty and gradation are recommended using an improved reward function that penalizes molecules similar to known structures and promotes the exploration of new chemical spaces. This includes a methodology for designing novel and graded structures as starting points for future generative work, and a variety of tools for designing, scoring, filtering, and ranking molecular conformations with a 3D reward function. Two SA scoring methods are implemented to prioritize molecular structures based on their overall accessibility. See also Figure 10B .
[0069] Figure 10B It shows that it can be used Figure 10A An example of the system's operational workflow. This workflow shows the input components and models used for de novo drug design; the 2D module includes first-line scoring and SOM, and the 3D module includes Confgen, descriptors, PP, shape, pocket, and structural deformation. Rewards and scoring are used to generate compounds with improved properties, provided as virtual hits. Figure 10C Showing the use Figure 10B Example timeline of the workflow.
[0070] Compound generation via RL involves two different levels of reward and scoring—2D modules and 3D modules. The 2D module set consists of two classes. The first class is the first-line filter group, which includes different filtering modules performing different filtering methods, such as by reward and scoring. Some of these modules are described below. The MCE-18 module is a unique molecular descriptor filter used to score structures based on medicinal chemistry evolution. The medicinal chemistry filter (MCF) module consists of approximately 460 MCFs and can be used to filter compounds that meet medicinal chemistry requirements, excluding unsuitable compounds containing structural alarms (e.g., reactivity, instability, toxicity, etc.). The Lipinski's Rule of Five (RO5) module performs filtering based on a higher probability of undesirable uptake or permeation when there are more than 5 hydrogen bond donors, 10 hydrogen bond acceptors, a molecular weight greater than 500, and a calculated LogP (CLog P) greater than 5. The T-index module constitutes a set of rules to eliminate structures with unbalanced numbers of carbon atoms and heteroatoms. The similarity scoring module can be used to assess the 2D similarity between the generated structure and the available chemical space. The medicinal chemistry or physicochemical (PC) feature module evaluates compounds using a set of molecular descriptors such as LogP, PSA, HBD, HBA, molecular weight (MW), etc. The drug similarity module filter is estimated using extended rules for drug similarity estimation. Other important filtering modules in this category are based on privileged fragments (PFs). A PF is a substructure that, statistically, is more likely to be molecularly active for a selected target or target class than for other biological targets. Privileged structure models feature automatic prioritization of fragments, which is useful for selected compound or biological target levels. The PF filter for analyzing PFs automatically prioritizes fragments considered essential for a specific level of compound or target. PFs are automatically identified based on the Hierarchical Active Molecules (HAM) dataset. HAMs are a broad set of reported in vitro activities (IC) against a wide range of targets. 50 A dataset of bioactive molecules (<10 μM, ~4 M records) was compiled and integrated into a second-class filter that includes self-organizing maps (SOMs). Comprehensive Accessibility (SA) scores and the Retrosynthetic Related Comprehensive Accessibility (ReRSA) module were used as filters to assess the accessibility of generated structures. ReRSA is an improved fragment-based SA estimation method. The idea of using SA scoring methods is to address the existence of fragments in large databases, but fragment-based methods differ significantly. From an organic synthesis perspective, ReRSA scores are based on more deliberate fragments, which contributes to more accurate SA estimation. A diversity module performs filtering via a main diversity module, which computes for all generated structures using a proprietary FP. A set of rules based on FP was also used to cluster the generated structures using a clustering module.
[0071] The second type of reward function can include many different 3D modules. In the example, the second type can include up to five classes of 3D modules. The first class of 3D modules includes ConfGen, FLEX scoring, and minimization. First, a set of different conformations is generated. Based on X-ray data, a set of rules with predefined substructure geometries is implemented. The FLEX scoring module is used to sort and select structures by stiffness (using a binding entropy factor). The second type of filter contains a set of 3D descriptors designed to provide an assessment of the 3D similarity between the generated structures and the reference molecule.
[0072] The third type of filter relates to pharmacophore hypothesis and search, and includes 3D pharmacophore hypothesis construction with potential binding sites, as well as distance, angle, and tolerance settings, followed by automatic assignment of potential structures and visualization. The fourth type contains filters that use shape similarity as a reference molecule for 3D shape similarity assessment. This includes shape scoring for pharmacophore alignment. The final type of filter focuses on pocket identification and scoring.
[0073] In some embodiments, the metabolic stability enhancer module is used to enhance the metabolic stability of the generated compound. Potential metabolic sites are identified, and these sites are replaced by more stable chemical structural fragments or groups.
[0074] In some embodiments, the 3D descriptor similarity module is used to coarsely assess the 3D similarity between the generated structure and the reference molecule based on a 3D molecular descriptor. This could be a shape similarity assessment used to evaluate the 3D shape similarity with the reference molecule. For pharmacophore arrangements, a shape score exists.
[0075] In some embodiments, the pharmacophore hypothesis module may include 3D construction and modeling of the target protein. The recognition and binding sites of the generated compounds are evaluated. Valuable distances, angles, or tolerances can be set for the generated compounds to filter out unacceptable compounds.
[0076] In some embodiments, the pocket identification and scoring module can be used to scan pockets of biological targets to identify potential binding sites. The scanner can be used for de novo or template-based mesh mapping. Ultrafast binding affinity assessments can be performed on the generated structures. Hydrogen bonding, collision, and repulsion features can be avoided. Forced binding sites can be identified. This module can be validated using X-ray eutectic data. Pocket results can be similar to data obtained from docking studies. Furthermore, volumetric increment scoring can be performed.
[0077] In some embodiments, the technology of the present invention includes the development and application of AI-based algorithms for de novo compound design. AI protocols allow for the comparison and estimation of model performance and the assessment of the properties of generated molecules. Comparing DL models is challenging and requires specific benchmarks and metrics because training parameters and architectural hyperparameters can significantly alter performance.
[0078] In some respects, Molecular Assembly of Elements (MOSES) can be used. MOSES can be designed as a benchmark platform to support deep learning (DL) studies in drug discovery by providing a set of molecular generation models and metrics for evaluating the novelty and quality of generated molecules. These methods are integrated within the AI platform described in this paper. These features allow users to record results, code, training data, and any necessary information to ensure that experiments can be replicated without manual saving. Working within such an environment is important, especially when conducting research experiments that require testing variations of the same model using different training datasets or investigating the relative performance of different models.
[0079] In some embodiments, the protocol described herein can be used to provide indices for estimating the novelty and gradation of generated molecules based on distance metrics in chemical space. The assessment of molecule novelty and gradation is performed using various methods, including automated preprocessing of training data that maximizes novelty and gradation efficiency through generative optimization rather than filtering, and methods that dynamically prioritize novelty and gradation.
[0080] In some embodiments, the protocol described herein can be implemented on a distributed cloud platform with a scalable cloud architecture, such as running on Amazon Web Services (AWS). This implementation integrates various features designed to optimize performance and improve user experience. These include Kubernetes cluster management, a variety of flexible workflows, automated CI / CD, and integrated monitoring and logging. The integration capabilities of the GENTRL platform allow for deployment in the cloud or on-premises. In the case of on-premises deployment, the platform can be easily integrated into existing workflows.
[0081] Figure 10AAn embodiment of a system for implementing a compound generation protocol is shown. A user computer 302 is connected to a compound generation system (CGS) platform 306 via a network 304. The CGS platform 306 can be configured as a small molecule generative chemistry platform that allows exploration of compounds using an automated machine learning platform and access to structure-based and ligand-based drug design, which can be used to generate novel and diverse molecules for protein targets of interest. The user computer 302 uploads user data to an input module 308 of the CGS platform 306. For example, the user data may include structural data of one or more compounds, whether related or different, which can serve as a starting point for compound generation. The user computer 302 uploads information about the desired result to a configuration module 308, which may include a summary of the desired result from the platform (e.g., via a web interface or API). The ideal result could be a molecule that interacts with a specific biological target (e.g., DDR1) or a molecule active for treating a specific disease (e.g., fibrosis). The configuration module 308 then interfaces with a pandomics 310 and obtains information about the relevant proteins (e.g., target identification). Pandomics 310 is a multi-omics target discovery and deep biological analysis engine that can leverage and decipher published omics data, analyze any type of omics data, evaluate drug target proteins, and obtain strategies for drug repurposing. Pandomics provides opportunities to acquire information from disease characteristics to potential targets and candidate molecules. Pandomics combines classic bioinformatics tools, such as signaling pathway analysis using iPANDA. The configuration module 308 can acquire data 312 of existing experimental results, for example, from a database or from the user computer 302. The input module can acquire available structure and building block data 314, for example, from a database or from the user computer 302. The input module 308 then provides the data to the generation module 316, which processes it according to the protocols described herein. The generation module 316 uses AI protocols to analyze chemical structures and substructures based on target proteins to generate suitable structures that can be processed using the medicinal chemistry and computational chemistry units. The AI protocols can provide AI-assisted de novo drug design. The medicinal chemistry and computational chemistry units can include models (training and benchmarking) and provide interpretable domain-specific analysis for generative chemistry tools. The generation module 316 provides one or more generation structures to the result analyzer module 318, which analyzes the results according to the protocol described herein. The result analyzer module 318 may include graphical representations, analysis, and comparisons of the generation structures, and may save information (e.g., SDF CSV) or provide data to a database. The result analyzer module 318 may provide data to the user computer 302 for analysis.For example, the results analyzer module 318 can utilize the platform to analyze the results, or export the results in a standard format from a web interface or via an API. The results analyzer module 318 can also provide new compound candidates 320 selected based on the protocols described herein. These new compound candidates 320 can then be provided to the user computer 302.
[0082] Configuration module 308 can be configured to utilize existing experimental results and available compound structures and substructures (e.g., building blocks). Configuration module 308 can include target identification data and desired properties of the generated compound; it can also use structural data and active compounds. Configuration module 308 can also receive input of generation options for the generated compound. In some aspects, configuration module 308 can be configured to process data from different targets, such as when data is available in the following formats: crystal structure only, crystal and ligand data, ligand data only, no crystal data or no ligand data, or desired properties.
[0083] To obtain the desired results from compounds that interact with target proteins, generation module 316 can be configured to perform generative exploration of chemical space. This can include molecular property optimization, structure optimization, and affinity optimization, all based on the protocols described herein. Molecular property optimization can include optimizing the structure to have desired molecular properties, such as stability, solubility, permeability, absorption, pharmacodynamics, pharmacokinetics, etc. Structure optimization can include optimizing the structure to have characteristics suitable for medicinal chemistry and to be synthesizable, preferably through simple synthesis. Affinity optimization can include optimizing molecular fit and binding to achieve affinity for the target protein. In some aspects, generation module 316 can include numerous models, such as those described herein, which can be provided with the platform.
[0084] In some embodiments, a user can upload their own model via user computer 302, thus enabling the user model to interface with and make requests to generation module 316. The user model can be trained as described herein. The model can then be used alone for compound generation, or in conjunction with other models in the system for compound generation.
[0085] Results Analyzer Module 318 can be configured to perform annotation and virtual screening of generated compounds. This can include screening compounds through a supplier database to determine the presence of the compound or its derivatives, or whether components of the compound are available for purchase for synthesis. Screening can also be targeted at clinical trial analyses, such as insilico clinical trials (e.g., Insilico), which allows for best practices in clinical trials by predicting clinical trial success rates, identifying weaknesses in trial design, and adopting capacity-building techniques to predict clinical trial outcomes. Molecular properties can also be screened to filter compounds with desired molecular properties. Reward-based grading (e.g., SOM) can be used to filter generated compounds during the analysis. The patentability or novelty of generated compounds can be assessed using Results Analyzer Module 318, which can filter molecules that have not been previously generated. Results Analyzer Module 318 can also filter generated compounds through a comprehensive accessibility analysis to determine the ease of synthesizing each specific compound and filter compounds that can be synthesized.
[0086] The results analyzer module 318 can also perform grading and priority sorting of the generated compounds. Grading and priority sorting can be performed as described herein to grade the compounds to identify those to be synthesized and validated. The selected compounds can then be provided (e.g., in a report) for visualization of the results (e.g., on a monitor or printed report) and validation. The selected compounds can be provided to the user's computer 302.
[0087] In some embodiments, the method for obtaining generated compounds can be performed using the CGS platform 306. The method may include a user identifying target proteins or target diseases and defining any output criteria that satisfy the generated compounds. The CGS platform performs computations for compound generation and identifies one or more generated compounds, which are generated, analyzed, and ranked as described herein. Compounds with the highest ranking that satisfy the output criteria and are then identified, along with the identification results and other information, are provided to the user (e.g., via the user's computer). The CGS platform can be configured to operate at the pathway activity level, enabling the processing of high-dimensional data. The CGS platform has an AI-driven toolkit including: a deep feature selection engine for pathway reconstruction, a pathway scoring engine, target association, a deep learning transcriptional response scoring engine, and an activation-based scoring engine. This multimodal approach, combining big data, chemistry, biology, and medicine, allows for a complete description of the molecular structure, properties, changes in biological samples, and interactions between drug responses required for compound generation, where the compounds are biologically active against target proteins (e.g., proteins in biological pathways).
[0088] In some embodiments, the computer-implemented method may include: receiving input of a biological target; receiving a generative tensor reinforcement learning (GENTRL) model or other generative model trained with a reference compound, wherein the reference compound includes: a general compound, a compound regulating a biological target, and a compound regulating biomolecules other than a biological target; generating a structure of the generative compound using the generative model; prioritizing the structures of the generative compound based on at least one criterion; processing the preferred chemical structures of the generative compound using the Sammon mapping protocol to obtain the hit structure; and providing the chemical structure of the hit structure. In some aspects, the method may include: receiving a reference compound; training a generative model with the reference compound. In some aspects, it includes at least one of: refining a structure having at least one reward function using the generative model; or performing pharmacophore modeling using the generative model.
[0089] In some embodiments, meeting at least one criterion for priority ranking is determined to be met by at least one of the following: performing molecular filtering on the generated compounds; performing a clustering / diversifying operation on the generated compounds; analyzing the supplier's compounds based on the generated compounds; performing reward priority ranking; performing root mean square deviation (RMSD) determination to make the generated compounds suitable for biological targets; or analyzing the novelty of the generated compounds based on published intellectual property documents.
[0090] In some embodiments, the method may include performing a structure refinement protocol using at least one Kohonen self-organizing map (SOM). In some aspects, the Kohonen SOM includes: a trend SOM that rewards structures that are newer than older structures based on a chronological timeline; a general biological target SOM that rewards a family of biological targets that includes the biological target but excludes other categories of structures that are not biologically active against the biological target family; and a specific biological target SOM that rewards structures that target a specific biological target.
[0091] In some embodiments, the method may include performing pharmacophore modeling with the generated compound and the biological target to analyze the scaffold structure or side-chain substituent structure of the generated compound. This may include structural analysis and docking analysis with the biological target.
[0092] In some embodiments, the method may include: generating a latent spatial manifold having multiple reference compounds in a trained generative model; and generating new compounds that do not exist in the latent spatial manifold using the trained generative model.
[0093] In some embodiments, the method may include prioritizing the structures of the generated compounds based on at least one criterion, including at least one of the following: filtering compounds to satisfy a predefined range of molecular descriptors; applying a medicinal chemistry filter to remove compounds with undesirable medicinal chemistry properties; applying Tanimoto-based clustering and settling to remove compounds with similar structures; applying a trend SOM, which selects newer structures based on a chronological timeline compared to older structures; applying a general biological target SOM, which selects structures that are biologically active against a family of biological targets including that biological target; applying a specific biological target SOM, which selects structures that target a specific biological target; or applying a pharmacophore filter to remove compounds that fail pharmacophore modeling.
[0094] In some embodiments, the method may include screening for generative compounds that are patentable.
[0095] In some embodiments, the method may include obtaining a dataset having at least one reference compound, comprising: patient data of a set of known compounds having specific functional activities, the known compounds being arranged by the confirmation date of the known compounds; and / or chemical structure data of a set of known compounds having specific functional activities.
[0096] In some embodiments, the method may include identifying an activity threshold that regulates a specific functional activity of a biological target, wherein compounds generated below the activity threshold are defined as inactive compounds, and compounds having an activity threshold or higher are defined as active compounds.
[0097] In some embodiments, the method may include: processing a dataset of reference compounds to exclude anomalous compounds; and normalizing the input chemical space by reducing the number of compounds with similar structures in each compound cluster.
[0098] In some embodiments, the method may include: identifying a plurality of first compounds, the plurality of first compounds being part of a learned manifold of the first compounds; training the structure of the parameterized learned manifold using tensors with partial known properties of the compounds, the partial known properties including known properties of the compounds; and generating a plurality of second compounds based on the plurality of first compounds, wherein the second compounds are generated compounds.
[0099] In some embodiments, the method may include: identifying a predefined range of molecular descriptors; obtaining the output of generated compounds from a generative model; and excluding compounds outside the predefined range to obtain a chemical space with generated compounds within the predefined range.
[0100] In some embodiments, the method may include randomly selecting a plurality of generating compounds from a Sammon map covering the chemical space to obtain a hit structure.
[0101] In some embodiments, the method of synthesizing the compound may include: obtaining a report of a provided chemical structure having a hit structure; selecting at least one compound from the provided chemical structure that has a specific functional activity for the biological target; and synthesizing the at least one compound.
[0102] In some embodiments, the method for verifying the bioactivity of the generated compound may include: obtaining at least one synthetic compound generated by a generative model; and verifying in a bioanalysis that the synthesized at least one compound has a specific functional activity.
[0103] In some embodiments, a method for generating compounds with specific functional activities may include: selecting specific functional activities; obtaining a reference compound; training a model with specific functional activities using the reference compound; generating the structure of the compound; processing the generated structure with a defined reward function; prioritizing a set of generated structures; selecting at least one compound from the generated structures with specific functional activities in a computer; synthesizing the at least one compound; and verifying that the synthesized at least one compound has specific functional activities.
[0104] In some embodiments, the method may include: avoiding mode collapse; utilizing partially labeled data; and inferring the chemical space of a compound from a manifold of a first compound, wherein the inferred chemical space of the compound includes a training library of second compounds and / or compounds that are not part of the manifold of the first compound. In some aspects, mode collapse is avoided by compressing the chemical space of the compounds into a latent distribution. In some aspects, the latent distribution is a Gaussian distribution.
[0105] In some embodiments, the method may include: parameterizing the latent space in a high-dimensional lattice at each node using a prior distribution of the compound with multiple multidimensional Gaussian distributions, wherein the parameterization associates latent codes with compound properties, and wherein the parameterization operates with omitted parameter values without explicitly inputting omitted parameter values.
[0106] In some embodiments, the method may include using a self-organizing map (SOM) to estimate the chemical structure relationship with a defined functional compound having a defined biological activity. In some aspects, the SOM includes a trend SOM, which is a Kohonen-based reward function that distinguishes unknown compounds from defined functional compounds. In some aspects, neurons with multiple unknown compounds are used to actively reward a compound generation model to obtain unknown compounds. In some aspects, the SOM includes a general function SOM, which is a Kohonen map that distinguishes compounds with a general function from those without, to obtain compounds with a general function. In some aspects, the general function is kinase inhibition, and the compound is a kinase inhibitor. In some aspects, unknown compounds with a general function greater than 1.3 are predicted using a Kohonen-based classifier protocol. In some aspects, the SOM includes a specific function SOM, which has been trained to identify compounds with a specific functional activity. In an example, the specific functional activity is DDR1 kinase inhibition. In some aspects, known compounds with known specific functional activities are distributed across multiple associated neurons. In some aspects, the method may include prioritizing the structures of the generated compounds using trend SOMs, general functional SOMs, and specific functional SOMs. In some aspects, the method may include identifying nodes of compounds with specific functional activities and preferentially using the identified nodes.
[0107] In some aspects, the method may include obtaining at least one dataset having: (A) small molecule compounds with specific functional activity; (B) a group of positive compounds with specific functional activity; (C) a group of negative compounds with different functional activities, optionally without specific functional activity; (D) patient data of a group of known compounds with specific functional activity, the known compounds being arranged by the date of identification of the known compounds (e.g., patent application date); or (E) chemical structure data of a group of known compounds with specific functional activity. In some aspects, the method may include identifying an activity threshold for specific functional activity, where compounds below the activity threshold are defined as inactive compounds, and compounds with an activity threshold or higher are defined as active compounds. In some aspects, the method may include processing the datasets of (A)-(D) to exclude anomalous compounds; and normalizing the input chemical space by reducing the number of compounds with similar structures contained in each compound cluster.
[0108] In some embodiments, the method may include: identifying a plurality of first compounds, the plurality of first compounds being part of a learned manifold of the first compounds; training the structure of the parameterized learned manifold with tensors using partially known properties of the compounds, the partially known properties including known properties of the compounds; and generating a plurality of second compounds based on the plurality of first compounds.
[0109] In some embodiments, the method may include training a generative model to generate specific functional activities using a set of compounds with general functional activities (e.g., kinase inhibitors) and a set of compounds with specific functional activities (e.g., DDR1 inhibitors). In some aspects, the method may include pre-training a generative model to generate specific functional activities using a general molecular dataset (e.g., the ZINC database).
[0110] In some embodiments, the method may include: identifying a predefined range of molecular descriptors; obtaining an output of generated molecules from a model; and excluding compounds outside the predefined range to obtain the chemical space of compounds within the predefined range. In some aspects, the method may include reducing the chemical space through clustering and sorting procedures. In some aspects, the method includes prioritizing structures using a universal activity SOM and a specific functional activity SOM. In some aspects, the method includes obtaining a 3D pharmacophore hypothesis from crystallographic data of compounds complexed with proteins (e.g., DDR1 kinase) that exhibit specific functional activity (e.g., inhibition of DDR1 kinase).
[0111] In some embodiments, the method includes prioritizing compounds by at least one of the following: Sammon mapping; and root mean square deviation values of 3D pharmacophore hypotheses. In some aspects, Sammon mapping is used to randomly select multiple chemical structures covering their chemical space.
[0112] Example
[0113] An example protocol is now described. A dataset of biological targets and known inhibitors of related targets (e.g., kinase inhibitors and DDR1 inhibitors) is used to fine-tune a generative GENTRL model, which is pre-trained on a large molecular dataset available in the ZINC database. The initial output of GENTRL is 30,000 small molecular structures as a “coarse mixture.” Compounds that do not meet the predefined range of the molecular descriptor are excluded. To remove molecules with structural alarms or active groups, the protocol applies a medicinal chemistry filter (MCF) containing more than 100 substructure queries. The resulting chemical space is then reduced through clustering and diversity sorting procedures.
[0114] The protocol then performs additional prioritization of structures for potential biotarget inhibitory (e.g., kinase inhibitory) activity and activity against specific biotargets (e.g., DDR1 kinase) using constructed general biotarget family SOMs (e.g., general kinase SOMs) and specific biotarget SOMs (e.g., specific kinase SOMs), respectively. While the generative model focuses on generating structures that are a priori classified as specific biotarget inhibitors (e.g., DDR1 inhibitors) via the aforementioned Kohonen-based reward function, the protocol also uses pharmacophore modeling to evaluate the generated structures. 3D pharmacophore hypotheses are constructed using available crystallographic data of small molecule compounds co-conjugated with biotargets (e.g., DDR1 kinases). These models are applied to score the obtained structures using root mean square deviation (RMSD, Å), which reflects the degree of fit to the developed pharmacophore (see [link to relevant documentation]). Figure 2A-2C (Example in the text).
[0115] Figure 2A The three-center pharmacophore hypothesis was demonstrated: Acc-hydrogen bond acceptor (r=2Å), Hyd|Aro-hydrophobic or aromatic center (r=2Å), and Hyd-hydrophobic center (r=2Å). Figure 2B The four-center pharmacophore hypothesis is shown: Acc-hydrogen bond acceptor (r=2Å), Hyd|Aro-hydrophobic or aromatic center (r=2Å), Hyd-hydrophobic center (r=2Å), Acc|specific-hydrogen bond acceptor or a fragment with similar spatial geometry (e.g., double or triple bonds, planar loops) (r=1.7Å). The distances not depicted are the same as those for a three-center pharmacophore. Figure 2C This demonstrates the 5-center pharmacophore hypothesis, which includes Figure 2B The similarities are highlighted, with additional hydrophobic features. The undescribed distances are the same as the distances between the 3-center and 4-center pharmacophores. This is based on a reported small molecule DDR1 inhibitor (PDB code: 5BVN).
[0116] In the final step of prioritization, the protocol uses the Sammon mapping method, which uses the same set of descriptors and pharmacophore modeling outputs to calculate RMSD values. When constructing the mapping, the protocol randomly selects a subset (e.g., 40) of structures that smoothly cover the obtained chemical space. Special attention is paid to the distribution of the RMSD values (see...). Figure 3 ). Figure 3 The nonlinear Sammon map of 40 selected molecules, marked with triangles, is displayed. The region of best pharmacophore matching is highlighted with a circle.
[0117] The execution of all priority-ordered procedures is summarized below. The physicochemical filter applies to compounds that meet a predefined range of molecular descriptors (e.g., the 12,147 selected compounds). The MCF filter applies to compounds containing alarm structures that are generally undesirable in medicinal chemistry (e.g., the 7,912 selected compounds). The clustering and diversity filter is based on Tanimoto-based clustering, then diversity is maximized for each cluster, removing similar compounds (≤5 compounds in a cluster) and chemically spatially normalized (e.g., the 5,542 selected compounds). The similarity filter is based on Tanimoto-based similarity to compounds available in the supplier's (MolPort, ZINC) inventory (threshold ≤0.5) (e.g., the 4,642 selected compounds). The universal biotarget family SOM (e.g., universal kinase SOM) filter categorizes structures into biotarget family inhibitors and biotarget family non-inhibitors (e.g., the 2,570 selected compounds). The specific biotarget SOM (e.g., DDR1 kinase) filter selects structures from neurons containing at least one specific target inhibitor (e.g., the 1,951 selected compounds) to overcome bias. The pharmacophore search filter is used for structures that successfully pass pharmacophore modeling (e.g., the 848 selected compounds). The Sammon mapping filter is used for compounds undergoing the Sammon learning process to randomly select the final set of structures (e.g., the 40 selected compounds).
[0118] In some embodiments, the protocol may include analysis of structures by chemists to facilitate synthesis. The synthesis of selected compounds, as well as cell-based analysis, can be performed.
[0119] In this example, six molecules were generated, selected, and synthesized, and submitted for biological testing (e.g., within 35 days). In the test sample, four compounds exhibited moderate to high activity (see [link to test sample]). Figure 4A (dose-response curves in the data). Compound 1 (e.g., INS015_036) and compound 2 (e.g., INS015_37) showed strong inhibition of DDR1 activity, IC50. 50 The values were 10 and 21 nM, respectively. Compound 3 (e.g., INS015_030) and compound 4 (e.g., INS015_032) showed moderate potency (1 μM and 278 nM, respectively), while compounds 5 (e.g., INS015_039) and 6 (e.g., INS015_038) showed no activity. Since the dose-response curves for compounds 2 and 4 appeared likely to be inconsistent, additional experiments were performed to confirm the activity of these molecules against DDR1 kinase (see [link to relevant documentation]). Figure 4B In these studies, compound 2 exhibited IC50. 50 The value was 37 nM, while compound 4 had 4 times lower activity (IC50).50 =156 nM). Therefore, nanomolar activities of both compounds were demonstrated in two different biochemical analyses. Compounds 1 and 2 were also evaluated against DDR2 kinase. Figure 4A The results described show that, for DDR1, compound 1 exhibits 23 times weaker activity, while compound 2 shows IC... 50 The value is 76 nM. Based on these results, it can be concluded that the two most active DDR1 inhibitors (compounds 1 and 2) are valuable structures for further research and optimization.
[0120] Figures 4A-4B The structures and dose-response curves of the generated molecules are shown. The six generated compounds were tested against DDR1 tyrosine kinase in a dose-dependent manner. Compounds 1 and 2 exhibited IC50 values in the low nanomolar range. 50 value( Figure 4A Additionally, compounds 2 and 4 were rescreened against DDR1 kinase using another biochemical assay (Thermo Fisher-PR6913A), showing IC50 values for each. 50 The values are 37.12 and 155.6 nM ( Figure 4B ). Figure 4C The structures of compounds 1-6 and ICs for DDR1 and DDR2 are shown. 50 Such as by their compound numbers.
[0121] Quantum mechanical (QM) calculations were performed to partially explain the results obtained from biological research. Figures 5A-5C The optimal conformation for the pharmacophore hypothesis and the rigid arrangement of the conformations predicted by QM calculations are shown. Using ab initio QM calculations based on the pharmacophore hypothesis, the predicted 3D conformation of compound 1 is very similar to the conformation verified in vacuum, and is more preferred and more stable. Figure 5A For compound 1, there may be a "lock and key" entropy-driven binding mechanism. For compound 3, a moderate RMSD value was observed. Figure 5B The activity of benzimidazole compound 4 is significantly lower than that of its close-structure analog compound 1 (not shown); however, QM calculations show that its conformation with an amino group at the para-methyl position is more stable and does not match the hypothesized hydrogen bond acceptor (HBA) characteristic. Furthermore, at a physiologically relevant pH of 7.4, the second basic nitrogen in the core of compound 4 is almost completely protonated.
[0122] In the 3D conformation pool of compound 5 verified by QM, none were found to conform to the pharmacophore hypothesis. Figure 5CCompound 6 did not exhibit any activity, possibly due to the large hydrophobic portion containing the chiral site. The trifluoromethyl fragment may be unsuitable for its position within the binding site. Furthermore, it can be inferred that the ortho-amide group is not located in a favorable binding position.
[0123] To overcome some of the problems of compounds that are unfavorable or do not meet certain criteria, the agreement may include obtaining a list of expert opinions on the selected structure from a professional medicinal chemistry team.
[0124] For compound 1, human experts noted its similarity to dasatinib (IC). 50 =9nM, DDR1 inhibitor). Compound 1 is a bioisosteric compound with an amide bond. From an AI perspective, compound 1 is very impressive. Compound 2 has a unique central linker (imidazolidine, whose chemical properties are very different from compound 1). Human experts noted that the stability of compound 2 should be examined. For compound 3, some issues regarding the potential metabolic stability of compounds (acetylene, pyrimidine, and dimethylamino) have been addressed. The tertiary amine of compound 3 is also problematic in terms of the activity of the DDR1 inhibitor and the observed SAR. For compound 4, human experts have found the chemical type presented to be fairly neutral or attractive, and patentability is considered a challenging issue; however, a recent patent search revealed a clear IP status. For compound 5, human experts noted that it is a chemical type of interest in kinase chemistry, with good physicochemical properties, especially solubility. According to expert opinion, compound 6 has a novel structure and contains Interesting and unique components, such as a unique hinge-binding core. According to human experts, almost all the compounds tested were considered novel and attractive for further biological testing. Some compounds have been classified as having unique structures. On the other hand, features identified by experts that require more work in further drug development, including metabolic instability, relatively poor overall accessibility, and the potential need for solubility modulation, have been aptly attributed to certain components. These are features that can be added to more sophisticated screening protocols or integrated into a second round of refinement. Therefore, the protocols for generating compounds may include steps involving human experts in chemistry analyzing the structure of compounds to facilitate the selection of lead compounds or compounds to be synthesized and validated for biological activity.
[0125] In some embodiments, the patentability of compounds can be analyzed using the platform. Notably, almost all selected compounds are considered to have good intellectual property status, as shown in Table 1. The protocol can perform preliminary scoring on these compounds using specific databases available for chemical patents. Theoretically, 39 structures are classified as novel structures because they fall outside the scope of any published patents or applications. A patent application describing a multi-kinase inhibitor claims protection for one of the generated structures, but the kinases listed therein do not include DDR1.
[0126] Table 1 – IP Summary
[0127]
[0128] * Perform a similarity search in the SciFinder database. A range of structural similarity is provided. The number of similar compounds within this range is enclosed in parentheses. If no molecules have a similarity >70%, it is displayed as <70%.
[0129] **Number of patents for compound-matched Markush structures.
[0130] ***The compounds present in the examples in the patent.
[0131] ****The generated compound and the molecules from the MolPort and ZINC databases show the greatest similarity.
[0132] Therefore, the generated compounds can be screened for patentability through protocols. This facilitates compound generation, thus avoiding the synthesis and verification of non-patentable compounds. This allows for a focus on generating commercially viable compounds.
[0133] In some embodiments, the protocols described herein can lead to examples of structures that are nontrivial potential bioisosteric substitutions and topological modifications of compounds. Figure 6 ). Figure 6 Representative examples of the resulting structures compared to the parent DDR1 inhibitor are shown. Overall, this computational deformation preserves well the inherent physicochemical properties of known DDR1 inhibitors and retains the key binding sites responsible for DDR1 affinity.
[0134] In some embodiments, the workflow protocol can lead to the generation of compounds with high selectivity. Selectivity issues are crucial for lead compound evaluation and can influence the potential deviation from the target that could affect the success of preclinical assessment. The protocol can be configured to assess the selectivity index (SI) of the most active generated compounds, for example, in enzyme assays, using the scanMAX Delta kinase panel via Eurofins. Compound 1, tested at a concentration of 10 µM, showed relatively high SIs for over 44 kinases, including serine / threonine protein kinases (e.g., CDK, PKCβ2, MAPKAPK3, TSSK, TTBK1, A-Raf, etc.), lipid and atypical kinases, and bispecific protein kinases, such as… Figure 7 As shown in the panel, the highest inhibitory potency against eEF-2K (INH%=37) is observed, while DDR1 kinase activity is completely inhibited at this concentration. DDR1 kinase is primarily expressed in epithelial cells, while DDR2 expression is typically observed in testicular interstitial cells. Although fibrosis prevention is the primary goal, the selectivity of DDR2 kinase is also a very important consideration. The inhibitory activity of molecular compounds 1 and 2 against the DDR2 kinase isoform was also evaluated. Compounds 1 and 2 were found to have good and moderate SI values: 23.4 and 3.6, respectively. Subsequent optimization of compound 1 through the synthesis of related structural analogs may lead to improved selectivity. Detailed selectivity curves are available in [the following section / section / etc.]. Figure 7 Presented in the middle.
[0135] In collagen-stimulated U2OS cells, the ability of compounds 1 and 2 to inhibit DDR1 autophosphorylation was investigated. The amount of activated DRR1 (Y543) was measured using Western blotting analysis, and the obtained data were normalized to HA and GAPDH protein levels. Dasatinib was used as a positive control and showed high potency, with an IC50 value of [missing value]. 50 The value was 1 nM. Compounds 1 and 2 were found to significantly block the autophosphorylation of DDR1 in a dose-dependent manner, IC50. 50 The values are 10.3 and 5.8 nM respectively. Figure 8A-8I These values are close to the activities observed in biochemical tests of the two compounds.
[0136] The antifibrotic activity of compounds 1 and 2 was evaluated using the MRC-5 cell line, such as... Figures 9A-9F As shown. Compared with two reference compounds, SB-525334. 44 (TGFBR1 inhibitor, IC50) 50 =5-15 nM) and dasatinib 45 (Non-selective kinase inhibitors, including DDR1 and DDR2, IC) 50Compared to ~15-30 nM, four antibodies were used to assess the antifibrotic activity of the selected molecules, two reference compounds were used as positive controls, and DMSO (0.1%) was used as a negative control. The obtained Western blot results are as follows: Figure 9A As described in [the text]. Figures 9B-9C As shown, dimethyl sulfoxide (DMSO) showed no effect in the test system; however, the addition of TGF-β (10 ng) resulted in enhanced α-actin expression (up to 9.3-fold) and enhanced CTGF expression (2-fold). For SB-525334, the observed dose-dependent effect is unclear, with maximum inhibition reaching near the negative control baseline at 10 µM (for dasatinib), while at 0.5 µM, we observed a 2-fold reduction in α-actin expression compared to TGF+DMSO stimulation. Conversely, dasatinib showed a significant stimulatory effect at 0.5 µM and higher concentrations. SB-525334 (10 µM) slightly reduced CTGF expression (by 1.5-fold), while at 0.5 µM it had no effect. Figure 9E-9F At all concentrations used, dasatinib showed only stimulatory effects, with no indication of inhibition. For compound 1, the maximum inhibitory potency against α-actin expression was reached at 10 µM, close to that determined for SB-525334 and dasatinib, while compound 2 showed lower activity. The maximum effect (1.5-fold) was observed at 0.37 µM. Compound 1 exhibited a strong dose-dependent effect compared to other molecules. In the CTGF assay, compound 1 showed no inhibitory potency at all concentrations used; however, at 0.013 µM, it showed antifibrotic activity equivalent to that determined for SB-525334 at a higher concentration (38-fold). Compound 2 showed near-maximal activity close to the negative control at a concentration of 0.041 µM, and it was more active than both SB-525334 and dasatinib.
[0137] In addition to the pulmonary fibrosis model, the antifibrotic effects of inhibitory compounds 1 and 2 were also investigated in the human hepatic stellate cell line LX-2. Western blot analysis was used to track collagen α1, α-SMA, CTGF, and GAPDH. DMSO and SB-525334 were used as negative and positive controls, respectively. Data not shown indicate the effects of compounds 1 and 2 on cellular fibrosis markers collagen α1, α-actin, and CTGF (normalized to GAPDH) in LX-2 cells.
[0138] Treatment with TGF-β induced the production of collagen α1, α-SMA, and CTGF in LX-2 cells. SB-525334 strongly inhibited the expression of collagen α1 and α-actin in the concentration range of 0.5–10 μM; however, at lower concentrations, we observed a significant decrease in activity. Data normalized to GAPDH levels clearly showed that compound 1 strongly inhibited collagen production in TGF-β-stimulated LX-2 cells in a dose-dependent manner, with an IC50 concentration of [missing data]. 50 The value was 13 nM. At a concentration of 41 nM, the highest inhibition of α-actin production was observed; however, compound 1 did not show inhibitory activity in the CTGF assay. Considering the nanomolar efficacy of the molecule in enzymatic, autophosphorylation, and fibrosis assays, the consistent conversion from biochemistry to cellular activity is clear and well-defined. Notably, the IC50 of compound 1 in the LX-2 assay was [not specified in the original text]. 50 The values significantly exceeded the cytotoxicity against the same cell line (CC50 = 3.3 μM). In collagen assays, compound 2 was found to have micromolar activity (IC50 = 3.3 μM). 50 >10 μM). At lower concentrations ranging from 3.3 to 0.014 μM, compound 2 did not inhibit the production of collagen α1. At a concentration of 14 nM, it blocked almost half of CTGF production (43%) and slightly inhibited α-SMA (15%). However, these effects diminished with increasing concentration. Compound 2 exhibited low cytotoxicity in LX-2 cells, with a CC50 value of 7.3 μM. Based on these preliminary results, it can be tentatively concluded that the new compound possesses good antifibrotic activity.
[0139] Because compound 2 contains an imidazolidine fragment, which is not commonly found in drug discovery, we experimentally evaluated the key properties of this molecule. Thus, compound 2 exhibited a kinetic solubility of 1.09 µg / ml at pH 7.4, while its thermodynamic solubility was <0.59 µg / ml, with a logD of 4.07 (TFA salt) and pKa of 6.99. The inhibitory activity of compounds 1 and 2 against a small fragment of a key cytochrome P450 (CYP450) isoform was also evaluated in vitro (Table 2).
[0140] Table 2 – IC50
[0141]
[0142] Compound 1 was found to inhibit the activity of CYP1A2, IC50. 50 The value is 7.36 µM; however, it is inactive against CYP2C9, CYP2C19, CYP2D6, and CYP3A4 (IC50). 50Compound 2 exhibited high inhibitory activity (>50µM), IC50. 50 The values were 10.6, 2.70, 6.56, 6.97, and 7.36 µM, respectively. Neither compound showed significant CYP450 inhibition. Their CYP450 activities were well below the target nanomolar potency, thus providing a good selectivity index comparable to that of many drugs, including kinase inhibitors. Detailed descriptions of the performed assays are provided in the supporting information. The metabolic stability of compound 2 (10 µM) was evaluated in liver microsomes of humans, SD rats, CD1 mice, and beagle dogs (Table 3).
[0143] Table 3 – Summary of microparticle stability results for compound 2
[0144]
[0145] *NCF: Abbreviation for No Cofactor. During a 60-minute incubation period, without adding the NADPH regeneration system to the NCF sample (using buffer instead), if the NCF residue is less than 60%, then NADPH-independent NCF occurs.
[0146] R 2 These are the correlation coefficients of the linear regression that determine the kinetic constants (see the original data worksheet).
[0147] T 1 / 2 It has a half-life, and CL int(微粒体) Internal clearance rate
[0148] C Lint(微粒体) =0.693 / half-life / mg microsomal protein per mL
[0149] C Lint(肝脏) =C Lint(微粒体) *mg microsomal protein / g liver weight*g liver weight / kg body weight mg microsomal protein / g liver weight: 45mg / g for 5 species
[0150] Liver weight: 88 g / kg for mice, 40 g / kg for rats, 32 g / kg for dogs, 30 g / kg for monkeys and 20 g / kg for humans.
[0151] In human microsomes, compound 2 exhibits a half-life of 12.8 min (t 1 / 2 The intrinsic clearance values were 97.3 mL / min / kg (human liver weight 20 g / kg, Clint / liver). For example, under the same conditions, testosterone, diclofenac, and propafenone showed the following values: t1 / 2 = 15.6, 10.7, 8.3 min, and C, respectively.lint / 肝脏 =79.7, 116.9, 149.7 mg / min / kg. Based on the results obtained, it can be concluded that compound 2 exhibits relatively good metabolic stability compared to the control molecule. It should be noted in particular that the metabolic reactions of all tested compounds were only performed in the presence of an NADPH regeneration system added to the samples (2.8% of compound 2 remained after 60 min of incubation); however, in the absence of NADPH, based on LC / MS / MS data, we observed that 88.7% of compound 2 was not modified. The residual amounts of testosterone, diclofenac, and propafenone were 6.9%, 1.9%, and 0.7%, respectively. This clearly demonstrates that compound 2 is quite stable under the experimental conditions. A detailed summary of the metabolic stability of compound 2 and the control molecule is presented in the supporting information. In addition, we evaluated the stability of compound 2 (10 µM) in phosphate buffer (50 mM, pH=7.4) and MOPS / EDTA (8 mM / 0.2 mM, pH=7.0). The samples were incubated for 0, 120, 240, 360, and 1440 min, and then immediately analyzed by LC / MS / MS. Therefore, under the experimental conditions, compound 2 was very stable (the residual amount of the compound was close to 100% at each time point, see Table 4).
[0152] Table 4 – Buffer stability results for compound 2
[0153]
[0154] Furthermore, the binding interactions of compound 1 in the target kinase were analyzed using DDR1 crystal structure (PDB code: 3ZOS) via molecular docking. The hypothesized binding mode revealed several features characterizing the type II inhibitory mechanism of protein kinase (data not shown). The molecular docking procedure was performed using Schrodinger Maestro. A conserved hinge interaction was found between the N1 of the imidazopyridazine scaffold and Met704. The C(2)H of the scaffold also binds to the backbone of Asp702 to form a pseudo-hydrogen bond. The 6-methyl-benzisoxazole moiety exhibits an orthogonal geometry to the hinge element via the ethynyl linker, which is the conformation required for the DFG-out pocket. The strong connection is established through the hydrogen bonding of isoxazole to the Asp784 of the DFG motif and the interaction of the π-cation with the catalytically active Lys655. Occupying a hydrophobic pocket opened via the DFE motif, the CF3 phenyl group of compound 1 further stabilizes in a complex via close contact with Ile675, Met676, Leu679, Ile684, Ile685, Leu757, and Ile782. The exocyclic amine hydrogen bond formed with Glu672 is part of a broad network of hydrogen bonds and / or charges composed of Lys655, Glu672, and Asp784. In summary, compound 1 forms multiple hydrogen bonds, favorable charge, and hydrophobic interactions with the active site residues of DDR1 kinase. The significant complementarity of compound 1 with the ATP site prerequisite confirms its strong inhibitory activity against DDR1.
[0155] This workflow protocol has demonstrated that deep generative networks with reinforcement learning can be used to generate novel active molecules with a fairly high hit rate. The complexity of achieving a balance of desired activity, specificity, selectivity, solubility, bioavailability, overall accessibility, and many other properties requires target- and therapeutic region-specific modeling, which can be performed using the workflow protocol described in this paper. Generative models can provide hit compounds with a higher level of complexity and novelty.
[0156] method
[0157] Refer to DDR1 small molecule inhibitors.
[0158] Using available databases, scientific publications, and patent records, a dataset of compounds with reported anti-DDR1 kinase activity was compiled: 63 molecules were obtained from Thomson Reuters Integrity. 49 and ChemBL 50 The database received 77 compounds from literature and 1230 structures manually collected from patent records. In total, the final dataset includes 1370 compounds.
[0159] pre-trained dataset
[0160] For the pre-training procedure, we have used data from the Zinc database. 51 The structure dataset was compiled using a clean navigation setup and proprietary databases from our partners. Structures containing unwanted atoms other than C, N, O, S, F, Cl, Br, and H were removed. Conventional medicinal chemistry filters were applied to exclude compounds with potentially toxic and reactive groups. The resulting dataset contains approximately 1.9 million structures in the form of canonical SMILES. The dataset was parsed to collect a vocabulary of 34 unique tokens, such as atomic symbols, brackets, and other SMILES-specific grammatical elements. The average string length is 36 tokens, and the maximum length is 58 tokens.
[0161] Kinase inhibitors and “negative” datasets
[0162] Using available data from the Integrity database, a dataset of active and inactive molecules against various kinases was compiled. Based on chemical hierarchies, the unique structures of over 56,000 kinase inhibitors were collected and analyzed. Standard clustering and diversity sorting procedures were used. 52 The dataset was reduced to a size of 23K unique structures. A second part of the dataset, based on the ChEMBL database, was compiled, containing 17,000 compounds with highly hierarchical structures and activity against other targets.
[0163] Compounds from the patent record (based on the priority date)
[0164] We used the Integrity database to collect data on large pharmaceutical companies (the top ten pharmaceutical companies in 2017). 53 A dataset of structures claimed as new pharmaceutical substances in patent records since 1950. Priority dates were assigned to all included compounds. To collect only unique records, a filtering procedure was performed, resulting in 22,000 compounds. The structures underwent the following preparation processes: salt removal, error correction, filtering (non-pharmaceutical elements, isotopes), clustering and cluster normalization, outlier removal, and duplicate removal. The final dataset contains 17,000 records.
[0165] Model
[0166] At the heart of our generator pipeline is GENTRL, a variational autoencoder with a rich prior distribution in the latent space. Figure 11The GENTRL model is a generative tensor reinforcement learning model. It comprises a learning and generation space 200 tailored to a specific biological target and a generation policy space 208. The GENTRL model includes a chemical dataset 202, which may include the types of compound groups described herein. The chemical dataset 202 is input (e.g., as an input vector) to an encoder 204 and processed to generate the compounds in a latent space 206. The latent space is modulated by a medicinal chemistry filter (MCF) 210, a medicinal chemistry evolution (MCE) 212, and pIC50 214. The latent space 206 utilizes generation policies through a generator 215 that generates compounds in the latent space 218, which are then filtered through a reward 220, such as a SOM reward (e.g., trend SOM, general SOM, and specific SOM). The latent space 206 outputs the results to a generator 216 to generate the generated compounds 222 in the chemical space. These generated compounds can then be synthesized, and their biological activity against the target can be verified. Therefore, the training procedure consists of two parts. In the first phase, we train conditional models to learn the relationship between molecular structure and properties. In the second phase, we explore the chemical space to discover promising molecules with high rewards.
[0167] The GENTRL model uses tensor decomposition to encode the relationship between molecular structure and its properties. The GENTRL model is trained in a semi-supervised manner, requiring no missing values. The code for the GENTRL model is available at github.com / insilicomedicine / gentrl.
[0168] Tensor Training Decomposition 54 High-dimensional tensors are approximated using a relatively small number of parameters. Discrete random variables Joint distribution of -1} It can be represented as n Elements of a dimensional tensor:
[0169] ,
[0170] Tensor It's the core, 1 m It is a vector of binary one's complement, and Z It is a standardization constant. With a larger core size, we can better approximate the distribution, although the number of parameters increases with the core size. m And it grows quadratically. In tensor training, we can effectively marginalize the distribution with respect to any variable:
[0171] ,
[0172] in It can be efficiently computed. For marginal distributions, we can compute the conditional distribution and sample using chain rule. Standardization constant. Z It is given by the following formula:
[0173] .
[0174] Since the generator autoencodes using continuous latent codes, we train the representation using continuous tensors. To simplify the notation, we assume the latent code... Z It is continuous, and its properties are... y It is continuous. We approximate the distribution. Gaussian distribution and component index A mixture. Regarding z and y The joint distribution is:
[0175] .
[0176] For conditional distribution We choose a fully factorized Gaussian distribution, which does not depend on y :
[0177] .
[0178] distributed The adjustable parameters are the core of tensor training. The average value of the Gaussian component and variance We store tensors in the form of tensor training. The resulting distribution becomes:
[0179] ,
[0180] Our model has a prior distribution. encoder and decoder A variational autoencoder. Consider training examples. ,in It is a molecule, and These are its known properties. The lower bound of evidence for our model is:
[0181] .
[0182] Since molecules determine their properties, we assume We also assume This indicates that the target is entirely defined by its underlying code. The lower bound of the resulting evidence is:
[0183]
[0184] .
[0185] in For the proposed joint distribution We can, based on the observed properties Analytically compute the posterior distribution with respect to the latent code.
[0186] By maximizing the lower bound of evidence, we train, autoencode, and prior on the aforementioned dataset: we sample molecules and their properties (e.g., conditions) from the dataset in SMILES format, including MCE-18 and pIC. 50 (IC) 50 The model uses the negative logarithm of the chemical space and binary features indicating whether a molecule passes through MCFs. We trained this model and obtained a mapping from chemical space to the latent code. This mapping is aware of the relationship between the molecule and its physicochemical properties.
[0187] In the next stage of training, we fine-tuned the model for inhibitors targeting specific biological targets, such as DDR1 kinase inhibitors. We used reinforcement learning (RL) to extend the latent manifold to novel inhibitors with reward functions described in the next section, such as the universal kinase SOM, specific kinase SOM, and trend SOM. We used REINFORCE. 55 Algorithms (also known as logarithmic derivative tricks) are used to directly optimize the model:
[0188] .
[0189] .
[0190] We reduced the variance of the gradient using a standard variance reduction technique called "baseline"—after calculating the reward for all molecules in the batch, we subtracted the average reward for the batch from each reward:
[0191] .
[0192] To preserve the mapping of the chemical space, we fixed the parameters of the encoder and decoder and trained only on the manifold distribution. We combined exploration and exploitation methods. For exploration, we sampled from outside the potential space currently being explored. ,in and It is the average of all dimensions and the sum The variance. For newly discovered areas, if the reward... If it is high, then the underlying manifold will extend toward it. Figure 11 Generation strategy 208).
[0193] Comparison of generative chemical models is crucial for the development of this emerging field, and several benchmark platforms are being developed. 56,57 We successfully compared the performance of GENTRL with previous methods (such as ORGAN). 38、39 RANC 18 and ATNC 17 The comparison was made, and training details are provided in the supporting information.
[0194] The proposed pipeline performs well for the generation of biological target inhibitors (e.g., DDR1 inhibitors). In some respects, the workflow protocol can perform appropriate hyperparameter tuning and exploration, and configure GENTRL to make it less prone to mode collapse. This can be achieved by tracking... The exploration speed is achieved by controlling the entropy, and if the entropy decreases too quickly, the exploration speed is increased. Furthermore, if some promising molecules are far from the exploration region... The model will never discover them. Using specific multi-objective reinforcement learning techniques to balance the performance of all rewards is also helpful. Finally, for some tasks, we found it beneficial to train the decoder together with the latent space during reinforcement learning training.
[0195] reward function
[0196] Based on Kohonen SOM ( Figure 12 ), and developed a reward function. Figure 12 A smooth representation of universal kinases and trend SOMs is shown. The algorithm was developed using Teuvo Kohone. 58 Introduced as a unique unsupervised machine learning dimensionality reduction technique, it can effectively reproduce the inherent topology and patterns hidden in the input chemical space in a faithful and unbiased manner. The input chemical space is typically described by molecular descriptors (input vectors); however, at the output, a two-dimensional or three-dimensional feature map is generated for visual inspection. A set of three SOMs is used as the reward function: the first SOM is trained to predict compound-resistant kinases (a universal kinase SOM, R0). 通用 The activity of ) (e.g., the universal biological target family SOM), a second SOM is developed to select compounds located in neurons associated with DDR1 inhibitors throughout the kinase map (specific kinase SOM, R 特定 (e.g., specific biological target SOM), and the last SOM is based on current trends in medicinal chemistry (trend SOM, R). 趋势 The training process assesses the novelty of chemical structures. During the learning process, generating models are rewarded when the generated structures are classified as molecules that act on kinases, and these models are localized to neurons attributed to DDR1 inhibitors, tending towards those that are relatively novel.
[0197] Trend SOM
[0198] As an additional reward function, we also used the Kohonen map, which was developed based on a training dataset of molecules from the top ten pharmaceutical companies, requiring priority dates for allocation in the intellectual property records. The following key molecular descriptors were selected to train the procedure: MW, LogP (lipophilicity, allocation coefficient calculated in the octanol-1 / water system), and LogS. w (Water-soluble), PSA (polar surface area, Å) 2 HBA, HBD, SS, and MCE-18 (Medicinal Chemistry Evolution). The MCE-18 function is closely related to chemical evolution and reflects the true nonplanarity of molecules and the simplicity and false positive sp. 3 - Rate. MCE-18 is sufficiently sensitive to techniques observed in modern medicinal chemistry based on novelty. Map size is 15×15 2D representation (random distribution threshold of 75 molecules per neuron), number of learning iterations: 2000, initial learning rate: 0.4 (linear decay), initial learning radius: 10, winning neuron determined using Euclidean distance. Initial weight coefficients are generated using random distribution. After the training process is complete, regions containing compounds claimed in different time periods are displayed ( Figure 12 a) in the example. Figure 12 As shown in Figure a, molecules described in relatively novel patent records (claimed between 2015 and 2018) are located in separate regions of the map, while the “Vecchio” structure is located primarily in distinct regions, providing statistically relevant separation. Within the map, we clearly observe intrinsic trends over many years, which can be described as a set of simple vectors. Neurons attributable to novel structures (the past decade) are used to reward our AI core, compared to neurons associated with “ancient” chemotypes.
[0199] Universal kinase SOM
[0200] The preprocessed training dataset comprised a total of 41,000 small molecule compounds: 24,000 kinase inhibitors and 17,000 molecules with reported anti-non-kinase target activity (concentration <1 µM). For the entire database, RDKit was used. 59 Mordred library 60 and SmartMining software 61,62More than 2,000 molecular descriptors were computed. The descriptors were ranked according to the t-values of bivariate students, and then nine descriptors were selected as the highest priority and theoretically most effective descriptors for distinguishing between kinase and non-kinase chemistry. The dataset includes: MW (molecular weight, t=-63.4), Q' (binary quadratic exponent, t=77.3), SS (ordinary electrotopological exponent, t=-69.3), S[>C<] (partial SS exponent, t=-50.3), 1Ka (first Kier topological exponent, t=-66.5), Hy (hydrophilicity exponent, t=-55.9), VDW. w Volume (weighted atom van der Waals volume, t=-70), HBA (number of hydrogen bond acceptors, t=-34.0), HBD (number of hydrogen bond donors, t=-8.5), and RGyr (radius of rotation, t=-55). The map size was 15×15 2D representation (random distribution threshold of 177 molecules per neuron), learning iterations: 2000, initial learning rate: 0.3 (linear decay), initial learning radius: 8, winning neurons were determined using Euclidean metrics, and initial weight coefficients were randomly distributed. After the training process was completed, the regions occupied by kinase inhibitors and molecules acting on other biological targets were highlighted. We observed that these two classes of compounds were mainly located in different regions within the map ( Figure 12 (bd in the middle). Then, neurons are prioritized based on the following privilege factor (PF): (%),in It is the percentage of kinase inhibitors located in the i-th neuron, however This represents the percentage of other molecules located within the same neuron, and vice versa. PF values greater than 1.3 are used as a threshold to assign neurons to one of the two neuron classes. No "dead" neurons are checked within the mapping. Using random thresholds, the average classification accuracy is 84%. Using this model, all generated structures are scored during the learning cycle and optimization step. Compounds classified as kinase inhibitors with the smallest error (Euclidean metric) are then assigned a specific kinase SOM.
[0201] Specific kinase SOM
[0202] A similar procedure was performed to construct the Kohonen map for identifying nodes in the kinase inhibitor chemical space assigned to DDR1 inhibitors. Structures classified as kinase inhibitors via the universal kinase SOM (see above) were used as the input chemical pool. The final set of molecular descriptors included: MW ( t =-44), Q' ( t =37), 1Ka ( t =-42), SS ( t =-52), Hy ( t =-30), VDW总 volume( t =-35), HBA ( t =-40) and HBD ( t =14). The mapping size is 10×10 2D representation (randomly distributed threshold of 245 molecules per neuron), and the learning settings are the same as those applied to the universal kinase SOM, except that the initial learning radius is equal to 6. After constructing the mapping, we identified neurons containing at least one DDR1 inhibitor. The formal average classification accuracy was 68%, and a bias towards DDR1 inhibitors was observed. Then, during the learning process, "active" neurons were used to select structures to reward the core GENTRL and to optimize the process. In this case, we did not use PF for selection decisions to overcome overtraining and enhance the novelty of generated structures operating near terrain.
[0203] pharmacophore hypothesis
[0204] Based on X-ray data available in the PDB database (PDB codes: 3ZOS, 4BKJ, 4CKR, 5BVN, 5BVO, 5FDP, 5FDX, 6GWR), we developed three pharmacophore models describing DDR1 inhibitors. To obtain ligand stacking, the complexes were arranged in 3D. These 3, 4, and 5-centered pharmacophore hypotheses contain key features for binding to the active site of DDR1 kinases, including a hydrogen-bonded acceptor in the hinge region, an aromatic / hydrophobic linker, and a hydrophobic center in a pocket near the DFG motif. For detailed information on pharmacophore features and distances, please refer to [link to relevant documentation]. Figure 2A-2C .
[0205] Nonlinear Sammon Mapping
[0206] To make the final selection, we used a Sammon-based mapping technique. 63The main goal of this algorithm is to approximate the hidden local geometric and topological relationships in the input chemical space on a visually understandable 2D or 3D dimensional map. The basic idea of this method is to drastically reduce the high dimensionality of the initial dataset to a low-dimensional feature space, and in this respect, it is similar to the SOM method and multidimensional scaling transformation. However, compared with other algorithms, the classic Sammon-based method allows scientists to construct a projection that reflects the global topographic relationships as pairwise distances between all objects in the entire space of the input vector sample. Structures that successfully pass all the above selection procedures are used as the input chemical space. To form the mapping, we used the same set of molecular descriptors applied to the SOM of a specific kinase, and added the RMSD values obtained during pharmacophore modeling as additional inputs. Euclidean distance was used as the similarity measure, stress threshold: 0.01, interaction number: 300, optimization step: 0.3, and structural similarity coefficient: 0.5. The resulting mapping ( Figure 3 This indicates that the structure follows a normal distribution in the Sammon plot.
[0207] Molecular generation and selection procedures
[0208] Using the GENTRL model, we learn the manifold by sampling the latent code. Neutralization is allocated from the decoder through the sampling structure. 30,000 unique and effective structures were generated. To select a batch of molecules for synthetic and biological research, we developed a priority sorting pipeline (see examples of rejected molecules). Figure 13). In the initial step, the dataset was reduced to a size of 12,147 compounds using the following molecular descriptor thresholds: -2 < logP < 7, 250 < MW < 750, HBA + HBD < 10, TPSA < 150, NRB < 10. Subsequently, 150 MCFs were applied to remove potentially toxic structures and compounds containing reactive and undesired groups. These include: substrates for 1,4-addition (Michael addition moieties) and other electrophilic species (such as para- or ortho-halogenated pyridines, 2-halogenated furans and thiophenes, haloalkanes, aldehydes, and acid anhydrides, etc.), disulfides, isatins, barbituric acids, strained heterocycles, fused polycyclic aromatic hydrocarbon systems, detergents, hydroxamic acids, and diazo compounds, peroxides, labile fragments, and sulfonate ester derivatives. Additionally, we used finer filtering rules, including the following: < 2 NO2 groups, < 3 Cl, < 2 Br, < 6 F, < 5 aromatic rings, undesired atoms (such as Si, Co, or P), to reasonably reduce the number of structures distributed across the entire chemical space to a more compact and drug-biased one. This process yielded 7,912 structures. Then, clustering analysis was performed based on Tanimoto similarity and the standard Morgan fingerprints implemented in the RDKit package. All compounds satisfying the 0.6 similarity threshold were assigned to the same cluster, with a minimum of 5 structures per cluster. Within each cluster, the compounds were classified according to their internal coefficient of variation 52 , to output the maximum gradability of the top 5 items in the structure. As a result, the dataset was reduced to 5,542 molecules. Then, we performed a similarity search using a collection of suppliers (MolPort 64 and ZINC 51 ), and an additional 900 compounds with similarity > 0.5 were removed to increase the novelty of the generated structures. Using the general kinase SOM and the specific kinase SOM, the compounds were prioritized based on their potential activity against the DDR1 kinase. Among the 2,570 molecules classified as kinase inhibitors by the general kinase SOM, 1,951 molecules were classified as DDR1 inhibitors by the specific kinase SOM and were used for pharmacophore-based virtual screening. For each molecule, 10 conformations were generated and minimized using the Universal Force Field (UFF) 65 implemented in RDKit. Using the developed hypotheses, a screening program was then executed, and a set of RMSD values for 848 molecules was obtained, which matched at least one pharmacophore hypothesis. Based on Sammon mapping, we performed a random selection procedure, paying particular attention to the RMSD value regions obtained for 4- and 5-center pharmacophores (Table 5). As a result, 40 molecules were selected for synthesis and subsequent biological evaluation.
[0209] Table 5 – RMSD values (Å) of 40 molecules matching the pharmacophore hypothesis
[0210]
[0211] 3c_rmsd – RMSD value of compounds matching the 3-center pharmacophore hypothesis
[0212] 4c_rmsd – RMSD value of compounds matching the 4-center pharmacophore hypothesis
[0213] 5c_rmsd – RMSD value of compounds matching the 5-center pharmacophore hypothesis
[0214] NA – Compounds that do not match the pharmacophore hypothesis
[0215] Detailed calculations from scratch
[0216] We performed first-principles calculations for the lowest conformational isomer, as predicted by the RDKit / UFF method described above. Geometric optimization was performed using a block-related coupled-cluster method, including single and double excitations (LCCSD) with a 6–31++G basis set. The final energies were calculated at the LCCSD(T) theoretical level. The initial molecular orbitals were obtained using the localized Pipek-Mezey procedure.
[0217] Docking simulation
[0218] In Maestro software 66 Molecular modeling was performed. The PDB structure was preprocessed using 3ZOS and its energy minimized using the Prep module. A binding site grid was generated around the ATP binding site with a buffer size of 20 Å. Docking postures were generated via ultraprecise (XP) Glide using the optimized ligand structure. The final model was selected based on its lowest docking fraction of -15 kcal / mol.
[0219] Cytochrome inhibition
[0220] Water used in experiments and analyses was purified using the ELGA laboratory purification system. Potassium phosphate buffer (PB) at a concentration of 100 mM and MgCl2 at a concentration of 33 mM were used. Working solutions (100×) of test compounds (compound 1 and compound 2) and standard inhibitors (α-naphthylflavonoid, sulfadiazine, 3-benzyl-5-ethyl-5-phenyl-2,4-imidazolidinedione, quinidine, ketoconazole) were prepared. Microsomes were removed from the -80°C freezer, thawed on ice, dated, and returned to the freezer immediately after use. 20 µL of matrix solution was added to the corresponding wells. 20 µL of PB was added to the blank wells. 2 µL of the test compound and positive control working solutions were added to the corresponding wells. Then, the HLM working solution was prepared. 158 µL of the HLM working solution was added to all wells of the culture plate. The culture plate was preheated in a 37°C water bath for approximately 10 minutes. Then, the NADPH cofactor solution was prepared. 20 µL of NADPH cofactor was added to all wells. The solutions were mixed and incubated in a 37 °C water bath for 10 min. At this time point, the reaction was terminated by adding 400 µL of cold stop solution (200 ng / mL tolbutamide and 200 ng / mL labetalol in ACN). The samples were centrifuged at 4000 rpm for 20 min to precipitate the proteins. 200 µL of the supernatant was transferred to 100 µL of HPLC water and agitated for 10 min. The percentage of the media control group to the test compound concentration was plotted using XL fit, and nonlinear regression analysis was performed on the data. The IC50 was determined using a three-parameter or four-parameter logistic equation. 50 Value. When the inhibition rate is less than 50% at the highest concentration (50 µM), IC50 value. 50 The value was reported as ">50µM". See Table 2.
[0221] Microparticle stability
[0222] The microsomal stability of compound 2 was assessed as follows: Working solutions of compound 2 and control compounds (testosterone, diclofenac, propafenone) were prepared. An appropriate amount of NADPH powder (β-nicotinamide adenine dinucleotide phosphate reduced form, tetrasodium salt, NADPH·4Na, supplier: Chem Impex International, catalog number 00616) was weighed and diluted in MgCl2 (10 mM) solution (working solution concentration: 10 units / mL; final concentration in the reaction system: 1 unit / mL). Appropriate concentrations of microsomal working solutions were prepared using 100 mM potassium phosphate buffer (human: HLM, catalog number 452117, Corning; SD rat: RLM, catalog number R1000, Beijing Bolide Biotechnology Co., Ltd.; CD-1 mouse: MLM, catalog number M1000, Beijing Bolide Biotechnology Co., Ltd.; Beagle: DLM, catalog number D1000, Beijing Bolide Biotechnology Co., Ltd.). Ice-cold acetonitrile (ACN) containing 100 ng / mL tolbutamide and 100 ng / mL labetalol as internal standards (IS) was used as a stop solution. 10 μL of the compound or control working solution / well was added to all culture plates (T0, T5, T10, T20, T30, T60, NCF60), except for the matrix blank. 80 μL / well of microparticle solution was added to each culture plate via Apricot dispensing, and the mixture of microparticle solution and compound was incubated at 37 °C for approximately 10 min. After preheating, 10 μL / well of NADPH regeneration system was dispensed to each culture plate via Apricot to initiate the reaction. The solution was then incubated at 37 °C. 300 μL / well of stop solution was added (cooled at 4 °C) to terminate the reaction. The sampling plate was shaken for approximately 10 min. The samples were centrifuged at 4000 rpm for 20 min at 4 °C. During centrifugation, load 300 μL of HPLC water into an 8× new 96-well plate, then transfer 100 μL of the supernatant and mix for LC / MS / MS. See Table 3.
[0223] Buffer stability
[0224] The stability of compound 2 was assessed in phosphate-buffered saline (pH=7.0 and pH=7.4). The test compound (10 μM) was incubated at 25 °C with 50 mM phosphate-buffered saline (pH=7.4), 8 mM MOPS (pH=7.0), and 0.2 mM EDTA (pH=7.0). The experiment was repeated twice. Samples were removed at time points (0, 120, 240, 360, and 1440 minutes) and immediately mixed with 50% ice-cold acetonitrile / water containing the internal standard (IS). Curcumin was used as a positive control under moderately alkaline conditions in this assay. The clearance of the test compound was assessed by LC / MS / MS analysis based on the analyte / IS peak area ratio (without a standard curve). See Table 4.
[0225] Hyperparameters of the GENTRL model
[0226] Architecture
[0227] encoder It is a 2-layer recurrent neural network with gated recurrent units (GRUs) of size 128 hidden. 69 Decoder It is a 7-layer stacked dilated convolution with 128 channels. Latent space z It has 50 dimensions, with 10 mixed components in each dimension. The core size for tensor training. m It's 30.
[0228] Autoencoder training:
[0229] We use multiple molecular properties We learn the mapping from chemical space to the potential manifold. For all training molecules, we compute MCE-18 and the binary flag MCF, which indicates whether a molecule successfully passes through the medicinal chemistry filter. For molecules from the kinase and "negative" datasets, we explicitly specify whether a molecule is a kinase (this value is unknown for molecules outside these datasets). For known DDR1 inhibitors, we explicitly specify pIC. 50 For each update, we constructed batches of 200 molecules: 60 active molecules from DDR1, 60 molecules from ZINC, 20 molecules from the kinase dataset, 20 molecules from the negative dataset, and 40 molecules from patent records. We used a learning rate of 10-1. -4 Adam 70 The optimizer runs optimization programs for 300,000 updates.
[0230] Reinforcement learning:
[0231] We trained the model using a reinforcement learning algorithm with a learning rate of 2. 10 -5 We update 2000 times with the Adam optimizer and a batch size of 200. We use a probability of 0.1 and a standard batch size... For exploratory batches Sampling. From a random sample of 50,000 molecules, it was estimated that 97.8% were valid and 73% were unique.
[0232] Biological research
[0233] Biochemical experiments
[0234] Execute experimental procedures using Eurofins. Use KinaseProfiler.67 The activity of the molecular inhibitors against hDDR1 and hDDR2 kinases was assessed. Enzyme samples were reacted with 8 mM MOPS buffer (pH 7.0), 0.2 mM EDTA, and 250 µM target protein, respectively. h DDR1 and h DDR2 kinase, 10 mM magnesium acetate / manganese chloride and [ - 33 P-ATP culture. At room temperature, this enzymatic reaction is carried out in Mg... 2+ The reaction was carried out for 40 minutes in the presence of cations and ATP, and terminated by the addition of phosphate. The reaction mixture (10 µL) was dropped onto a P30 filter and washed four times with 0.425% phosphate and once with methanol. All compounds were prepared in 100% DMSO. Astrococcus was used as a reference inhibitor and added to each culture plate at the estimated concentration leading to complete inhibition. In Dundee Eurofins, the Scan Max Delta Panel [10 µM ATP] Kinase Profiler was used. 68 Perform a selective biological assessment of non-target kinases.
[0235] Autophosphorylation
[0236] The human DDR1b gene with an HA tag was cloned into the pCMV Tet-On vector (cloning technology), and the stable induced cell line established in U2OS was used for IC 50 Testing. Cells were seeded in 12-well culture plates and DDR1b expression was induced for 48 h at 37°C in a humidity-controlled incubator with 5% CO2 and 10 µg / ml doxycycline hydrochloride (Selleckchem#S4163) before DDR1 activation via rat tail type I collagen (sigma#11179179001). Cells were isolated using trypsin and transferred to 15 ml tubes. Cells were then treated with the compound for 1.5 h at 37°C in the presence of 10 µg / ml rat tail type I collagen after 0.5 h of pretreatment. At the end of treatment, each sample was washed once with ice-cold PBS and lysed for 20 min at 4°C in RIPA buffer containing protease and phosphatase inhibitors (Sigma#0278, Sigma#P5726, and Sigma#P0044). Lysates were removed by centrifugation, and the supernatant was analyzed by Western blotting to obtain the exfoliated activated human DDR1b (Y513) (cell signal #14531S), total DDR1b (HA tag, sigma #H9658), and GAPDH. The integrated intensity of each band was quantified, and the IC50 of the evaluation compounds was calculated on a 10.3-fold dilution series.50 value.
[0237] MRC-5 fiberization test
[0238] MRC-5 cells were grown in Eagle (Sigma, M2279) medium, which provided 1% MEM non-essential amino acids (Merck Life Sciences, Inc., 11140-050), 10% fetal bovine serum (Hyclone, SV30087.03), penicillin (100 U / mL)-streptomycin (100 µg / mL) (Merck Millipore, TMS-AB2-C), and 2 mM L-glutamine (Merck Life Sciences, Inc., 25030-001). After 24 hours of growth in 12-well plates, the cell culture medium was replaced with the same medium as above (except for the use of 2% fetal bovine serum). After 20 hours of growth in serum-reduced medium, the cells were treated with the specified dose of the compound for 30 minutes. Subsequently, the cells were stimulated with 10 ng / mL TGF-β (R&D system, 240-B-002) for 48 or 72 hours. Cells were washed twice with DPBS, and then harvested at 4°C with 100 mL of RIPA buffer (Sigma, R0278) supplemented with a protease inhibitor mixture (Roche, 04693132001). Cells were analyzed using a BCA protein assay kit (Pierce). TM The total protein content in each sample was quantified, and an equal amount of total protein from each sample was loaded onto the WES fully automated Western Blot system (ProteinSimple, Bio-techne) following the manufacturer's instructions. The antibodies used were mouse anti-α-actin (SPM332) (sc-365970) and mouse anti-CTGF (E5) (sc-365970) from Santa Cruz Biotechnologies; and mouse anti-GAPDH (6C5) (EMDMillipore, MAB374).
[0239] LX-2 Fiberization Test
[0240] Human hepatic stellate cells (LX-2) were grown in DMEM (Ingenie Biotech, Inc., 11960), which provided 1% MEM non-essential amino acids (Ingenie Biotech, Inc., 11140-050), 2% fetal bovine serum (Hyclone, SV30087.03), penicillin (100 U / mL)-streptomycin (100 μg / mL) (Merck Millipore, TMS-AB2-C), and 2 mM L-glutamine (Ingenie Biotech, Inc., 25030-001). After growing the cells in 12-well plates for 24 hours, the cell culture medium was replaced with the same medium as above (except for the use of 0.4% fetal bovine serum). After growing in serum-reduced medium for 20 hours, the cells were treated with the specified dose of the compound for 30 minutes. Subsequently, the cells were stimulated with 4 ng / mL TGF-β (R&D system, 240-B-002) for 48 hours. Cells were washed twice with DPBS, and then harvested at 4°C with 100 μL of RIPA buffer (Sigma, R0278) supplemented with a protease inhibitor mixture (Roche, 04693132001). Cells were analyzed using a BCA protein assay kit (Pierce). TM The total protein content in each sample was quantified, and an equal amount of total protein from each sample was analyzed by Western blotting. The antibodies used were mouse anti-α-actin (SPM332) (sc-365970), mouse anti-CTGF (E5) (sc-365970), and mouse anti-collagen α1 (3G3) (sc-293182) from Santa Cruz Biotechnology; and mouse anti-GAPDH (6C5) (EMD Millipore, MAB374).
[0241] Cytotoxicity
[0242] In the presence of the compound, LX-2 cells were seeded into 96-well culture plates and allowed to grow for 72 hours. LX-2 cell viability was then assessed using CellTiter-Glo® luminescence assay according to the manufacturer's instructions. Cytotoxicity (CC) was calculated for 10 series of 3-fold compound dilutions using GraphPad Prism software. 50 ).
[0243] The physicochemical properties of compounds 1 and 2 are shown in Table 6.
[0244] Table 6
[0245]
[0246] *Predicted
[0247] **LE=1.4*(pIC)50 ) / N, N——Number of heavy atoms
[0248] Results of a complete mouse PK study of intravenous (IV) administration of compound 1 at 10 mg / kg. Formulation: 5 mg / mL, clear solutions in NMP / PEG400 / H2O = 1:7:2 are shown in Table 7.
[0249] Table 7
[0250]
[0251] ND = Undetermined (Parameters are undetermined because the final elimination phase is not sufficiently defined)
[0252] BQL = Below the lower limit of quantification (LLOQ)
[0253] If the adjusted rsq (linear regression coefficient of the final phase concentration value) is less than 0.9, it may be impossible to accurately estimate T. 1 / 2 .
[0254] If AUC 额外 If the percentage is greater than 20%, the AUC may not be accurately estimated. 0-inf Cl, MRT 0-inf and Vd ss .
[0255] If AUMC 额外 If the percentage is greater than 20%, it may be impossible to accurately estimate the MRT. 0-inf and Vd ss .
[0256] If the linear regression coefficient of the adjusted final phase concentration is less than 0.9, it may be impossible to accurately estimate T. 1 / 2 .
[0257] a: Using AUC 0-inf (AUC) 额外 %<20%) or AUC 0-最后 (AUC) 额外 Bioavailability (%) is calculated based on the dosage (%>20%) and the administered dose.
[0258] Results of a complete mouse PK study of oral (PO) administration of 15 mg / kg of compound 1. Formulation: 3 mg / mL, in NMP / PEG400 / H2O = 1:7:2, clear solutions are shown in Table 8.
[0259] Table 8
[0260]
[0261] ND = Undetermined (Parameters are undetermined because the final elimination phase is not sufficiently defined)
[0262] BQL = Below the lower limit of quantification (LLOQ)
[0263] If the adjusted rsq (linear regression coefficient of the final phase concentration value) is less than 0.9, it may be impossible to accurately estimate T. 1 / 2 .
[0264] If AUC 额外 If the percentage is greater than 20%, the AUC may not be accurately estimated. 0-inf Cl, MRT 0-inf and Vd ss .
[0265] If AUMC 额外 If the percentage is greater than 20%, it may be impossible to accurately estimate the MRT. 0-inf and Vd ss .
[0266] If the linear regression coefficient of the adjusted final phase concentration is less than 0.9, it may be impossible to accurately estimate T. 1 / 2 .
[0267] a: Using AUC 0-inf (AUC) 额外 %<20%) or AUC 0-最后 (AUC) 额外 Bioavailability (%) is calculated based on the dosage (%>20%) and the administered dose.
[0268] Case Studies
[0269] In the first case study, a generative tensor reinforcement learning (GENTRL) model was used to generate in vivo active DDR1 and DDR2 inhibitors (see Figures 1, 12, and 15). As described herein, a GENTRL model (e.g., a model or combination of models) is a high-level VAE with RL optimization, prioritizing compounds by synthetic feasibility, their effectiveness against biological targets, and how they can be distinguished from other molecules. Model training and validation were performed using a set of molecules obtained from the ZINC database. Structures containing atoms other than carbon, nitrogen, oxygen, sulfur, fluorine, chlorine, bromine, and hydrogen were removed. MCF was applied to excluded compounds with potentially toxic and active groups. Additionally, a dataset including known DDR1 kinase inhibitors, a set of common kinase inhibitors (positive group), a set of molecules acting on non-kinase targets (negative group), and patent data for bioactive molecules already claimed by pharmaceutical companies were used. The Integrity database was used to collect a dataset of structures of new drugs claimed in patent records and 3D structures of DDR1 inhibitors.
[0270] The GENTRL model generated 30,000 unique and effective structures, which were prioritized as follows. First, the set was reduced to 12,147 compounds using a molecular descriptor threshold, and MCF was applied to remove potentially toxic structures and compounds containing reactive and undesirable groups. Then, filtering rules were used to exclude undesirable atoms such as silicon, cobalt, or phosphorus. This process yielded a subset of 7,912 structures. Subsequent clustering analysis based on Tanimoto similarity and standard Morgan fingerprints produced a set of 5,542 molecules. These were then prioritized using universal kinase SOMs and specific kinase SOMs, prioritizing them by their activity against the DDR1 kinase. 1,951 molecules were classified as DDR1 inhibitors and used for pharmacophore-based virtual screening. Random selection was performed based on the Sammon map, with particular attention paid to RMSD value regions obtained from four and five central pharmacophores.
[0271] A series of filtration steps further reduced the number of compounds considered. Based on the final scoring results, 40 molecules were selected for subsequent biological evaluation. Six of these were selected based on their selectivity and SA (anti-inflammatory response). The in vitro inhibitory activity of these six compounds (e.g., compounds 1-6) was tested in an enzyme kinase assay. The results showed that two compounds with favorable physicochemical properties strongly inhibited DDR1 activity while exhibiting selectivity for DDR1. The pharmacokinetics of the most promising compounds were tested in vivo using a rodent model.
[0272] Second Case Study
[0273] The second case study (see Figures 1, 12, and 15) aimed to generate a non-covalent inhibitor of the SARS-CoV-2 3C-like protease (Mpro). The reason for initiating the project to design non-covalent inhibitors was that most initial efforts within the scientific community focused on the development of covalent inhibitors of SARS-CoV-2 Mpro, while the possibility of identifying non-covalent inhibitors remained relatively unexplored. However, covalent inhibitors are not always suitable for treating patients due to the high probability of side effects, and non-covalent inhibitors, with their higher therapeutic index, offer another possibility, albeit with reduced potency.
[0274] The methodology and timeframe are similar to the DDR1 case study. The first step is to select, filter, and process the datasets to be used for training and validating the DL model. The main protease dataset consists of antiprotease active molecules extracted from the Integrity database in enzyme assays. Duplicate structures were excluded and salt portions were removed during the normalization process. MCF was applied to exclude non-drug-like molecules and cyclic structures containing more than 8 atoms and peptides. The final dataset contains 60,293 unique structures. To adapt the scoring metrics and reward function to the types of structures to be generated, common peptidoglycan structures were queried using SMARTS to extract protease peptidoglycan datasets from the protease dataset. The final protease peptidoglycan dataset contains 5,891 compounds.
[0275] To generate novel molecular structures, ligand-based and pocket-based methods were employed. For this purpose, pocket and ligand features were obtained from the amino acid environment at the binding site and from the co-crystallized fragment 6W63 from the same PDB record, respectively.
[0276] In total, 28 ML models were used to generate molecular structures. These structures were optimized using reward-based RL, which is a weighted sum of multiple intermediate rewards. These included medicinal chemistry and drug similarity scores, activity chemistry scores, structural scores, novelty and grading scores, and SA scores. The ML models utilized different molecular representations, including fingerprints, string representations, and graphs. Each model optimized its reward function to explore the chemical space, investigate promising clusters, and generate molecular structures. These structures shared a common structural pattern with peptide-like compounds. Following a series of filtering steps, 20 compounds were selected for synthesis and further biological evaluation, including enzyme and cell-based antiviral assays.
[0277] The generated structure:
[0278] The system and methods described in this paper have been used to generate many new structures, including those with ChEMBL similarity scores <700. Some examples are shown in this paper, such as in... Figure 14 middle.
[0279] For the models, processes, and methods disclosed herein, the operations performed in the processes and methods may be implemented in different orders. Furthermore, the operations outlined are provided merely as examples, and some operations may be optional, combined into fewer operations, canceled, supplemented by further operations, or extended into additional operations without departing from the essence of the disclosed embodiments.
[0280] This invention is not limited to the specific embodiments described herein, which are intended to be illustrative of various aspects. Many modifications and variations can be made without departing from its spirit and scope. In addition to the methods and apparatuses listed herein, functionally equivalent methods and apparatuses within the scope of the invention are also possible from the foregoing description. Such modifications and variations are intended to fall within the scope of the appended claims. This invention is limited only by the terms of the appended claims and the full scope of their equivalents. The terminology used herein is merely for describing particular embodiments and is not intended to be limiting.
[0281] In one embodiment, the method of the present invention may include aspects that are executed on a computing system. Therefore, the computing system may include a memory device having computer-executable instructions for performing the method. The computer-executable instructions may be part of a computer program product comprising one or more algorithms for performing any of the methods described in any one of the claims.
[0282] In one embodiment, any operation, process, or method described herein may be performed or caused to be performed in response to the execution of computer-readable instructions stored on a computer-readable medium and executable by one or more processors. The computer-readable instructions may be executed by processors of a variety of computing systems, including desktop computing systems, portable computing systems, tablet computing systems, handheld computing systems, and network elements and / or any other computing device. The computer-readable medium is not transient. A computer-readable medium is a physical medium having computer-readable instructions stored therein, so that a computer / processor can physically read from the physical medium.
[0283] Various media (e.g., hardware, software, and / or firmware) are available for carrying out the processes and / or systems and / or other technologies described herein, and the preferred media may vary depending on the deployment environment of the processes and / or systems and / or other technologies. For example, if the implementer determines that speed and accuracy are paramount, the implementer may primarily choose hardware and / or firmware media; if flexibility is paramount, then the implementer may primarily choose a software implementation; or, alternatively, the implementer may choose some combination of hardware, software, and / or firmware.
[0284] The various operations described herein can be implemented individually and / or collectively by a wide variety of hardware, software, firmware, or virtually any combination thereof. In one embodiment, several portions of the subject matter described herein can be implemented via an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a digital signal processor (DSP), or other integrated form. However, some aspects of the embodiments disclosed herein can be implemented, either wholly or partially equivalently, in an integrated circuit as one or more computer programs running on one or more computers (e.g., as one or more programs running on one or more computer systems), as one or more programs running on one or more processors (e.g., as one or more programs running on one or more microprocessors), as firmware, or as virtually any combination thereof, and it is possible to design circuit systems and / or write code for software and / or firmware according to this disclosure. Furthermore, the mechanisms of the subject matter described herein can be distributed as program products in a variety of forms, and the illustrative embodiments of the subject matter described herein apply regardless of the specific type of signal-bearing medium used to actually perform the distribution. Examples of physical signal-carrying media include, but are not limited to, the following: recordable media, such as floppy disks, hard disk drives (HDDs), compact optical discs (CDs), digital versatile optical discs (DVDs), digital magnetic tapes, computer memory, or any other non-transitory or non-transferable physical media. Examples of physical media with computer-readable instructions omit transient or transferable media, such as digital and / or analog communication media (e.g., fiber optic cables, waveguides, wired communication links, wireless communication links, etc.).
[0285] Devices and / or processes are typically described in the manner set forth herein, and are subsequently integrated into data processing systems using engineering practice. That is, at least a portion of the devices and / or processes described herein can be integrated into a data processing system through a reasonable amount of experimentation. A typical data processing system typically includes one or more of the following: a system unit enclosure; a video display device; memory, such as volatile and non-volatile memory; a processor, such as a microprocessor and a digital signal processor; computing entities, such as an operating system, drivers, a graphical user interface, and applications; one or more interactive devices, such as a touchpad or screen; and / or a control system, including feedback loops and control motors (e.g., feedback for sensing position and / or speed; control motors for moving and / or adjusting the number and quantity of components). A typical data processing system can be implemented using any suitable commercially available components, such as those commonly found in data computing / communication and / or network computing / communication systems.
[0286] The topics described herein sometimes illustrate different components contained in or connected to different other components. Such depicted architectures are merely exemplary, and in fact, many other architectures can be implemented to achieve the same functionality. Conceptually, any arrangement of components that achieves the same functionality is effectively “associated” to achieve the desired function. Therefore, any two components combined herein to achieve a particular function can be considered “associated” with each other to achieve the desired function, regardless of the architecture or intermediate components. Similarly, any two components so associated can also be considered “operably connected” or “operably coupled” to each other to achieve the desired function, and any two components that can be so associated can also be considered “operably coupled” to each other to achieve the desired function. Specific examples of being operablely coupled include, but are not limited to: components that are physically compatible and / or physically interact with each other, and / or components that can wirelessly interact and / or wirelessly interact with each other, and / or components that logically interact and / or are logically interdependent.
[0287] Figure 15 An exemplary computing device 600 (e.g., a computer) is shown, which can be arranged in some embodiments to perform the methods (or portions thereof) described herein. In a very basic configuration 602, the computing device 600 typically includes one or more processors 604 and system memory 606. A memory bus 608 may be used for communication between the processor 604 and the system memory 606.
[0288] Depending on the desired configuration, processor 604 can be of any type, including but not limited to: microprocessor (µP), microcontroller (µC), digital signal processor (DSP), or any combination thereof. Processor 604 may include one or more levels of cache, such as L1 cache 610 and L2 cache 612, processor core 614, and registers 616. Exemplary processor core 614 may include an arithmetic logic unit (ALU), floating-point unit (FPU), digital signal processing core (DSP core), or any combination thereof. Exemplary memory controller 618 may also be used with processor 60, or in some implementations, memory controller 618 may be an internal part of processor 604.
[0289] Depending on the desired configuration, system memory 606 can be of any type, including but not limited to: volatile memory (e.g., RAM), non-volatile memory (e.g., ROM, flash memory, etc.), or any combination thereof. System memory 606 may include operating system 620, one or more application programs 622, and program data 624. Application program 622 may include determining application 626, which is configured to perform operations as described herein, including operations related to the methods described herein. Determining application 626 may acquire data, such as pressure, flow rate, and / or temperature, and then determine changes in the system to alter the pressure, flow rate, and / or temperature.
[0290] Computing device 600 may have additional features or functions, as well as additional interfaces, to facilitate communication between basic configuration 602 and any desired devices and interfaces. For example, bus / interface controller 630 may be used to facilitate communication between basic configuration 602 and one or more data storage devices 632 via storage interface bus 634. Data storage device 632 may be removable storage device 636, non-removable storage device 638, or a combination thereof. Examples of removable and non-removable storage devices include, to name a few, disk devices such as floppy disk drives and hard disk drives (HDDs); optical disk drives such as CD drives or digital versatile disk (DVD) drives; solid-state drives (SSDs) and tape drives. Exemplary computer storage media may include volatile and non-volatile, removable and non-removable media implemented for any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data.
[0291] System memory 606, removable storage device 636, and non-removable storage device 638 are examples of computer storage media. Computer storage media include, but are not limited to: RAM, ROM, EEPROM, flash memory or other storage technologies, CD-ROM, digital versatile disc (DVD) or other optical storage devices, magnetic tape cassettes, magnetic tape, disk storage devices or other magnetic storage devices, or any other medium that can be used to store desired information and is accessible by computing device 600. Any such computer storage medium may be part of computing device 600.
[0292] The computing device 600 may also include an interface bus 640 for facilitating communication from various interface devices (e.g., output device 642, peripheral interface 644, and communication device 646) to the basic configuration 602 via a bus / interface controller 630. An exemplary output device 642 includes a graphics processing unit 648 and an audio processing unit 650, which may be configured to communicate with various external devices (e.g., displays or speakers) via one or more A / V ports 652. An exemplary peripheral interface 644 includes a serial interface controller 654 or a parallel interface controller 656, which may be configured to communicate with external devices such as input devices (e.g., keyboards, mice, pens, voice input devices, touch input devices, etc.) or other peripheral devices (e.g., printers, scanners, etc.) via one or more I / O ports 658. An exemplary communication device 646 includes a network controller 660, which may be configured to facilitate communication with one or more other computing devices 662 over a network communication link via one or more communication ports 664.
[0293] A network communication link can be an example of a communication medium. A communication medium can typically be embodied in computer-readable instructions, data structures, program modules, or other data (such as a carrier wave or other transmission mechanism) in a modulated data signal, and can include any information delivery medium. A “modulated data signal” can be a signal having one or more characteristics set or altered in a manner that encodes information in the signal. By way of example and not limitation, a communication medium can include wired media, such as wired networks or direct wired connections; and wireless media, such as acoustic, radio frequency (RF), microwave, infrared (IR), and other wireless media. The term computer-readable medium, as used herein, can include storage media and communication media.
[0294] The computing device 600 can be implemented as part of a miniaturized portable (or mobile) electronic device, such as a mobile phone, personal data assistant (PDA), personal media player device, wireless smartwatch device, personal headset device, dedicated device, or a hybrid device including any of the above functions. The computing device 600 can also be implemented as a personal computer, including both laptop and non-laptop computer configurations. The computing device 600 can also be any type of network computing device. The computing device 600 can also be an automated system as described herein.
[0295] The embodiments described herein may include the use of a dedicated or general-purpose computer, including various computer hardware or software modules.
[0296] Embodiments within the scope of this invention also include computer-readable media for carrying or having computer-executable instructions or data structures stored thereon. Such computer-readable media can be any available medium accessible to a general-purpose or special-purpose computer. By way of example, and not limitation, such computer-readable media can include RAM, ROM, EEPROM, CD-ROM or other optical disc storage, disk storage devices or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code meaning in the form of computer-executable instructions or data structures and is accessible to a general-purpose or special-purpose computer. When information is transmitted or provided to a computer via a network or other communication connection (hardwired, wireless, or a combination of hardwired and wireless), the computer appropriately considers the connection as a computer-readable medium. Therefore, any such connection is appropriately referred to as a computer-readable medium. Combinations of the above should also be included within the scope of computer-readable media.
[0297] Computer-executable instructions include, for example, instructions and data that cause a general-purpose computer, a special-purpose computer, or a special-purpose processing device to perform a function or group of functions. Although the subject matter has been described in language specific to structural features and / or methodological actions, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are disclosed as examples of implementing the claims.
[0298] In some embodiments, a computer program product may include a tangible non-transient storage device having computer-executable instructions that, when executed by a processor, cause the execution of a method, the method comprising: providing a dataset having object data for objects and condition data for conditions; processing the object data of the dataset with an object encoder to obtain potential object data and potential object-condition data; processing the condition data of the dataset with a condition encoder to obtain potential condition data and potential condition-object data; processing the potential object data and potential object-condition data with an object decoder to obtain generated object data; processing the potential condition data and potential condition-object data with a condition decoder to obtain generated condition data; comparing the potential object-condition data with the potential condition data to determine a difference; processing the potential object data and potential condition data, and either the potential object-condition data or the potential condition-object data, with a discriminator to obtain a discriminator value; selecting a selected object from the generated object data based on the generated object data, the generated condition data, and the difference between the potential object-condition data and the potential condition-object data; and providing a report to the selected object with a recommendation to verify the physical form of the object. The tangible non-transient storage device may also have other executable instructions for any of the methods or method steps described herein. In addition, instructions can be for performing non-computational tasks, such as the synthesis of molecules and / or experimental protocols for verifying molecules. Other executable instructions may also be provided.
[0299] Regarding the use of virtually any plural and / or singular terms in this document, those skilled in the art may, where appropriate, convert from plural to singular and / or from singular to plural depending on the context and / or application. For clarity, various singular / plural permutations may be explicitly described herein.
[0300] Those skilled in the art will understand that, in general, the terms used herein, and especially in the appended claims (e.g., the body of the appended claims), are typically intended as “open” terms (e.g., the term “including” should be interpreted as “including but not limited to,” the term “having” should be interpreted as “at least having,” the term “include” should be interpreted as “including but not limited to,” etc.). Those skilled in the art will further understand that if there is an intent to have a specific number of introduced claim statements, such intent will be expressly stated in the claims; if no such statements are present, such intent does not exist. For example, to aid understanding, the appended claims may include the introductory phrases “at least one” and “one or more” to introduce claim statements. However, the use of such phrases should not be construed as implying that introducing a claim statement with the indefinite article “a (a)” or “an” limits any particular claim containing such a claim statement to an embodiment containing only one such statement, even when the same claim contains the introductory phrase “one or more” or “at least one” and an indefinite article such as “a (a)” or “an” (e.g., “a (a)” and / or “an” should be interpreted as “at least one” or “one or more”); the same applies to the use of definite articles to introduce a claim statement. Furthermore, even when a specific number of claim statements is explicitly stated, those skilled in the art will recognize that such a statement should be interpreted as indicating at least the number of stated statements (e.g., in the absence of other modifiers, simply stating “two statements” means at least two statements or two or more statements). Furthermore, in the common use of phrases like "at least one of A, B, and C," this structure is intended to have a meaning that those skilled in the art will understand (e.g., "a system having at least one of A, B, and C" will include, but is not limited to, systems having a single A, a single B, a single C, A and B together, A and C together, B and C together, and / or A, B, and C together, etc.). Those skilled in the art will further understand that virtually any separating words and / or phrases presenting two or more alternative terms, whether in the specification, claims, or drawings, should be understood to account for the possibility of including one of the terms, either of the terms, or both of the terms. For example, the phrase "A or B" should be understood to include the possibility of including "A" or "B" or "A and B."
[0301] Furthermore, when the features or aspects of the invention are described according to the Markush group, those skilled in the art will recognize that the invention can also be described according to any individual member or subgroup of the Markush group.
[0302] As those skilled in the art will understand, for any and all purposes, such as for the purpose of providing a written description, all scopes disclosed herein also encompass any and all possible subscopes and combinations thereof. Any listed scope can be readily appreciated as adequately descriptive and capable of being decomposed into at least equal halves, thirds, quarters, fifths, tenths, etc. As a non-limiting example, each scope discussed herein can be readily decomposed into a lower third, a middle third, and an upper third, etc. As those skilled in the art will also understand, all language such as “not greater than”, “at least”, etc., includes the stated numbers and refers to a scope that can subsequently be decomposed into subscopes as discussed above. Finally, as those skilled in the art will understand, a scope includes each individual member. Thus, for example, a group having 1-3 cells means a group having 1, 2, or 3 cells. Similarly, a group having 1-5 cells means a group having 1, 2, 3, 4, or 5 cells, and so on.
[0303] In summary, it should be understood that various embodiments of the invention have been described herein for illustrative purposes, and various modifications may be made without departing from the scope and spirit of the invention. Therefore, the various embodiments disclosed herein are not intended to be limiting, and their true scope and spirit are indicated by the following claims.
[0304] abbreviation
[0305] DCM dichloromethane
[0306] DMF N,N-dimethylformamide
[0307] DMSO dimethyl sulfoxide
[0308] EA ethyl acetate
[0309] HPLC (High Performance Liquid Chromatography)
[0310] PBS phosphate buffer
[0311] THF tetrahydrofuran
[0312] TFA (trifluoroacetic acid)
[0313] TLC (Thin-Layer Chromatography)
[0314] Pyridine
[0315] EDCI 1-(3-Dimethylaminopropyl)-3-ethylcarbodiimide
[0316] ACN acetonitrile
[0317] T3P 1-propylphosphoric anhydride solution
[0318] Pd(dppf)Cl2 [1,1'-bis(diphenylphosphino)ferrocene]palladium(II) dichloride
[0319] NIS N-iodosuccinimide
[0320] TEA Triethylamine
[0321] TFA (trifluoroacetic acid)
[0322] HATU 1-[bis(dimethylamino)methylene]-1H-1,2,3-triazolo[4,5-b]pyridine 3-oxyhexafluorophosphate
[0323] DIEA diisopropylethylamine
[0324] NMP (N-methyl-2-pyrrolidone)
[0325] XPhos Pd G3(2-dicyclohexylphosphino-2′,4′,6′-triisopropyl-1,1′-biphenyl)[2-(2′-amino-1,1′-biphenyl)]palladium(II)methanesulfonate
[0326] MTBE methyl tert-butyl ether
[0327] CDI 1,1'-carbonyldiimidazole
[0328] This patent cross-references the following U.S. Application No. 16 / 015,990, filed June 2, 2018; U.S. Application No. 16 / 134,624, filed September 18, 2018; U.S. Application No. 62 / 727,926, filed September 6, 2018; U.S. Application No. 62 / 746,771, filed October 17, 2018; and U.S. Application No. 62 / 809,413, filed February 22, 2019; which are incorporated herein by reference in their entirety.
[0329] All references cited in this article are incorporated herein in their entirety through specific citations.
[0330] 1. Paul, SM et al. How to improve R&D productivity: thepharmaceutical industry's grand challenge. Nat. Rev. Drug Discov. 9, 203–214(2010).
[0331] 2. Avorn, J. The $2.6 Billion Pill — Methodologic and PolicyConsiderations. N. Engl. J. Med. 372, 1877–1879 (2015).
[0332] 3. Fleming, N. How artificial intelligence is changing drugdiscovery. Nature 557, S55 (2018).
[0333] 4. Chen, H., Engkvist, O., Wang, Y., Olivecrona, M.&Blaschke, T. Therise of deep learning in drug discovery. Drug Discov. Today 23, 1241–1250(2018).
[0334] 5. Butler, K. T., Davies, D. W., Cartwright, H., Isayev, O.&Walsh, A.Machine learning for molecular and materials science. Nature 559, 547–555(2018).
[0335] 6. Segler, M. H. S., Preuss, M.&Waller, M. P. Planning chemicalsyntheses with deep neural networks and symbolic AI. Nature 555, 604–610(2018).
[0336] 7. De Fauw, J. et al. Clinically applicable deep learning fordiagnosis and referral in retinal disease. Nat. Med. 24, 1342–1350 (2018).
[0337] 8. Wang, K., Wan, M., Wang, R.-S.&Weng, Z. Opportunities for Web-based Drug Repositioning: Searching for Potential Antihypertensive Agentswith Hypotension Adverse Events. J. Med. Internet Res. 18, e76 (2016).
[0338] 9. Mamoshina, P., Vieira, A., Putin, E.&Zhavoronkov, A. Applicationsof Deep Learning in Biomedicine. Mol. Pharm. 13, 1445–1454 (2016).
[0339] 10. Baskin, I. I., Winkler, D.&Tetko, I. V. A renaissance of neuralnetworks in drug discovery. Expert Opin. Drug Discov. 11, 785–795 (2016).
[0340] 11. Zhang, L., Tan, J., Han, D.&Zhu, H. From machine learning to deeplearning: progress in machine intelligence for rational drug discovery. DrugDiscov. Today 22, 1680–1685 (2017).
[0341] 12. Goodfellow, I. et al. Generative Adversarial Nets. in Advances inNeural Information Processing Systems 2672–2680 (2014).
[0342] 13. Sanchez-Lengeling, B.&Aspuru-Guzik, A. Inverse molecular designusing machine learning: Generative models for matter engineering. Science361, 360–365 (2018).
[0343] 14. Kadurin, A. et al. The cornucopia of meaningful leads: Applyingdeep adversarial autoencoders for new molecule development in oncology.Oncotarget 8, 10883–10890 (2017).
[0344] 15. Kadurin, A., Nikolenko, S., Khrabrov, K., Aliper, A.&Zhavoronkov,A. druGAN: An Advanced Generative Adversarial Autoencoder Model for de NovoGeneration of New Molecules with Desired Molecular Properties in Silico. Mol.Pharm. 14, 3098–3104 (2017).
[0345] 16. Gómez-Bombarelli, R. et al. Automatic Chemical Design Using aData-Driven Continuous Representation of Molecules. ACS Cent. Sci. 4, 268–276(2018).
[0346] 17. Putin, E. et al. Adversarial Threshold Neural Computer forMolecular de Novo Design. Mol. Pharm. 15, 4386–4397 (2018).
[0347] 18. Putin, E. et al. Reinforced Adversarial Neural Computer for deNovo Molecular Design. J. Chem. Inf. Model. 58, 1194–1204 (2018).
[0348] 19. Harel, S.&Radinsky, K. Prototype-Based Compound Discovery UsingDeep Generative Models. Mol. Pharm. 15, 4406–4416 (2018).
[0349] 20. Polykovskiy, D. et al. Entangled Conditional AdversarialAutoencoder for de Novo Drug Discovery. Mol. Pharm. 15, 4398–4405 (2018).
[0350] 21. Kuzminykh, D. et al. 3D Molecular Representations Based on theWave Transform for Convolutional Neural Networks. Mol. Pharm. 15, 4378–4385(2018).
[0351] 22. Segler, Marwin HS and Kogej, Thierry and Tyrchan, Christian andWaller, Mark P. Generating focused molecule libraries for drug discovery withrecurrent neural networks. ACS central science 4.1 (2017): 120-131.
[0352] 23. Merk, Daniel and Friedrich, Lukas and Grisoni, Francesca andSchneider, Gisbert. De novo design of bioactive small molecules by artificialintelligence. Molecular informatics 37.1-2 (2018): 1700153
[0353] 24. Merk, Daniel and Grisoni, Francesca and Friedrich, Lukas andSchneider, Gisbert. Tuning artificial intelligence on the de novo design ofnatural-product-inspired retinoid X receptor modulators. CommunicationsChemistry 1.1 (2018): 68.
[0354] 25. Sutton, R. S.&Barto, A. G. Reinforcement Learning: AnIntroduction. (MIT Press, 1998).
[0355] 26. Sutton, R. S., McAllester, D., Singh, S.&Mansour, Y. Policygradient methods for reinforcement learning with function approximation. inProceedings of the 12th International Conference on Neural InformationProcessing Systems 1057–1063 (MIT Press, 1999).
[0356] 27. Yu, L., Zhang, W., Wang, J.&Yu, Y. SeqGAN: Sequence GenerativeAdversarial Nets with Policy Gradient. Preprint at arxiv.org / abs / 1609.05473(2017).
[0357] 28. Sohn, K., Lee, H.&Yan, X. Learning Structured OutputRepresentation using Deep Conditional Generative Models. in Advances inNeural Information Processing Systems 3483–3491 (2015).
[0358] 29. Tomczak, J. M.&Welling, M. VAE with a VampPrior. Proceedings ofthe 21st International Conference on Artificial Intelligence and Statistics(AISTATS) (2017).
[0359] 30. Imaizumi, M., Maehara, T.&Hayashi, K. On Tensor Train RankMinimization : Statistical Efficiency and Scalable Algorithm. in Advances inNeural Information Processing Systems 3930–3939 (2017).
[0360] 31. Novikov, A., Rodomanov, A., Osokin, A.&Vetrov, D. Putting MRFs ona Tensor Train. In International Conference on Machine Learning 811–819(2014).
[0361] 32. Oseledets, I. V. Tensor-Train Decomposition. SIAM J. Sci. Comput.33, 2295–2317 (2011).
[0362] 33. Kingma, D. P.&Welling, M. Auto-Encoding Variational Bayes.Preprint at arxiv.org / abs / 1312.6114 (2013).
[0363] 34. Makhzani, A., Shlens, J., Jaitly, N., Goodfellow, I.&Frey, B.Adversarial Autoencoders. Preprint at arxiv.org / abs / 1511.05644 (2015).
[0364] 35. Kingma, D. P., Mohamed, S., Rezende, D. J.&Welling, M. Semi-supervised Learning with Deep Generative Models. in Advances in NeuralInformation Processing Systems 3581– 3589 (2014).
[0365] 36. Leo, A. J. Calculating log Poct from structures. Chem. Rev. 93,1281–1306 (1993).
[0366] 37. Bickerton, G. R., Paolini, G. V., Besnard, J., Muresan, S.&Hopkins, A. L. Quantifying the chemical beauty of drugs. Nat. Chem. 4, 90–98(2012).
[0367] 38. Guimaraes, G. L., Sanchez-Lengeling, B., Farias, P. L. C.&Aspuru-Guzik, A. Objective- Reinforced Generative Adversarial Networks (ORGAN) forSequence Generation Models. Preprint at arxiv.org / abs / 1705.10843 (2017).
[0368] 39. Sanchez-Lengeling, B., Outeiral, C., Guimaraes, G. L.&Aspuru-Guzik, A. Optimizing distributions over molecular space. An Objective-Reinforced Generative Adversarial Network for Inverse-design Chemistry(ORGANIC). Preprint at chemrxiv.org / articles / ORGANIC_1_pdf / 5309668 (2017).
[0369] 40. Kohonen, T. The self-organizing map. Proc. IEEE 78, 1464–1480(1990).
[0370] 41. Wassermann, A. M., Camargo, L. M.&Auld, D. S. Composition andapplications of focus libraries to phenotypic assays. Front. Pharmacol. 5,164 (2014).
[0371] 42. Szymański, P., Markowicz, M.&Mikiciuk-Olasik, E. Adaptation ofHigh-Throughput Screening in Drug Discovery—Toxicological Screening Tests.Int. J. Mol. Sci. 13, 427–452 (2011).
[0372] 43. Morgan, P. et al. Impact of a five-dimensional framework on R&Dproductivity at AstraZeneca. Nat. Rev. Drug Discov. 17, 167–181 (2018).
[0373] 44. Cong, L., Xia, Z.-K.&Yang, R.-Y. Targeting the TGF-β receptorwith kinase inhibitors for scleroderma therapy. Arch. Pharm.347, 609–615(2014).
[0374] 45. Liu, L. et al. Synthesis and biological evaluation of noveldasatinib analogues as potent DDR1 and DDR2 kinase inhibitors. Chem. Biol.Drug Des. 89, 420–427 (2017).
[0375] 46. Drug Discovery | boehringer-ingelheim.com.Available at:boehringer- ingelheim.com / innovation / drug-discovery. (Accessed: 30th October2018).
[0376] 47. Scannell, J. W., Blanckley, A., Boldon, H.&Warrington, B.Diagnosing the decline in pharmaceutical R&D efficiency. Nat. Rev. DrugDiscov. 11, 191–200 (2012).
[0377] 48. Munos, B. Lessons from 60 years of pharmaceutical innovation.Nat. Rev. Drug Discov. 8, 959–968 (2009).
[0378] 49. Clarivate Analytics Integrity. Available at: integrity.thomson-pharma.com / integrity / xmlxsl / . (Accessed: 30th October 2018).
[0379] 50. EBI Web Team. ChEMBL. Available at: ebi.ac.uk / chembl / . (Accessed:30th October 2018).
[0380] 51. Irwin, J. J., Sterling, T., Mysinger, M. M., Bolstad, E. S.&Coleman, R. G. ZINC: a free tool to discover chemistry for biology. J. Chem.Inf. Model. 52, 1757–1768 (2012).
[0381] 52. Trepalin SV, E. al. New diversity calculations algorithms usedfor compound selection. J. Chem. Inf. Comput. Sci. 42, 249–258 (2002).
[0382] 53. Top 25 Global Pharma Companies by Market Cap (end of 2017) |GlobalData Plc. (2018). Available at: globaldata.com / top-25-global-pharma-companies-market-cap-end-2017 / . (Accessed: 30th October 2018).
[0383] 54. Oseledets, I. V. Tensor-Train Decomposition. SIAM J. Sci. Comput.33, 2295–2317 (2011).
[0384] 55. Williams, R. J. Simple statistical gradient-following algorithmsfor connectionist reinforcement learning. Mach. Learn. 8, 229–256 (1992).
[0385] 56. Brown N, Fiscato M, Segler M.H.S., Vaucher A.C. GuacaMol:Benchmarking Models for De Novo Molecular Design arXiv preprint arXiv:1811.09621 (2018).
[0386] 57. Polykovskiy, D et al., Molecular Sets (MOSES): A BenchmarkingPlatform for Molecular Generation Models. arXiv preprint arXiv:1811.12823(2018).
[0387] 58. Ritter, H.&Kohonen, T. Self-organizing semantic maps. Biol.Cybern. 61, 241–254 (1989).
[0388] 59. Landrum, G. RDKit. Available at: rdkit.org. (Accessed: 30thOctober 2018).
[0389] 60. Moriwaki, H., Tian, Y.-S., Kawashita, N.&Takagi, T. Mordred: amolecular descriptor calculator. J. Cheminform. 10, 4 (2018).
[0390] 61. Ivanenkov, Y. A.&Khandarova, L. M. Advanced ArtificialIntelligence Methods Used in the Design of Pharmaceutical Agents. inPharmaceutical Data Mining 457–489 (Wiley, 2009).
[0391] 62. Pletnev, I. V., Ivanenkov, Y. A.&Tarasov, A. V. DimensionalityReduction Techniques for Pharmaceutical Data Mining. in Pharmaceutical DataMining 423–455 (Wiley, 2009).
[0392] 63. Sammon, J. W. A Nonlinear Mapping for Data Structure Analysis.IEEE Trans. Comput. C- 18, 401–409 (1969).
[0393] 64. Molport. Available at: molport.com / . (Accessed: 30th October2018).
[0394] 65. Rappe, A. K., Casewit, C. J., Colwell, K. S., Goddard, W. A.&Skiff, W. M. UFF, a full periodic table force field for molecular mechanicsand molecular dynamics simulations. J. Am. Chem. Soc. 114, 10024–10035(1992).
[0395] 66. Schrödinger | Maestro Suite. Available at: schrodinger.com.(Accessed: 30th October 2018).
[0396] 67. Eurofins Pharma Discovery Services. Available at:eurofinsdiscoveryservices.com / catalogmanagement / viewitem / DDR1-Human-TK-Kinase-Enzymatic-Radiometric-Assay-10-uM-ATP-KinaseProfiler / 14-942KP10.(Accessed: 30th October 2018).
[0397] 68. Eurofins Pharma Discovery Services. Available at:eurofinsdiscoveryservices.com / catalogmanagement / viewitem / scanMAX-Delta-Panel-10uM-ATP / 50-100KP10. (Accessed: 30th October 2018).
[0398] 69. Cho, Kyunghyun; van Merrienboer, Bart; Gulcehre, Caglar;Bahdanau, Dzmitry; Bougares, Fethi; Schwenk, Holger; Bengio, Yoshua. LearningPhrase Representations using RNN Encoder-Decoder for Statistical MachineTranslation. Proceedings of the 2014 Conference on Empirical Methods inNatural Language Processing (EMNLP), pages 1724–1734 (2014).
[0399] 70. Kingma D. P., Ba J. Adam: A method for stochastic optimization,3rd International Conference for Learning Representations, 2015.
[0400] 71. Ivanenkov YA, Zagribelnyy BA, Aladinskiy VA. Are We Opening theDoor to a New Era of Medicinal Chemistry or Being Collapsed to a ChemicalSingularity? J Med Chem. 2019;62: 10026–10043.
Claims
1. Computer-implemented methods, including: Receive input from biological targets or ligands; Receive input on the properties of the generated compound; Receive at least one generative model trained with a reference compound; The generation of a compound is based on the properties of the generated compound as input, and the generated compound has a structure for each generation model, wherein the generated compound is designed to interact with the biological target and / or be associated with the structural features of the ligand. The generation includes refining the generated compound using a reward function and performing pharmacophore modeling. Based on at least one reward criterion, the structures of the generated compounds of each generative model are ranked by priority; The preferred chemical structure of the generated compound is processed using the Sammon mapping protocol to obtain the hit structure, wherein the Sammon mapping protocol uses the root mean square deviation (RMSD) value of the molecular descriptor and pharmacophore modeling used in the reward function; and Provides the chemical structure of the hit structure.
2. The method according to claim 1, further comprising: Receive reference compound; and Generative models are trained using reference compounds.
3. The method according to claim 1, further comprising at least one of the following: Use generative models to refine structures with at least one reward function; or Generative models are used for pharmacophore modeling.
4. The method according to claim 1, wherein, The at least one reward criterion is determined to satisfy at least one of the following: Perform molecular filtration on the generated compound; Perform clustering / singleization operations on the generated compounds; The supplier's compounds are analyzed based on the generated compounds; Execution rewards are sorted by priority; The root mean square deviation value is determined to ensure that the generated compound is suitable for the biological target; or The novelty of the generated compounds was analyzed based on published intellectual property documents.
5. The method of claim 1, wherein prioritizing the structures includes performing a structure refinement protocol using at least one Kohonen self-organizing map (SOM).
6. The method of claim 5, wherein the Kohonen SOM comprises: Trend SOM, whose reward is based on a time-series schedule compared to a newer structure; A general biological target SOM rewards a family of biological targets of a class of structures, said family of biological targets including said biological targets but excluding other classes of structures that are biologically inactive to said biological target family; and A specific biological target SOM rewards a specific structure that targets the biological target.
7. The method of claim 1, further comprising pharmacophore modeling of the generated compound to analyze the scaffold structure or side chain substituent structure of the generated compound, wherein the generated compound is modeled to interact with a biological target and / or modeled to correlate with the structural features of a ligand.
8. The method according to claim 1, further comprising: In the trained GENTRL model, a latent spatial manifold with multiple of the aforementioned reference compounds is generated; and Generate new compounds that do not exist in the latent spatial manifold using a trained GENTRL model.
9. The method of claim 1, wherein prioritizing the structures of the generated compounds based on at least one reward criterion includes at least one of the following: Filter compounds to meet the molecular descriptor range of a predefined range; Apply pharmaceutical chemical filters to remove compounds with undesirable pharmaceutical chemical properties; Applying Ribinsky's five rules to compounds; The T-index is applied to compounds to eliminate structures with an unbalanced number of carbon atoms and heteroatoms; A similarity filter is applied to the compound based on a 2D similarity assessment between the compound and the available chemical space. Applying physicochemical characteristic filters to compounds based on their physicochemical properties; Apply drug similarity filters to compounds based on drug similarity estimation; Privileged structure filters are applied to compounds to prioritize fragments that are useful to selected categories of compounds or biological targets. Based on the comprehensive accessibility assessment of the compounds, a comprehensive accessibility filter is applied to the compounds; Apply diversity filters to compounds based on privileged fragments of the structure; Applying clustering filters to compounds based on privileged fragments of their structure; Tanimoto-based clustering and grouping are applied to remove compounds with similar structures; The application trend SOM is selected based on a time-series schedule compared to an older structure; The application of a general biological target SOM, wherein the SOM is selected to be a structure that is biologically active against a family of biological targets, the family of biological targets including the biological targets; Applying a specific biological target SOM, wherein the SOM selects a specific structure that targets the biological target; or Apply a pharmacophore filter to remove compounds that fail to model pharmacophores.
10. The method of claim 1, further comprising screening for patentable generating compounds.
11. The method of claim 1, further comprising obtaining at least one dataset having a reference compound, the reference compound comprising: Common compounds; Compounds that regulate biological targets; Compounds that regulate biomolecules other than biological targets; Patient data from a set of known compounds that have specific functional activities on biological targets; or Chemical structure data for a set of known compounds that possess the aforementioned specific functional activities for biological targets.
12. The method of claim 1, further comprising identifying an activity threshold for regulating a specific functional activity of a biological target, wherein compounds generated below the activity threshold are defined as inactive compounds, and compounds having an activity threshold or higher are defined as active compounds.
13. The method according to claim 1, further comprising: Process the dataset of reference compounds to exclude outlier compounds; and Normalize the input chemical space by reducing the number of compounds with similar structures in each compound cluster.
14. The method according to claim 1, further comprising: Identify a plurality of first compounds, wherein the plurality of first compounds are part of a learned manifold of the first compounds; The structure of a manifold is learned by tensor training using partially known properties of the compound, where the partially known properties include known properties of the compound. and A plurality of second compounds are generated, the second compounds being based on the plurality of first compounds, wherein the second compounds are the generated compounds.
15. The method according to claim 1, further comprising: Identify the predefined range of molecular descriptors; The output of the generated compound is obtained from the generative model; and Exclude compounds outside the predefined range to obtain a chemical space with generated compounds within the predefined range.
16. The method of claim 1, further comprising randomly selecting a plurality of generating compounds from a Sammon map covering the chemical space to obtain the hit structure.
17. A method for synthesizing a compound, the method comprising: A report is provided on the chemical structure having the hit structure obtained by claim 1; Select at least one compound from the provided chemical structure, which has specific functional activity against the biological target; and Synthesize at least one of the compounds.
18. A method for verifying the biological activity of a generated compound, the method comprising: To obtain at least one compound synthesized in claim 17; and In bioanalysis, at least one synthesized compound is verified to have specific functional activity for the biological target.
19. One or more non-transitory computer-readable media storing instructions that are responsive to being executed by one or more processors to cause a computer system to perform operations, including performing the method of claim 1.
20. A computer system, including: One or more processors; and One or more non-transitory computer-readable media storing instructions that are responsive to being executed by one or more processors to cause a computer system to perform operations, including performing the method of claim 1.
Citation Information
Patent Citations
Mutual information adversarial autoencoder
US11403521B2
Subset conditioning using variational autoencoder with a learnable tensor train induced prior
US20200090049A1
Methods, systems and non-transitory computer readable media for automated design of molecules with desired properties using artificial intelligence
WO2019018780A1