Computational framework for drug discovery of bioactive therapeutics
The computational framework using STONED and QSAR modeling addresses the lack of directed drug design by generating flavonoid derivatives with enhanced bioactivity and safety, facilitating efficient drug discovery through in-silico methods.
Patent Information
- Application Number
- PCT/CA2025/050402
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-21
- Filing Date
- 2025-03-21
- Publication Date
- 2025-09-25
AI Technical Summary
Current drug design methods lack streamlined approaches for directed drug design based on existing molecule scaffolds, particularly for natural products-like structures with enhanced bioactivities, and there is a need for efficient in-silico processes to guide modification and prediction of bioactivity.
A computational framework utilizing the STONED algorithm for generating novel molecules with guided modifications, combined with QSAR modeling and machine learning models like XGBoost, to predict and optimize flavonoid derivatives for bioactivity and ADMET properties, enabling the generation of compounds suitable for biosynthesis and in vivo evaluation.
The framework efficiently explores chemical space to generate novel flavonoid derivatives with enhanced bioactivity and safety profiles, reducing the need for extensive laboratory experimentation and accelerating drug discovery by providing targeted therapeutic candidates.
Smart Images

Figure CA2025050402_25092025_PF_FP_ABST
Abstract
Description
COMPUTATIONAL FRAMEWORK FOR DRUG DISCOVERY OF BIOACTIVE THERAPEUTICSFIELD
[0001] The present disclosure relates to methods and systems for generating novel molecules and optimizing same.BACKGROUND
[0002] Polyphenols are a diverse group of naturally occurring compounds found in plants, many of which have potential health benefits. Known for their antioxidant properties, many polyphenols have been linked to a range of health advantages, including anti-inflammatory and anti-cancer properties. Among the various subclasses of polyphenols, flavonoids stand out. With their distinct chemical structure, flavonoids contribute to the vibrant colors of many fruits and vegetables. Beyond their role in pigmentation, flavonoids have been associated with numerous health benefits, including anti-inflammatory, anti-cancer, and cardio-protective effects. Their significance in pharmacology is noteworthy, as they have demonstrated potential in the prevention of various diseases. The intricate relationship between polyphenols, particularly flavonoids, and human health underscores the importance of exploring these types of compounds for their pharmacological applications.
[0003] Flavonoids are characterized by a 15 -carbon skeleton structure arranged in two phenyl rings (A and B) and a heterocyclic ring (C). Among the various types of flavonoids, the subclasses include flavones, flavanones, flavonols, flavanols, isoflavones, and anthocyanins. Each subclass possesses unique chemical structures and biological activities, contributing to antioxidant properties and a range of protective effects and health benefits. Understanding the diverse types of flavonoids is desirable for unlocking their full potential in both plant biology and human health. Derivatives of flavones, flavanones, and flavonols, for instance, are often recognized for their ability to modulate several cellular enzyme functions, and their potential therapeutic effects are primarily attributed to their structural configuration. Their chemical behavior and biological activity can be significantly altered through modifications like hydroxylation and methylation on the ring structures. Hydroxylation, the addition of hydroxyl groups to theflavonoid skeleton, typically enhances the compound's polarity, potentially augmenting its antioxidant capacity and interaction with other biomolecules (1, 2). On the other hand, methylation, the substitution of hydrogen atoms by methyl groups, tends to reduce polarity, which might affect their solubility, bioavailability, and interaction with biological targets. These modifications allow for a nuanced tuning of a flavonoid’s properties, showcasing the potential for structural optimization in the quest for targeted therapeutic effects. By understanding and manipulating such chemical modifications, researchers can explore a wider spectrum of bioactive compounds, each with distinct properties and potential applications in disease prevention and treatment.
[0004] In the field of drug discovery and molecular design, computational programs and algorithms have emerged as invaluable tools, transforming the conventional trial- and-error methodology. These advanced tools exploit the capabilities of computational chemistry and bioinformatics to predict and refine molecular structures with remarkable accuracy. For instance, Graph-based methods have gained prominence in the field of cheminformatics and drug discovery. These methods represent molecules as graphs, where atoms are nodes and bonds are edges. Each atom and bond can have associated features, such as atom type, bond type, and more. Additionally, machine learning algorithms have assumed a pivotal role, assimilating knowledge from extensive datasets to predict molecular properties and propose novel compounds. Not only do these tools accelerate the drug discovery process, but they also enhance cost-effectiveness by substantially diminishing the need for extensive laboratory experimentation. One example is Graph Convolutional Networks (GCNs), which have been employed for capturing the structural information of molecules and are used in generative models to create new molecular structures. The idea is to train a model on a dataset of existing molecules and then use the learned patterns to generate novel molecules (3). Another similar tool is GraphAF, which offers efficient parallel computation during training and has the capability to generate chemically valid molecules even in the absence of chemical knowledge rules while achieving 100% validity when such rules are applied. Using reinforcement learning, GraphAF attains a high level of performance in the generaloptimization of chemical properties (4). Other machine learning algorithms that have gained prominence in drug discovery are DeepChem and IBM Watson for Drug Discovery, which leverages artificial intelligence to analyze vast datasets, predict molecular properties, and propose novel drug candidates. De novo drug design programs like LigBuilder and Schrodinger's Maestro platform also allow for the generation of entirely new molecular structures with desired properties.
[0005] The STONED (Superfast Traversal, Optimization, Novelty, Exploration, and Discovery) algorithm has also emerged as a novel tool in the domain of molecular generation, especially in the context of drug discovery (5). At its core, it uses SELFIES (Self-Referencing Embedded Strings) representation, which surpasses the commonly used SMILES (Simplified Molecular Input Line Entry System) representation by alleviating issues of validity and robustness during genetic operations. Unlike SMILES, where molecules are defined as a chain of atoms written as strings with a complex grammar, SELFIES ensures a grammar-less, self-contained representation, significantly reducing the possibility of generating invalid molecular structures during operations like mutation and crossover. SELFIES representation defines separate enclosed tokens within square brackets, thus providing discrete meaningful tokens that are robust to various operations (6, 7, 8, 9). While there are alternative generative methods, such as Variational Autoencoders (VAE) and Generative Adversarial Networks (GAN), that have found applications in molecular generation, STONED stands out for its unique advantages in scenarios with constrained substructure requirements and limited computational resources. Unlike methods that require Graphic Processing Units (GPUs) for training deep learning models, STONED operates efficiently with the machine’s CPU, offering an accessible and cost-effective solution.
[0006] Various programs and algorithms that streamline the complex process of identifying hit candidates to specific biological targets also accelerate the drug discovery process by reducing the reliance on serendipity and intuition. One notable program is AutoDock, widely employed for molecular docking studies to predict how small molecules interact with target proteins. Other molecular dynamics simulations areprograms like GROMACS, and AMBER, which provide insights into the dynamic behavior of biological molecules, aiding in the understanding of their structural changes over time. Virtual screening programs like DOCK and Surflex-Dock sift through large chemical libraries to identify potential ligands that bind to specific targets. QSAR (Quantitative Structure-Activity Relationship) modeling, implemented by programs such as Dragon and MOE, has been extensively used in predicting the biological activity or other properties of molecules based on their chemical structure. It relies on mathematical models to draw correlations between chemical structures and pharmacological activities quantitatively. QSAR is utilized in drug design (10), adhering to the fundamental notion that structural variations in molecules result in differing biological activities. In medicinal chemistry, QSAR is pivotal for several applications, including drug discovery and lead optimization, particularly in scenarios lacking 3D structures of specific drug targets (11). For instance, it's employed in molecular design, biological function assessment, lead compound refinement, virtual screening, and the analysis of drug candidates' toxicity (12). Some biological activities explored in QSAR studies encompass enzyme activity, minimum effective dose, and toxicity. These models are built with molecular descriptors like dipole moment, atomic volume, and the number of aromatic rings. QSAR modeling has been pivotal in validating the antioxidant activity of flavonoids identifying which atoms in a molecule have larger radical scavenging abilities (13). In another study on traditional Chinese medicine, QSAR modeling was instrumental in understanding the role of flavonoids in voltage-gated calcium channels, establishing a statistical correlation between the biochemical features and mechanisms of action for 24 compounds, including Kaempferol and Taxifolin (14). In two other studies, QSAR models were employed to evaluate the anti-oxidant, anti-inflammatory, and anti-cancer activities of flavonoids, among others (15, 16).SUMMARY
[0007] In one example, an apparatus for generating at least one novel molecule having predefined characteristics, comprising: processing circuitry;a memory device coupled to the processing circuitry, the memory device comprising instructions executable by the processing circuitry, wherein the processing circuitry is configured to: receive at least one input molecule, and randomize the at least one input molecule to generate a plurality of randomized molecules; generate a mutation of each of the plurality of randomized molecules iteratively until at least one candidate molecule having a predetermined substructure is obtained; standardize different representations or descriptions of the same molecule to a single, unique form to form a canonical SMILES string; store the at least one candidate molecule in the memory device to form a first training dataset for a neural network; compute feature vector representation of the at least one candidate molecule, wherein the feature vector representation comprises at least one of a descriptor, molecular structure and substructure; with a pretrained model, predict whether the at least one candidate molecule possesses at least one desired property; determine whether the at least one candidate molecule can be synthesized to produce a physical molecule.
[0008] In one example, a method for generating at least one novel molecule, the method comprising the steps of: at processing circuitry executing instructions stored in a memory device, receiving at least one input molecule, and randomizing the at least one input molecule to generate a plurality of randomized molecules; generating a mutation of each of the plurality of randomized molecules iteratively until at least one candidate molecule having a predetermined substructure is obtained; standardizing different representations or descriptions of the same molecule to a single, unique form to form a canonical SMILES string;storing the at least one candidate molecule in the memory device to form a first training dataset for a neural network; computing feature vector representation of the at least one candidate molecule, wherein the feature vector representation comprises at least one of a descriptor, molecular structure and substructure; predicting whether the at least one candidate molecule possesses at least one desired property; and determining whether the at least one candidate molecule can be synthesized to produce a physical molecule.
[0009] In one example, a non-transitory computer readable medium having instructions stored thereon, the instructions being executable by the processing circuitry to configure the processing circuitry to at least: receive at least one input molecule, and randomize the at least one input molecule to generate a plurality of randomized molecules; generate a mutation of each of the plurality of randomized molecules iteratively until at least one candidate molecule having a predetermined substructure is obtained; standardize different representations or descriptions of the same molecule to a single, unique form to form a canonical SMILES string; store the at least one candidate molecule in the memory device to form a first training dataset for a neural network; compute feature vector representation of the at least one candidate molecule, wherein the feature vector representation comprises at least one of a descriptor, molecular structure and substructure; with a pretrained model, predict whether the at least one candidate molecule possesses at least one desired property; and determine whether the at least one candidate molecule can be synthesized to produce a physical molecule.
[0010] In another example, the present disclosure addresses a critical gap in current techniques for drug design, particularly in the realm of directed drug design. Despite the range of existing approaches, there is a notable absence of streamlined methods that offer directed drug design based on existing molecule scaffolds. The novelty of this disclosure lies in its provision of a series of methods for in vitro molecule optimization, specifically targeting the creation of natural products-like structures with enhanced bioactivities, all facilitated by efficient in-silico processes. This contribution to the field emphasizes guided modification of scaffold molecules for bioactivity prediction, thereby streamlining the drug design process. Moreover, the disclosure favors the generation of novel compounds that can be readily produced through simple biosynthesis methods, facilitating their in vivo evaluation, underscoring its potential for broad practical application and further advancement in drug discovery.BRIEF DESCRIPTION OF THE DRAWINGS
[0011] Figure 1 shows a top-level diagram of an overall system architecture upon which any one or more of the techniques for generating novel molecules can be performed;
[0012] Figure 2 shows a flow diagram of the molecular generation pipeline of flavonoid-derivatives;
[0013] Figure 3 shows a flowchart with example steps for molecular generation;
[0014] Figure 4 shows exploration of luteolin derivatives in the chemical space using t-SNE visualization plot; the generated derivatives are shown in the varying coloured clusters depending on structural similarities;
[0015] Figure 5 shows distributions of predicted negative log IC50 values for protein targets histograms showing the predicted negative log IC50 values for (A) 5LO, (B) CASP1, (C) mPGES, (D) PLD1, (E) TRKA, and (F) TRKB protein targets, in which the higher values indicate greater potency against the target;
[0016] Figure 6 shows top-ranking known molecules obtained across the QSAR models, shown in no particular order, in which the molecules’ common names or if no common name exists, their IUPAC names are noted; (A) acacetin, (B) primetin, (C)tectochrysin, (D) 5 ,7-dihydroxy-2-(4-hydroxy-3 -methylphenyl) -4h- 1 -benzopyran-4- one, (E) 2-(2,4-dihydroxyphenyl)-5 -hydroxy chromen-4-one, (F) 2-(3,4- dihydroxyphenyl)-7-methylchromen-4-one, (G) 5,3',4'-trihydroxyflavone, (H) 3', 4'- dimethoxyflavone, (I) 5-methoxy-2-(4-methoxyphenyl)-7-methylchromen-4-one, (J) 2- (3,4-dimethoxyphenyl)-7-methyl-4h-chromen-4-one;
[0017] Figure 7 shows top-ranking molecules obtained across the QSAR models that have not been extensively studied and possess limited associated or published bioactivity data. (A) 5,7-dihydroxy-2-(4-hydroxy-3-methylphenyl)-4h-l-benzopyran-4- one, (B) 2-(2,4-dihydroxyphenyl)-5-hydroxychromen-4-one, (C) 2-(3,4- dihydroxyphenyl)-7-methylchromen-4-one, (D) 5-methoxy-2-(4-methoxyphenyl)-7- methylchromen-4-one, (E) 2-(3,4-dimethoxyphenyl)-7-methyl-4h-chromen -4-one; and
[0018] Figure 8 shows a block diagram of an example machine upon which any one or more of the techniques (e.g., methodologies) discussed herein may be performed. DETAILED DESCRIPTION
[0019] The following detailed description refers to the accompanying drawings. Wherever possible, the same reference numbers are used in the drawings and the following description to refer to the same or similar elements. While embodiments of the disclosure may be described, modifications, adaptations, and other implementations are possible. For example, substitutions, additions, or modifications may be made to the elements illustrated in the drawings, and the methods described herein may be modified by substituting, reordering, or adding stages to the disclosed methods. Accordingly, the following detailed description does not limit the disclosure. Instead, the proper scope of the disclosure is defined by the appended claims.
[0020] Moreover, it should be appreciated that the particular implementations shown and described herein are illustrative of the invention and are not intended to otherwise limit the scope of the invention in any way. Indeed, for the sake of brevity, certain subcomponents of the individual operating components, and other functional aspects of the systems may not be described in detail herein. Furthermore, the connecting lines shown in the various figures contained herein are intended to represent exemplary functionalrelationships and / or physical couplings between the various elements. It should be noted that many alternative or additional functional relationships or physical connections may be present in a practical system.
[0021] Referring to Figure 1, there is shown a top-level diagram of an overall system architecture 20 upon which any one or more of the techniques (e.g., methodologies) discussed herein can be performed. For example, system 20 may be employed for implementing methods for producing defined modifications in selected molecules with the goal of generating a chemical catalog of novel molecules that can be screened for the design of drugs.
[0022] In one example, the method produces novel molecules similar to natural polyphenols such as flavonoids and stilbenes, and includes an adapted STONED algorithm to generate guided mutations. Along with using QSAR modeling, subsequent algorithms are then used to filter the catalog of generated molecules according to their chemical and ADMET properties, as an indication of drugability, and to analyze potential drug-target interactions. The STONED algorithm includes a set of algorithms where a single step of molecular generation is carried out based on initial seed structures. Each of these algorithms uses incremental changes within the SELFIES representation of a molecule, enabling rapid exploration of chemical space without a pre-trained generative model or set of reaction rules. This methodology bypasses the need for large amounts of data and training times usually required by deep generative models and also removes the need for structural validity checks (5, 9, 17). Moreover, the algorithm's design facilitates targeted modifications in molecular structures, as described in the applications for directed drug design presented in the present disclosure.
[0023] The system and methods may be used to generate analogs of polyphenolic compounds known in nature such as flavonoids, anthocyanins, stilbenes, bibenzyls, phenanthrenes, or cinnamic acids, as well as other natural phytochemicals such as terpenoids, including monoterpenes, diterpenes, triterpenes and tertraterpenes (or carotenoids), and alkaloids. For example, the system and methods may be used to generate flavone analogs with a conservative structure that only allows certain functionalgroups, such as hydroxyl groups (-OH) or methyl (-CH3), to be added or substituted onto the structure. Other possible modifications that may be added to a molecule skeleton may include adding other alkyls groups such as ethyl (-C2H5), propyl (-C3H7), butyl (-C4H9), or pentyl (-C5H11); prenyl groups (-C5H9, -C10H14); carbonyl (>C=O); carboxyl (-COOH), acetyl (CH3CO-), amino (-NH2), sulfhydryl (-SH), nitro (-NO2), or halogen (e.g., -Cl, - Br, -I, -F) groups. The presence and location of these functional groups will affect the bioactivity of the generated molecules. As an example, the methylation of flavonoids, via their free hydroxyl groups or carbon atoms, dramatically enhances their metabolic stability and membrane transport, facilitating absorption and greatly improving oral bioavailability (18). Methylation can also reduce their reactivity and augment their antimicrobial activity (19). On the other hand, functional hydroxyl groups in flavonoids mediate their antioxidant effects by scavenging free radicals and / or chelating metal ions, which is crucial for preventing radical generation that damages target biomolecules (20). Furthermore, the presence of a 2'-hydroxyl group in a flavonoid regulates the conformational freedom between its molecular structures, favoring molecule planarity and accounting for its biological properties (21).
[0024] The method also includes QSAR models which may be trained using open- source data from the PubChem Database with data reported in IC50, which is the concentration of a substance needed to inhibit a particular biological process by 50% (22). However, the use of the negative logarithm of IC50 (pIC50) instead of IC50 can be used for QSAR modeling to align with the logarithmic relationship between bioactivities and compound concentration. This transformation yields a uniform data distribution, which is desirable for stable statistical analyses and modeling while offering numerically larger data values, which are beneficial in computational contexts. Additionally, pIC50 encourages a logarithmic view of in vitro assay data, facilitating linear dose-response curve plotting, which is crucial for understanding compound potency and efficacy (23).
[0025] The QSAR models are trained to predict flavonoids,, however since the PubChem datasets comprise diverse molecular scaffolds, in order to attain greater accuracy for predicting on flavonoids, an integrated approach of using scaffold splittingand stratified splitting can be used to balance the training, test, and validation datasets. Scaffold splitting is a technique leveraged in model validation to segregate datasets based on the structural scaffolds inherent in molecules. This method is anchored on an algorithm devised by Bemis and Murcko, which facilitates the grouping of ligands into common frameworks (24). Unlike random splitting, scaffold splitting offers a more realistic and challenging approach. It simulates a real-world scenario where a model, trained on past data, is used to make predictions on new, structurally distinct data that may arise in the future. Conversely, stratified splitting adopts a different approach towards dataset segregation. This method aims at ensuring a uniform distribution of a particular variable, often the target variable, across the various splits of the dataset. In a classification setting, stratified splitting is often chosen to ensure that the train and test sets have a similar percentage of samples for each target class, mirroring the distribution in the entire dataset. This method helps maintain the dataset's statistical characteristics across the different splits, allowing for a more precise assessment of the model's performance (25). Combining scaffold and stratified splitting creates a method where datasets are split based on molecular structures while keeping the target variable evenly distributed across these splits. This balanced method is desirable for training robust models, and combines the benefits of both scaffold and stratified splitting, which improves the model's ability to work well with different molecular structures while keeping the target variable distribution even. This combination is useful in situations like this project, with a wide variety of molecular structures and where the distribution of the target variable is key for training and evaluating models.
[0026] XGBoost (extreme Gradient Boosting) is recognized for its adeptness in handling various QSAR modeling projects due to its precise prediction generation capabilities. XGBoost is particularly suitable for QSAR modeling because it can handle sparse data, manage missing values, and model non-linear relationships inherent in molecular data. One study showcased a comparative analysis, pitting the efficacy of XGBoost against other algorithms like random forest and single-task deep neural nets across 30 in-house QSAR data sets. The findings showed that despite XGBoost harboringnumerous adjustable parameters, it could be fine-tuned to deliver highly accurate predictions without imposing a computational burden (26).
[0027] Generally, the effectiveness and safety of a drug may be determined by based on the Absorption, Distribution, Metabolism, Excretion, and Toxicity (ADMET) properties. These properties show how a substance acts in the body and are important for passing the safety standards needed to move forward in drug development. Data to model the ADMET tasks can be obtained from the Therapeutics Data Commons (TDC), an open-science initiative aimed at advancing the discovery and development of drugs by providing AI / ML-ready datasets. Some of these include the Martins' Blood-Brain Barrier (BBB) dataset, which contains 1,975 compounds, and the AstraZeneca Lipophilicity dataset, encompassing 4,200 compounds (27). In a recent study, researchers employed XGBoost for accurate ADMET prediction by applying an ensemble of features, including fingerprints and descriptors. The XGBoost model showcased significant performance in the TDC ADMET benchmark group. Among 22 tasks, the model ranked first in 18 tasks and was within the top 3 in 21 tasks, highlighting its accuracy in predicting ADMET properties, which are vital for drug discovery.
[0028] Virtual screening is also employed in drug discovery for rapid assessment of vast compound libraries to identify those molecules most likely to have desired therapeutic effects. The essence of virtual screening lies in predicting the binding affinity of multiple molecules to a given target, leveraging molecular similarities. In recent times, the framework DeepPurpose has revolutionized virtual screening by incorporating deep learning methodologies to automate the screening of potential drug-target pairs. Using pre-trained models can effectively rank compounds based on their predicted binding scores, paving the way for more focused experimental validation. Such an approach not only reduces the need for exhaustive lab testing but also accelerates the pace of drug discovery by providing preliminary insights into potential lead compounds.
[0029] The system 20 may be configured to train one or more machine learning models to produce defined modifications in selected molecules with the goal ofgenerating a chemical catalog of novel molecules that can be screened for the design of drugs. In one example, the system 20 may use the trained machine learning models to: predicts ADMET properties on unseen and structurally different drugs; use the features of the molecules to predict the activity or property of interest for different targets; screen molecules against an array of target proteins, or against the paired drug-target combinations; ranking the molecules based on QSAR, ADMET, and molecular properties; and determine the feasibility of producing each molecule for future testing.
[0030] As shown in Figure 1, the overall system architecture 20 comprises a machine 21 with a processing circuitry 22, memory device 23 and input / output interface module 24, interconnected by communications bus 25. In one example, memory device 23 is capable of storing machine executable instructions 26, and data 27, such as, input data, training data, raking data, output data; target data, molecular properties data; machine learning models, data process models, among others. Further, the processor 22 is capable of executing the instructions 26 stored in memory device 23 to implement aspects of processes described herein. For example, processor 22 may be embodied as an executor of software instructions 26, wherein the software instructions 26 may specifically configure the processor 22 to perform algorithms and / or operations described herein when the software instructions 26 are executed. Alternatively, the processor 22 may execute hard-coded functionality. Machine 21 also comprises a graphical user interface (GUI) 28 and database 29 are also coupled to tenant user device 21 via I / O interface module 24.
[0031] System 20 may also comprise data storage 31, which is configured to maintain one or more datasets, including data structures storing linkages and other data, such as medical images, libraries, models, and rules. Data storage 31 may be a relational database, a flat data storage, a non-relational database, among others. Coupled to the machine 21 is one or more databases 32 which may be communicatively coupled directly to the machine 21 or via network 33.
[0032] The machine 21 may be communicatively coupled to another machine 34 via the network 33. Similar to machine 21, the other machine 34 comprises one or moreprocessors 35, memory device 36 storing data 37, including data models and process models, and instructions 38, I / O interface module 39, interconnected by communications bus 40.
[0033] At machine 21, memory device 23 comprise several modules with instructions 26 stored therein which are executable by the processing circuitry 22 for the generation of molecules. The modules may include a random molecular generator module 41; a mutation module 42; a canonical module 44; a data preparation module 45; a prediction module 46; a screening module 48; a ranking module 50; and an annotation module 52. Memory device 23 may also stores data 27,
[0034] Figure 2 shows an overview workflow diagram 100 for a process for generating novel molecules.
[0035] Molecule Generation
[0036] In one example, polyphehol derivaties such as flavonoid-derivatives are generated. In step 102, an input molecule 60, such as luteolin, goes through steps of molecule modification (step 104) and a plurality of potential novel molecules are generated via data processing step (step 106). Next, the properties of the potential novel molecules are computed (step 108), and the descriptors of the potential novel molecules are generated (step 110). In step 112, the ADMET properties are predicted and are used to filter out non-viable candidates. In step 114, the top molecules are chosen based on ranking results from QSAR modeling.
[0037] In more detail, Figure 3 shows a flowchart 200 with example steps for molecular generation using a flavonoid as an example. The choice of molecules in these examples was driven by the desire to examine deeper into the structural intricacies of the flavonoids and explore the potential of optimizing their properties for targeted therapeutic applications. In one example, a luteolin molecule is chosen as the input molecule. In step 202, instructions associated with the random molecular generator module 40 are executed by the processing circuitry 22 to randomize the representation of the molecule 60 and generate a specified number of randomized molecules, e.g. between 100 and 500. Theserandomized molecules are represented in SELFIES, which provides a foundation for ensuring the chemical validity of the molecules throughout the process.
[0038] Next, in step 204 mutated versions of the randomized molecules are generated. Each of these randomized molecules is taken, one by one, and subjected to a mutation process. This process aims to perform a single mutation on each molecule, either substituting a character at a random or specified index with a character from a defined alphabet. As an example scenario, the defined alphabet may consist of hydroxyl and methyl groups, performing either substitution or addition mutations. The mutation process is designed with a check -and-balance mechanism to ensure the preservation of the flavone substructure in each mutated molecule. Accordingly, the mutation process is performed up to a predetermined number of times until a valid mutation is found. Each attempt at mutation is scrutinized for substructure violations against a defined substructure pattern. This check ensures that the flavone scaffold remains intact in each mutated molecule. The mutated molecules that pass the substructure check are collected for further processing. The SMILES strings of these valid mutated molecules are updated in a set of unique mutants and a dictionary of similarity scores to the input molecule. This data is then recorded in files for further analysis. The orchestration of the above steps is done through iterations. In each iteration, a new set of randomized molecules is generated, which then go through the mutation process. The data collection is updated with any new unique mutants found in each iteration. This iterative process continues for a specified number of iterations, each time expanding the molecular space with new unique mutants centered around the starting molecule. When all iterations have been completed, the total count of unique mutants obtained is presented. All the unique mutants, along with their similarity scores, are saved to files, providing a comprehensive record of the molecular space generated around the molecule. Through this process, a comprehensive molecular space centered around the molecule scaffold is generated. This process, thus, provides a systematic and controlled approach to exploring and understanding the molecular space of luteolin, generating a wealth of data for further analysis and discovery.
[0039] In step 206, instructions are executed by the processing circuitry 22 to improve the large number of results produced. The instructions may comprise a modified STONED algorithm for molecular generation, which is be adapted to tailor a series of modifications to address specific constraints and requirements as needed. As such, the adapted STONED algorithm allows for tailored generation of molecular structures that are not only novel but also feasible for testing. Accordingly, STONED algorithm is modified to preserve a scaffold of the molecules generated. In this tailored molecular generation process, initially, the necessary configurations, such as the sample size and number of iterations, are defined from a configuration file. The modified STONED algorithm for molecular generation starts creating molecules from different starting points, leading to a wide variety of valid SMILES strings. These strings, while being valid and unique, could potentially represent a single molecular entity. The process starts with the combining of SMILES strings from the generation portion of the pipeline, accumulating a substantial count of over 500,000, over one million, over 1.5 million, and over 2 million strings to form a SMILES strings dataset.
[0040] Due to the inherent multiplicity of SMILES strings for a singular molecular structure, a function that canonicalizes the SMILES is applied to the SMILES strings dataset. Accordingly, instructions associated with the canonical module 44 are executed by the processing circuitry 22 to standardize different representations or descriptions of the same molecule to a single, unique form, typically known as a canonical SMILES string. In one example, a RDKit toolkit which distills the data down to a unique molecular set of about 10,000, about 20,000, about 30,000, or about 50,000 canonical SMILES strings may be employed.
[0041] Data Acquisition
[0042] In step 208, following the establishment of canonical SMILES strings, the system 20 gathers molecule and drug data from a plurality of sources for input into the training engine 54. For example, the system may obtain pharmaceutical data from one or more databases or data sources, such as the chEMBL database, and the PubChem database, the Therapeutics Data Commons (TDC), as non-limiting examples. Forexample, input data may include QSAR datasets for QSAR modelling, in which it is desirable to obtain, clean, and preprocess bioactivity data for a target protein. In one example, an RDKit-powered Python library, such as Datamol, is used for data acquisition.. The data gathering starts by identifying the UniProt accession number for the target protein of interest. Then, using the ChEMBL web interface, the UniProt accession number or the name of the target protein is used as an input to search the database. The results page, under the 'Targets' tab, displays a list of compounds that have been tested against the target, along with their bioactivity data. Using the database fdters, the list of compounds is narrowed down based on their IC50 values. For example, setting a range to view compounds with IC50 values less than 100 nM, less than 50 mM, or less than 10 nM. Once satisfied, the data is downloaded, which encompassed molecules, their IC50 values against the target protein, and additional chemical and assay information.
[0043] In addition, various therapeutic tasks and datasets, including target discovery, activity modeling, efficacy and safety, and manufacturing, may be gathered. For examples, such datasets may be used ADMET predictions. For example, the TDC offers a unified benchmark for a fair comparison between different machine learning models. For each ADMET prediction task, the TDC divides the dataset into 80% for training and 20% fortesting using a scaffold split. This simulates a real -world scenario where a trained machine -learning model predicts ADMET properties on unseen and structurally different drugs. The properties evaluated under ADMET include absorption, distribution, metabolism, excretion, and various toxicities. These properties play a crucial role in determining the drug efficacy and safety, with parameters such as octanol / water partition coefficient, solubility, dissociation constant, intestinal absorption, Caco-2 permeability, human bioavailability, plasma protein binding, volume of distribution, metabolism, halflife, excretion, urinary excretion, clearance, and toxicity being among the key ADMET properties assessed.
[0044] Data Processing
[0045] In step 210, once a dataset is obtained following the above-noted data acquisition steps, instructions associated with the data preparation module 45 areexecuted by the processing circuitry 22 to receives the datasets and clean and pre-process the datasets. In one example, the instructions comprises an algorithm for adjusting the nitrogen aromaticity, and for correcting faulty valence issues in nitrogen-containing aromatic rings. The molecules are validated by converting them from mols to smiles and back to the mol format. This conversion helps rectify inconsistencies in molecular representations, and also addresses any valence discrepancies arising from incorrect atomic charges; however, this step does not remove charges but rather corrects valence issues. In one example, a Sanifix algorithm within Datamol is used.
[0046] Next, the data is standardized by creating canonical SMILES for the molecules to achieve a uniform representation of each one. If some compounds are reported in salt forms or are associated with metal ions, such metal ions are disconnected thereby focussing on the organic part of the molecule, which is typically the primary concern in drug discovery. The dataset is assembled from multiple sources, so further processing may be required, for example, removing columns not needed for downstream machine learning and calculating or converting values. The datasets are first restructured by renaming relevant columns for clarity, and their molecular weight is calculated using SMILES strings. Next, any values in the 'IC50' column that do not exactly match the given standards are adjusted by introducing random noise. This is done primarily to larger values to introduce some variance so that the machine learning model would perform better. Then, the negative log IC50 for the IC50 value column is calculated. Lastly, scaffolds are generated for the molecules, and any singleton scaffolds, which are unique or rare scaffolds appearing only once in a dataset, are filtered out. After all these transformations, an exploratory data analysis (EDA) is conducted to ensure data quality. EDA helps in understanding the underlying structure of the data, spotting anomalies, checking assumptions, and ensuring the dataset is ready for downstream machine learning analyses.
[0047] In step 212, the datasets for each of the ADMET prediction tasks are divided into 80% fortraining and 20% fortesting using a scaffold split. In one example, a function is applied to the 'smiles' column of the QSAR dataset to identify molecules containing aspecific structure known as the scaffold. A new column is generated, indicating whether each molecule contains this scaffold. Following this, the data is split into two sets: a training set and a temporary set. 70% of the data is allocated to the training set, while the remaining 30% is placed in a temporary set. This split is stratified based on the presence of the scaffold to ensure a balanced representation of scaffold-containing molecules in both sets. Subsequently, the temporary set is further divided into validation and test sets. Of the remaining data, 20% is allocated to the validation set and 10% to the test set. Similar to the primary split, this division is also stratified to maintain a balanced representation of flavone -containing molecules. As a result of the workflow, three datasets are created: The training set, validation set, and test set. The training set, comprising 70% of the original data, is used to train the machine learning (ML) model. The validation set, making up 20% of the original data, is utilized to tune the model parameters and provide an unbiased evaluation of model fit during the training phase. Lastly, the test set, which comprised 10% of the original data, is used to provide an unbiased evaluation of the final model fit, ensuring no overlap with the data used in the training and validation phases. This structured approach allows for a fair representation of the key molecular scaffold of interest across all three datasets, which is desirable for unbiased training, validation, and testing of the QSAR model.
[0048] Feature Generation
[0049] Following the dataset standardization, the SMILES data is converted into RDKifs molecular objects (mols). In step 214, instructions associated with the feature generation module 47 are executed by the processing circuitry 22 to extract a particular set of features from extract a particular set of features from the molecular objects and group the extracted features, such as in one or more vectors, to generate the training data. In one example, the instructions comprise a featurization function which iterate over a list of mols, calculating a feature set for each mol, using six featurizers from DeepChem. Once the featurization function is initiated, the first of the six feature sets to be computed is the MACCS key, employing 166 predefined structural keys to generate a binary string reflective of a molecule's inherent structural nuances, encapsulating the presence orabsence of specific substructures. The Extended Connectivity Fingerprints (ECFPs), formulates a bit vector by segmenting a molecule into its circular neighborhoods, capturing the environment around each atom. Mol2Vec fingerprints, inspired by the word2vec method in natural language processing, use an unsupervised machine learning technique to create vector representations of molecules, treating each fragment of a molecule as a "word" and the entire molecule as a "sentence." PubChem fingerprints, which consist of 881 structural keys, provide a comprehensive representation of substructures and molecular features. Mordred descriptors are comprehensive, calculating over 1800 molecular characteristics, including parameters like aromatic atom counts and topological polar surface area. Lastly, RDKit descriptors, derived from the RDKit suite of cheminformatics tools, offer a variety of chemical descriptors, such as molecular weight and radical electron counts. For each molecule, these fingerprints and descriptors are computed. After computation, they are concatenated to create a combined feature set for every molecule. Given the data's high -dimensional nature, some values may end up being undefined (NaN) or approaching infinity. These anomalies are managed by replacing NaN values with zeros and capping excessively large values. In the final step, the generated features and the corresponding target values, specifically 'IC5 O' values, are saved to the designated directory. This ensures that the featurized data are accessible for future QSAR modeling steps, whether that be data normalization, model training, or validation.
[0050] QSAR Modeling
[0051] In step 216, instructions are executed by the processing circuitry 22 to determine the optimal hyperparameters for prediction model, such as a XGBoost regressor model which is desirable given its implementation of gradient-boosted decision trees designed for speed and performance. For example, a hyperparameter optimization framework, such as Optuna from Preferred Networks, Japan, may be used. Generally, the hyperparameter optimization finds a set of hyperparameters that minimizes or maximizes a predefined objective function, which, in this case, is defined to evaluate the QA2 statistic during optimization. The QA2 statistic serves as a quantitative measure of the model'spredictive capability. A 5-fold cross-validation scheme is deployed to evaluate the performance of different hyperparameter combinations. Cross-validation is a technique used to assess how the results of a statistical analysis will generalize to an independent data set. It is crucial in preventing overfitting and ensuring the model's ability to generalize well to unseen data. In the case of 5 -fold cross-validation, the data is partitioned into five equal subsets, with each subset serving as the validation set once while the rest serve as the training set. Optuna is configured to maximize the QA2 statistic, hence focusing the search towards hyperparameters that would yield higher predictive accuracy. The search is performed over a predefined range of values for each hyperparameter, ensuring a thorough yet focused search in the parameter space. Upon completion of the optimization process, the best hyperparameters that yield the highest QA2 value, are saved as a text file on the file system for future reference and usage.
[0052] Model Training and Evaluation
[0053] Post hyperparameter optimization, in step 218, the training engine 54 may be configured to train a model for molecular generation. Generally, the training data set and the feature vectors are used to fully train one or more predictive models. In one example, different machine learning classifiers or algorithms are used for building the predictive models, such as, supervised learning algorithms, unsupervised learning algorithms and reinforcement learning algorithms. Examples of supervised learning algorithm systems include support vector machine, decision tree, linear regression, logistic regression, naive Bayes, Uncarcst neighbor, random forest, AdaBoost, XGBoost, and neural network methods. Examples of unsupervised learning algorithm systems include K-means, mean shift, affinity propagation, hierarchical clustering, DBSCAN (density-based spatial clustering of applications with noise), Gaussian mixture modeling, Markov random fields, ISODATA (iterative self-organizing data), and fuzzy C-means systems. Examples of reinforcement learning algorithm systems include Maja and Teaching-Box systems. Generally, training the predictive models involves optimizing the parameters of a predictive system to minimize the loss function. In addition to the training step, the predictive models also undergo validation using test datasets.
[0054] As such, in one example, the XGBoost regressor model is trained to use the best hyperparameters obtained. The trained model is then saved to the file system for future use, especially for making predictions on new data. The evaluation phase starts with making predictions on the validation and test sets. The model's performance is evaluated using various metrics, including Root Mean Square Error (RMSE), Mean Absolute Error (MAE), Pearson Correlation, RA2, and Concordance Correlation Coefficient (CCC). These metrics provide different lenses through which the model's predictive performance can be assessed. Subsequently, Y-scrambling is performed to evaluate the robustness of the model. Y-scrambling is a technique used to check the model's dependency on the structure of the data by randomly permuting the target values and re-training the model. A comparison of the performance before and after Y- scrambling provides insights into the model's robustness and the likelihood of overfitting. The Y-scrambled model should perform quite poorly when the Y values are scrambled. The model performing well suggests that the model did not train but rather memorized the values. The feature importance is calculated for the different types of features used in training the model. Understanding feature importance helps interpret the model and understand which features drive the predictions. The results of the feature importance analysis can be saved to a text file for documentation and further analysis.
[0055] QSAR Predictions
[0056] In step 220, instructions associated with the prediction module 46 are executed by the processing circuitry 22 to determine how each molecule would theoretically perform in an assay specific to a given protein target. For the QSAR predictions, a dedicated function is defined to use the trained model. The QSAR prediction function is designed to load the featurized data; however, if the featurized data is not available, it re- featurizes the data. Featurization is a crucial step in QSAR modeling as it translates molecular structures into numerical values or features that can be used by the machine learning model. Once the featurized data is loaded or generated, predictions are made using the trained model for different targets. The model uses the features of the molecules to predict the activity or property of interest for different targets. The predictions arecollated and saved to a DataFrame, which provides a structured and tabular format for storing and analyzing the prediction results. This process facilitates efficient QSAR predictions using these trained models to determine how each molecule would theoretically perform in an assay specific to a given protein target.
[0057] ADMET predictions
[0058] In step 222, instructions associated with the prediction module 46 are executed by the processing circuitry 22 to determine the efficacy, safety, and manufacturing feasibility for each molecule. The model training for ADMET predictions is similar to the process for the QSAR modeling. In one example, the model training is initiated by executing a script that begins by importing necessary libraries and data from the TDC benchmark group. The data associated with the selected property is already split into training / validation and test sets. An XGBoost regressor object is initialized with the 'gpu_hisf tree method to leverage GPU acceleration. A set of hyperparameters specifically tailored for a particular task is defined - these parameters are taken from an open-source repo by Tao Research Group2. The parameters are fed into the XGBoost regressor to perform a randomized search for the best hyperparameters on the training / validation data. The best set of hyperparameters is then saved. To build a more robust model, training can be carried out five times with different random states but with the same optimal hyperparameters. In each iteration, the model is trained on the training / validation data, and predictions are made on the test data. These predictions are collected into a list for later evaluation. Simultaneously, the feature importance of each model is extracted and collected. The models are then saved to the file system for potential later use. The script culminates in the evaluation of the model's predictions against the ground truth in the test data using the evaluation function provided by the TDC library. The mean feature importance across the five models is computed for different types of features, which provided insights into how much each type of feature contributed to the model's predictions. This protocol leads to a well-tuned model, ready for making ADMET predictions, with insights into the importance of different feature types.
[0059] Virtual Screening
[0060] In step 224, instructions associated with the screening module 48 are executed by the processing circuitry 22 to virtual screen a set of molecules against an array of target proteins, or against the paired drug-target combinations. In one example, screening module 48comprises a DeepPurpose framework, whose primary objective is to expedite drug discovery by automating the screening of potential drug-target pairs, thereby obtaining predicted binding scores. This automation serves as a tool for identifying promising drug candidates without engaging in labor-intensive laboratory testing. The script carries out a systematic virtual screening process on a set of molecules against an array of target proteins, making use of the pre-trained models from DeepPurpose. When running the virtual screening script, a step-by-step process is followed using each pretrained model from DeepPurpose. For each model, protein sequences are gathered and carefully matched with the molecules. After matching, a pre-trained model is loaded to start the virtual screening against the paired drug -target combinations. In the last step, the results from the virtual screening are collected into a data table. Virtual screening is done through a tool like DeepPurpose to automate the screening of an extensive space of potential drug -target pairs. This automation is a solution to cost and time efficiency, significantly accelerating the drug discovery process.
[0061] Ranking
[0062] In step 226, instructions associated with the ranking module 50 are executed by the processing circuitry 22 to rank and filter molecules based on a predefined criteria. In one example, the ranking procedure for molecules is a multi-step process that ranks and filters molecules using a combination of statistical measures, sensitivity analysis, and heuristic ranking. Initially, columns are defined for QSAR, ADMET, and molecular properties, which are crucial for the ranking process. These columns contained values pertinent to the drug-likeness and efficacy of the molecules. Following this, the QSAR values are ranked, and weights are computed based on two metrics, Test CCC and Q2, to derive a weighted score for each molecule across various assays. Separate lists of top molecules for each assay are generated based on the highest scores and compiled into a data frame. As the procedure progresses, different ranking strategies like Weighted BordaCount, Weighted Average Rank, Statistical Measures, Sensitivity Analysis, and Heuristic Rank are applied to the ranked QSAR values, each offering a unique perspective on the relative standing of the molecules. A consensus ranking is then computed as an average of the various ranking measures to provide a balanced view of the molecule standings. The computed ranks are appended to the original dataset, and the top 100, the top 200, the top 300, or the top 500 molecules based on the consensus rank can be selected. These top molecules are saved to a CSV fde for further analysis. In the subsequent phase, molecules from the primary ranking are matched with another dataset to find common molecules. The datasets are filtered, consolidated, and deduplicated to ensure a unique set of molecules is obtained. A check against a molecule database is performed to differentiate between existing and unknown molecules. Drug -likeness evaluations are carried out on both existing and unknown molecules to gauge their pharmaceutical potential. Finally, the data frames containing the existing and unknown molecules, complete with their drug-likeness evaluations, are outputted for further analysis. This ranking and filtering protocol supplies a thorough examination of the molecules and a robust groundwork for identifying potential drug candidates. The consensus ranking combines different ranking philosophies to mitigate biases and offer a well-rounded view of the molecule’s standings. Furthermore, the drug -likeness evaluations and database checks add a layer of validation to the rankings, aiding in the careful selection of molecules for downstream analyses or experiments.
[0063] Annotating
[0064] Following the virtual screening process, in step 228 instructions associated with the annotation module 52 are executed by the processing circuitry 22 to determine the likelihood of success of the molecule with respect to production of same. In one example, the top 100, or the top 200, or the top 300, orthe top 500 molecules are subjected to manual selection. Using RDKit, visual representations of each molecule are generated for further examination. This step is to determine the feasibility of producing each molecule for future testing. Next, a method is used to check if these molecules are listed on PubChem, a database of chemical molecules, and to get their identifiers from there.This is carried out using Python and asynchronous programming to efficiently query the PubChem database. The code uses the aiohttp library for handling HTTP requests and asyncio for managing the asynchronous execution of these requests. A function is defined within the script to process the molecules in batches, using each molecule's SMILES notation to construct the query URL. The code aims to fetch the PubChem CID (Compound Identifier) for each molecule, marking molecules as "unknown" if they are not found in the database. Upon successful retrieval of PubChem CIDs, manual online searches are conducted to gather more information regarding the properties, uses, assay data, and potential research implications of each molecule. This additional step provides a broader understanding and verification of the data that can be obtained from PubChem, ensuring a detailed analysis of the molecular candidates. Lastly, a drug-likeness evaluation is performed on the molecules using a defined function within the script. This function uses various drug-likeness rules such as Lipinski's Rule of Live, Veber's rules, Ghose filter, and a subset of Muegge's criteria to identify potential issues that could affect the drug development process. The function checks the molecular weight, the number of hydrogen bond donors and acceptors, the logarithm of the octanol-water partition coefficient (LogP), the total polar surface area, and the number of rotatable bonds among other parameters, and logs any violations of the rules. This evaluation further refines the list of molecules, ensuring the selected candidates adhere to accepted drug-likeness principles, thus increasing the likelihood of success.
[0065] EXAMPLES
[0066] Example 1
[0067] This example describes the assembly of flavonoid derivatives. Since the first derivatives of interest only have a few possible combinations of mutations and functional groups, it is necessary to cluster the molecules based on their chemical properties, figure 4 is a visualization diagram of the derivatives obtained from the flavones luteolin and its derivatives in the chemical space using the t-SNE visualization plot, where those derivatives with similar structural similarities are clustered together. This enables the identification of the most promising derivatives for further analysis.
[0068] Example 2
[0069] This example describes QSAR Distributions. In cheminformatics, understanding the distribution of bioactivity scores, such as IC50, is important. The distribution of negative log IC50 values is chartered for a set of molecules, spanning 13 distinct protein targets. Analyzing these distributions holds two-fold significance. First, it offers insight into the inherent bioactivity landscape of the generated molecules. Second, by contrasting these distributions against those of the datasets employed to train our predictive models, there is a clear picture of how the generated molecule scores align or deviate from the training set. This comparison not only validates the predictions made by the trained models but also provides a deeper understanding of the potential bioactivity profile of the generated molecules. In Figure 5, distributions of QSAR models for five protein targets are shown, 5 -Lipoxygenase (5LO), caspase 1 (CASP1), microsomal prostaglandin E synthase-1 mPGES (mPGES), phospholipase DI (PLD1), and tyrosine protein kinase receptors A and B (TRKA and TRKB) with the dataset of molecules generated in Example 1. For instance, the enzyme 5LO has been explored for its interactions with flavonoids, which have been shown to inhibit its activity. Inhibition of leukotrienes formation in human neutrophils delineates a clear structure -activity relationship concerning 5-LO inhibition. Other computational approaches and comparative molecular analysis have also been utilized to delve into the inhibitory activity of flavonoid inhibitors against 5-LO. Some flavonoids have also been identified to interact with the CASP1 protein both in vivo and in vitro. Despite the previously documented flavonoid-induced caspase activation, certain inhibitory flavonoids suggest a potential non-classical apoptosis pathway, that is not dependent on caspase activity. Given their antioxidant properties, flavonoids may offer protective effects against conditions where inflammation plays a pivotal role, thus potentially interacting with caspase-1 in inflammatory scenarios.
[0070] Example 3
[0071] This example describes known molecules from Example 2 that are ranked well across the QSAR models and are known to exist in public databases. Figure 6 shows thetop ten known ranking molecules. However, due to the high similarity between most structures, it is not possible to computationally distinguish which one is preferred over the other molecules. Out of these top-ranked molecules, five of them can be distinguished by available information, (i) Acacetin is a known flavonoid exhibiting bioactive properties. It has been investigated for potential anti-inflammatory, anti-cancer, and antioxidant properties. Acacetin can affect plant growth and photosynthesis light reactions as observed in a study with Lolium perenne, Echinochloa crus-galli, and Physalis ixocarpa seedlings (28). (ii) Primetin is a natural flavonoid present in the flowers and leaves of the primrose plant, Primulae veris. In terms of bioactivity, primetin has shown promising antioxidant, anti-inflammatory, and anticancer properties, though research on primetin is still in its early stages. In vitro studies reveal that it can scavenge free radicals, inhibit the production of pro-inflammatory cytokines, and induce apoptosis in cancer cells. Furthermore, in vivo studies have demonstrated its potential in reducing oxidative stress and inflammation in animal models for various diseases such as diabetes, arthritis, and Alzheimer's disease. The safety and toxicity of primetin are yet to be fully evaluated, although it's shown a low toxicity profile in animal models even at high doses, (iii) Tectochrysin is a flavonoid compound known for its wide array of bioactive properties. Its anticancer activity is notably significant, as demonstrated through its antiproliferative effects on human colorectal adenocarcinoma cells, achieved by inducing the expression of DR3, DR4, Fas, and proapoptotic proteins while inhibiting the activity of NF-KB . It also has shown anti-inflammatory properties, although the exact mechanisms remain to be fully elucidated. The compound also exhibits notable anti -Alzheimer activity, particularly in alleviating Api-42-induced impairments of spatial learning and memory in mice, hinting at a promising application in Alzheimer's disease management. Its anti-leishmanial, antioxidant, and anti-diarrheal activities further broaden its therapeutic scope, catering to a variety of health conditions including Leishmaniasis and gastrointestinal issues. Tectochrysin is abundant in Chinese medicine from Alpinia oxyphylla Miq., which has been associated with a plethora of health benefits (29-32). (iv) 5,3',4'-Trihydroxyflavone demonstrates significant larvicidal activity as per a studyexploring its molecular action on the mosquito enzyme glutathione S -transferase Noppera-bo (Nobo) in Aedes aegypti. Alongside 7,3'4'-trihydroxyflavone, another luteolin derivative, it exhibited inhibitory action against AeNobo, comparable to luteolin's effect. This activity suggests potential larvicidal implications against mosquito larvae, possibly by interfering with the crucial enzyme for ecdysone biosynthesis, an insect steroid hormone. This research emphasizes the value of flavonoids in mosquito control, aligning with the pressing demand for safe and efficacious larvicidal agents, thereby broadening the scope of bioactive flavonoids in pest control applications, (v) 3', 4'- Dimethoxyflavone is a lipophilic flavone, which means it's soluble in lipids or fats. It can be isolated from the leaves of Primula veris and has various biological activities. It's known to reduce the synthesis and accumulation of poly (ADP-ribose) polymerase (PARP) and protect cortical neurons against cell death. This molecule acts as an aryl hydrocarbon receptor antagonist in human breast cancer cells, which could potentially play a role in cancer treatment. Moreover, 3 ',4' -Dimethoxyflavone can promote the proliferation of human hematopoietic stem cells, which are the stem cells that give rise to other blood cells. Among its biological activities, it's known to exhibit antioxidant, anticancer, anti-inflammatory, anti-atherogenic, hypolipidaemic, and neuroprotective or neurotrophic effects-. Further, the molecule is a good substrate for CYP1B1, leading to O-demethylation and ring-oxidation products, which could potentially affect its activity or the activity of other compounds when co-administered (34).
[0072] Example 4
[0073] This example describes uncharacterized existing molecules. The structures of the top-ranking uncharacterized molecules from Example 2 that were described in the literature but with no information regarding their bioactivity or other relevant properties are shown in Figure 7. Future investigations should consider more analysis on analyzing and determining the characteristics of these compounds. Likewise, the QSAR modeling analysis also produced a short number of unknown molecules.
[0074] Example 5
[0075] This example describes the assessment of QSAR model results for the distinct protein targets using multiple metrics. This evaluation consisted of an evaluation of QSAR models spanning 13 distinct protein targets, emphasizing their performance on a validation set. For the sake of rigor and clarity, multiple metrics were employed: Root Mean Square Error (RMSE), Mean Absolute Error (MAE), Pearson Correlation, coefficient of determination (R2), and Concordance Correlation Coefficient (CCC). The performance results of QSAR modeling for the different assays are noted in Table 1.
[0076] Table 1
[0077] Among the QSAR models, several assays emerged as particularly strong performers. The MAGL model had a high Pearson Correlation of 0.92, coupled with an R2value of 0.8435. Similarly, the LCK model showed its efficacy, with a Pearson Correlation of 0.9052 and an accompanying R2of 0.8187. The TRKB model, with its Pearson Correlation of 0.8743 and an R2of 0.7613, further validated the robustness ofthe QSAR modeling approach adopted in this study. However, the study also highlighted areas that warrant further scrutiny. The FLAP assay, for instance, presented a modest R2value of 0.1774. This relatively low coefficient suggests that the model, in its current state, may not be capturing all the underlying data intricacies for this particular assay. A similar situation is seen in the TYK2 model, with an R2of 0.5139, indicating potential areas of improvement. Whether these limitations stem from feature selection, lack of training data, model architecture, or the inherent complexity of the target protein remains a topic for further exploration. A summary of key evaluation metrics used for assessing model performance is shown in Table 2.
[0078] Table 2
[0079] This table shows the full form of each metric alongside a detailed explanation of their significance and ideal values for optimal model performance. These metrics offer insights into the model's ability to accurately predict and represent the underlying data trends. The metrics presented contain both validation and test datasets, highlighting the Root Mean Square Error (RMSE), Mean Absolute Error (MAE), Pearson Correlation, coefficient of determination (R2), and Concordance Correlation Coefficient (CCC). Additionally, Y-scrambling metrics are provided to assess model robustness and validity by comparing the model's performance with that of a randomized dataset. The relative importance of different molecular descriptors (maccskeys, circular, mol2vec, mordred,and rdkit) used in XGBoost-based QSAR models for 13 distinct targets are discussed in Table 3. The percentages represent the contribution of each descriptor to the model's predictive capability, thereby providing insights into which molecular properties are most influential for each target's prediction.
[0080] Table 3
[0081] Example 6
[0082] This example describes the ADMET modeling results. The predicted ADMET property values for the three selected molecules, tectochrysin, genkwanin, and hydroxygenkwanin, were assessed to evaluate their physicochemical characteristics, absorption, distribution, metabolism, excretion, and potential toxicity profdes. The results of this prediction analysis are shown in Table 4.
[0083] Table 4
[0084] Regarding their physicochemical properties, all three molecules displayed similar physicochemical profiles with molecular weights ranging from 268.07 to 300.06 kg / mol. The number of heteroatoms increased progressively from tectochrysin (4) to hydroxygenkwanin (6), as did the number of hydrogen acceptors (HAs) and hydrogen donors (HDs). Each molecule possessed 2 rotatable bonds and 3 rings. The logarithm of the octanol-water partition coefficient (log KOW) showed a decreasing trend from tectochrysin (3.17) to hydroxygenkwanin (2.59). Regarding their absorption properties,the Caco-2 permeability predictions were negative for all molecules, with hydroxygenkwanin exhibiting the lowest permeability (-5.29 log(cm / s). The human intestinal absorption (HIA) was consistent across the molecules at 74.15%. P- glycoprotein (Pgp) inhibition decreased slightly from tectochrysin to hydroxygenkwanin. The log D7.4 values, representing the distribution between the organic phase and the buffer phase at pH 7.4, decreased marginally across the molecules. Aqueous solubility predictions were in the same range, with genkwanin displaying the highest solubility (- 4.65 log(mol / L)). Oral bioavailability was highest for tectochrysin (44.61%) and lowest for hydroxygenkwanin (42.87%). In relation to their distribution, blood-brain barrier (BBB) penetration predictions declined for hydroxygenkwanin (29.08%) compared to tectochrysin (35.44%). The plasma protein binding rate (PPBR) was highest for genkwanin (48.33%) and lowest for tectochrysin (45.45%). The volume of distribution at steady state (VDss) was highest for hydroxygenkwanin (3.91 L / kg). Regarding their metabolism, all three molecules were predicted to be moderate to strong inhibitors of cytochrome P450 enzymes (CYP2C9, CYP2D6, and CYP3A4), with genkwanin and hydroxygenkwanin showing higher inhibition percentages for CYP3A4. Similarly, all were substrates for these enzymes, with tectochrysin being the most likely substrate for CYP2C9 and CYP2D6, while hydroxygenkwanin was the most likely substrate for CYP3A4. On excretion, the predicted half-life was longest for hydroxygenkwanin (66.92 hr) and shortest for tectochrysin (45.14 hr) . Clearance rates in hepatocytes (CL-Hepa) and microsomes (CL-Micro) were highest for tectochrysin and lowest for hydroxygenkwanin. On toxicity, the potential to block the human Ether-a-go-go-Related Gene (hERG) channels, a predictor of cardiotoxicity, was similar across the molecules. Mutagenicity, as indicated by the Ames test, was highest for tectochrysin (47.69%) and lowest for genkwanin (43.5%). The drug -induced liver injury (DILI) potential was highest for hydroxygenkwanin (51.04%). The LD50 values, indicative of acute toxicity, were lowest for genkwanin (2.28 -log(mol / kg) and highest for tectochrysin (2.53 -log(mol / kg). In summary, while the three molecules shared similar physicochemical properties and absorption profiles, they exhibited some variations in their metabolism, excretion, andtoxicity profiles. These variations could influence their potential therapeutic applications and safety margins.
[0085] Figure 8 illustrates a block diagram of an example machine 300, such as machine 21 or another machine 34, upon which any one or more of the techniques (e.g., methodologies) discussed herein may be performed. Machine 300 (e.g., computer system) may include a hardware processor 301 (e.g., a central processing unit (CPU), a graphics processing unit (GPU), a hardware processor core, or any combination thereof), a main memory 302 and a static memory 306, connected via an interlink 303 (e.g., link or bus), as some or all of these components may constitute hardware for systems or related implementations discussed above.
[0086] Generally, the hardware processor 301 may, for example, include at least one of a Central Processing Unit (CPU), a Reduced Instruction Set Computing (RISC) Processor, a Complex Instruction Set Computing (CISC) Processor, a Graphics Processing Unit (GPU), a Digital Signal Processor (DSP), a Tensor Processing Unit (TPU), a Neural Processing Unit (NPU), a Vision Processing Unit (VPU), a Machine Ueaming Accelerator, an Artificial Intelligence Accelerator, an Application Specific Integrated Circuit (ASIC), a Field Programmable Gate Array (FPGA), a Radio-Frequency Integrated Circuit (RFIC), a Neuromorphic Processor, a Quantum Processor, or any combination thereof. A processor circuit may further be a multi-core processor having two or more independent processors (sometimes referred to as "cores") that may execute instructions contemporaneously. Multi-core processors contain multiple computational cores on a single integrated circuit die, each of which can independently execute program instructions in parallel. Parallel processing on multi -core processors may be implemented via architectures like superscalar, VUIW, vector processing, or SIMD that allow each core to run separate instruction streams concurrently. A processor circuit may be emulated in software, running on a physical processor, as a virtual processor or virtual circuit. The virtual processor may behave like an independent processor but is implemented in software rather than hardware.
[0087] Specific examples of main memory 302 include Random Access Memory (RAM), and semiconductor memory devices, which may include storage locations in semiconductors such as registers. Specific examples of static memory 306 include nonvolatile memory, such as semiconductor memory devices (e.g., Electrically Programmable Read-Only Memory (EPROM), Electrically Erasable Programmable Read-Only Memory (EEPROM)) and flash memory devices; magnetic disks, such as internal hard disks and removable disks; magneto-optical disks; RAM; or optical media such as CD-ROM and DVD-ROM disks.
[0088] The machine 21, 34 may further include a display device 310, an input device 312 (e.g., a keyboard), and a user interface (UI) navigation device 314 (e.g., a mouse). In an example, the display device 310, input device 312, and UI navigation device 314 may be a touch-screen display. The machine 21, 34 may include a mass storage device 316 (e.g., drive unit), a signal generation device 318 (e.g., a speaker), a network interface device 320. The machine 21, 34 may include an output controller 323, such as a serial (e.g., universal serial bus (USB), parallel, or other wired or wireless (e.g., infrared (IR), near field communication (NFC), etc.) connection to communicate or control one or more peripheral devices (e.g., a printer, card reader, etc.).
[0089] The mass storage device 316 may comprise a machine-readable medium 322 on which is stored one or more sets of data structures or instructions 304 (e.g., software) embodying or utilized by any one or more of the techniques or functions described herein. The instructions 304 may also reside, completely or at least partially, within the main memory 302, within static memory 306, or within the hardware processor 301 during execution thereof by the machine 21, 34. In an example, one or any combination of the hardware processor 301, the main memory 302, the static memory 306, or the mass storage device 316 comprises a machine -readable medium.
[0090] Specific examples of machine -readable media include, one or more of nonvolatile memory, such as semiconductor memory devices (e.g., EPROM or EEPROM) and flash memory devices; magnetic disks, such as internal hard disks and removable disks; magneto-optical disks; RAM; or optical media such as CD-ROM and DVD-ROMdisks. While the machine-readable medium is illustrated as a single medium, the term "machine readable medium" may include a single medium or multiple media (e.g., a centralized or distributed database, or associated caches and servers) configured to store the one or more instructions 304.
[0091] The term “machine readable medium” includes, for example, any medium that is capable of storing, encoding, or carrying instructions for execution by the machine 21, 34 and that cause the machine 21, 34 to perform any one or more of the techniques of the present disclosure or causes another apparatus or system to perform any one or more of the techniques, or that is capable of storing, encoding or carrying data structures used by or associated with such instructions. Non-limiting machine -readable medium examples include solid-state memories, optical media, or magnetic media. Specific examples of machine-readable media include: non-volatile memory, such as semiconductor memory devices (e.g., Electrically Programmable Read-Only Memory (EPROM), Electrically Erasable Programmable Read-Only Memory (EEPROM)) and flash memory devices; magnetic disks, such as internal hard disks and removable disks; magneto -optical disks; Random Access Memory (RAM); or optical media such as CD-ROM and DVD-ROM disks. In some examples, machine readable media includes non-transitory machine- readable media. In some examples, machine readable media includes machine readable media that is not a transitory propagating signal.
[0092] The instructions 304 may be transmitted or received, for example, over a communications network 305 using a transmission medium via the network interface device 320 utilizing any one of a number of transfer protocols (e.g., frame relay, internet protocol (IP), transmission control protocol (TCP), user datagram protocol (UDP), hypertext transfer protocol (HTTP), etc.). Example communication networks include a local area network (LAN), a wide area network (WAN), a packet data network (e.g., the Internet), mobile telephone networks (e.g., cellular networks), Plain Old Telephone (POTS) networks, and wireless data networks (e.g., Institute of Electrical and Electronics Engineers (IEEE) 802.11 family of standards known as Wi-Fi®), IEEE 802.15.4 family of standards, a Long Term Evolution (LTE) 4G or 5G family of standards, a UniversalMobile Telecommunications System (UMTS) family of standards, peer-to-peer (P2P) networks, satellite communication networks, among others.
[0093] In an example, the network interface device 320 includes one or more physical jacks (e.g., Ethernet, coaxial, or other interconnection) or one or more antennas to access the communications network 305. In an example, the network interface device 320 includes one or more antennas to wirelessly communicate using at least one of singleinput multiple-output (SIMO), multiple-input multiple-output (MIMO), or multiple -input single-output (MISO) techniques. In some examples, the network interface device 320 wirelessly communicates using Multiple User MIMO techniques. The term “transmission medium” shall be taken to include any intangible medium that is capable of storing, encoding or carrying instructions for execution by the machine 600, and includes digital or analog communications signals or other intangible medium to facilitate communication of such software.
[0094] Examples, as described herein, can include, or can operate on, logic or a number of components, modules, or mechanisms (all referred to hereinafter as “modules”). Modules are tangible entities (e.g., hardware) capable of performing specified operations and is configured or arranged in a certain manner. In an example, circuits are arranged (e.g., internally or with respect to external entities such as other circuits) in a specified manner as a module. In an example, the whole or part of one or more computer systems (e.g., a standalone, client or server computer system) or one or more hardware processors are configured by firmware or software (e.g., instructions, an application portion, or an application) as a module that operates to perform specified operations. In an example, the software can reside on a non-transitory computer readable storage medium or other machine -readable medium. In an example, the software, when executed by the underlying hardware of the module, causes the hardware to perform the specified operations.
[0095] Accordingly, the term “module” is understood to encompass a tangible entity, be that an entity that is physically constructed, specifically configured (e.g., hardwired), or temporarily (e.g., transitorily) configured (e.g., programmed) to operate in a specifiedmanner or to perform part or all of any operation described herein. Considering examples in which modules are temporarily configured, each of the modules need not be instantiated at any one moment in time. For example, where the modules comprise a general-purpose hardware processor configured using software, the general-purpose hardware processor is configured as respective different modules at different times. Software can accordingly configure a hardware processor, for example, to constitute a particular module at one instance of time and to constitute a different module at a different instance of time.
[0096] Implementations of the subject matter and the functional operations described in this specification can be implemented in digital electronic circuitry, in tangibly embodied computer software or firmware, in computer hardware, including the structures disclosed in this specification and their structural equivalents, or in combinations of one or more of them. Implementations of the subject matter described in this specification can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible, non-transitory computer-storage medium for execution by, or to control the operation of, data processing apparatus. Alternatively or in addition, the program instructions can be encoded on an artificially generated propagated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal that is generated to encode information for transmission to suitable receiver apparatus for execution by a data processing apparatus. The computer-storage medium can be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or a combination of one or more of them.
[0097] A computer program, which may also be referred to or described as a program, software, a software application, a module, a software module, a script, or code can be written in any form of programming language, including compiled or interpreted languages, or declarative or procedural languages, and it can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program may, but need not, correspond to a file in a file system. A program can be stored in a portion of a file thatholds other programs or data, e.g., one or more scripts stored in a markup language document, in a single file dedicated to the program in question, or in multiple coordinated fries, e.g., files that store one or more modules, sub-programs, or portions of code. A computer program can be deployed to be executed on one computer or on multiple computers that are located at one site or distributed across multiple sites and interconnected by a communication network. While portions of the programs illustrated in the various figures are shown as individual modules that implement the various features and functionality through various objects, methods, or other processes, the programs may instead include a number of sub-modules, third-party services, components, libraries, and such, as appropriate. Conversely, the features and functionality of various components can be combined into single components, as appropriate.
[0098] The processes and logic flows described in this specification can be performed by one or more programmable computers executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows can also be performed by, and apparatus can also be implemented as, special purpose logic circuitry, e.g., a CPU, a GPU, an FPGA, or an ASIC.
[0099] The term “graphical user interface,” or “GUI,” may be used in the singular or the plural to describe one or more graphical user interfaces and each of the displays of a particular graphical user interface. Therefore, a GUI may represent any graphical user interface, including but not limited to, a web browser, a touch screen, or a command line interface (CUI) that processes information and efficiently presents the information results to the user. In general, a GUI may include a plurality of user interface (UI) elements, some or all associated with a web browser, such as interactive fields, pull-down lists, and buttons operable by the user. These and other UI elements may be related to or represent the functions of the web browser.
[0100] Implementations of the subject matter described in this specification can be implemented in a computing system that includes a back-end component, e.g., as a data server, or that includes a middleware component, e.g., an application server, or that includes a front-end component, e.g., a client computer having a graphical user interfaceor a Web browser through which a user can interact with an implementation of the subj ect matter described in this specification, or any combination of one or more such back-end, middleware, or front-end components. The components of the system 20 can be interconnected by any form or medium of wireline and / or wireless digital data communication, e.g., a communications network 305.
[0101] Various Notes
[0102] Each of the non-limiting aspects in this document can stand on its own or can be combined in various permutations or combinations with one or more of the other aspects or other subject matter described in this document.
[0103] The above detailed description includes references to the accompanying drawings, which form a part of the detailed description. The drawings show, by way of illustration, specific embodiments in which the invention can be practiced. These embodiments are also referred to generally as “examples.” Such examples can include elements in addition to those shown or described. However, the present inventors also contemplate examples in which only those elements shown or described are provided. Moreover, the present inventors also contemplate examples using any combination or permutation of those elements shown or described (or one or more aspects thereof), either with respect to a particular example (or one or more aspects thereof), or with respect to other examples (or one or more aspects thereof) shown or described herein.
[0104] In this document, the terms “a” or “an” are used, as is common in patent documents, to include one or more than one, independent of any other instances or usages of “at least one” or “one or more.” In this document, the term “or” is used to refer to a nonexclusive or, such that “A or B” includes “A but not B,” “B but not A,” and “A and B,” unless otherwise indicated. In this document, the terms “including” and “in which” are used as the plain-English equivalents of the respective terms “comprising” and “wherein.” Also, in the following claims, the terms “including” and “comprising” are open-ended, that is, a system, device, article, composition, formulation, or process that includes elements in addition to those listed after such a term in a claim are still deemed to fall within the scope of that claim. Moreover, in the following claims, the terms “first,”“second,” and “third,” etc., are used merely as labels, and are not intended to impose numerical requirements on their objects.
[0105] Method examples described herein can be machine or computer-implemented at least in part. Some examples can include a computer-readable medium or machine- readable medium encoded with instructions operable to configure an electronic device to perform methods as described in the above examples. An implementation of such methods can include code, such as microcode, assembly language code, a higher-level language code, or the like. Such code can include computer readable instructions for performing various methods. The code may form portions of computer program products. Such instructions can be read and executed by one or more processors to enable performance of operations comprising a method, for example. The instructions are in any suitable form, such as but not limited to source code, compiled code, interpreted code, executable code, static code, dynamic code, and the like.
[0106] Further, in an example, the code can be tangibly stored on one or more volatile, non-transitory, or non-volatile tangible computer-readable media, such as during execution or at other times. Examples of these tangible computer-readable media can include, but are not limited to, hard disks, removable magnetic disks, removable optical disks (e.g., compact disks and digital video disks), magnetic cassettes, memory cards or sticks, random access memories (RAMs), read only memories (ROMs), and the like.
[0107] The above description is intended to be illustrative, and not restrictive. For example, the above-described examples (or one or more aspects thereof) may be used in combination with each other. Other embodiments can be used, such as by one of ordinary skill in the art upon reviewing the above description. The Abstract is provided to allow the reader to quickly ascertain the nature of the technical disclosure. It is submitted with the understanding that it will not be used to interpret or limit the scope or meaning of the claims. Also, in the above Detailed Description, various features may be grouped together to streamline the disclosure. This should not be interpreted as intending that an unclaimed disclosed feature is essential to any claim. Rather, inventive subject matter may lie in less than all features of a particular disclosed embodiment. Thus, the followingclaims are hereby incorporated into the Detailed Description as examples or embodiments, with each claim standing on its own as a separate embodiment, and it is contemplated that such embodiments can be combined with each other in various combinations or permutations. The scope of the invention should be determined with reference to the appended claims, along with the full scope of equivalents to which such claims are entitled.
[0108] References1. Ullah A, Munir S, Badshah SL, Khan N, Ghani L, Poulson BG, Emwas AH, Jaremko M. Important flavonoids and their role as a therapeutic agent. Molecules, 25(22):5243, 2020. doi: 10,3390 / molecules25225243.2. Liu Y, Qian J, Li J, Xing M, Grierson D, Sun C, Xu C, Li X, Chen K.Hydroxylation decoration patterns of flavonoids in horticultural crops: chemistry, bioactivity, and biosynthesis. Horticulture Research, 9:uhab068, 2022. DOI: 10. 1093 / hr / uhab068. https: / / doi.org / 10.1093 / hr / uhab0683. Sun M, Zhao S, Gilvary C, Elemento O, Zhou J, Wang L. Graph convolutional networks for computational drug development and discovery. Brief Bioinform. 21(3):919-935, 2020. doi: 10.1093 / bib / bbz042.4. Shi C, Xu M, Zhu Z, Zhang W, Zhang M, Tang J. Graphaf: a flow-based autoregressive model for molecular graph generation. arXiv preprint arXiv:2001.09382, 2020.5. Nigam A, Pollice R, Krenn M, Gomes GDP, Aspuru-Guzik A. Beyond generative models: superfast traversal, optimization, novelty, exploration and discovery (STONED) algorithm for molecules using SELLIES. Chemical Science, 12(20): 7079-7090, 2021. https: / / doi.org / 10, 1039%2Ldlsc00231g6. Rajan K, Steinbeck C, Zielesny A. Performance of chemical structure string representations for chemical image recognition using transformers. ChemRxiv,(DOI: 10.33774 / chemrxiv-2021-7c9wf-v2), 2021. htps: / / doi.org / 10.1039 / DlDD00Q13F Krenn M, Hase F, Nigam AK, Friederich P, Aspuru-Guzik A. Self-Referencing Embedded Strings (SELFIES): A 100% robust molecular string representation. Machine Learning: Science and Technology, 1:045024, 2020. htps: / / doi.org / 10.48550 / arXiv.1905.13741 Krenn M, Ai Q, Barthel S, Young A, Yu, R Aspuru-Guzik A. SELFIES and the future of molecular string representations. Patern Recognition, (DOI:10. 1016 / j .pater.2022. 100588). https: / / doi.Org / 10.1016 / i.pater.2022.100588 Wellawate GP, Seshadri A, White AD. Model agnostic generation of counterfactual explanations for molecules. Chemical Science, 13(13):3697-3705, 2022. DOI: 10.1039 / dlsc05259d. https: / / doi.org / 10.1039 / DlSC05259D Kwon S, Bae H, Jo J, Yoon S. Comprehensive ensemble in QS AR prediction for drug discovery. BMC Bioinformatics, 20(521): 1-19,2019.htps: / / doi.org / 10.1186 / sl2859-019-3135-4 Wang T, Wu MB, Lin JP, Yang LR. Quantitative structure-activity relationship: promising advances in drug discovery platforms. Expert Opinion on Drug Discovery, 10(9): 1283-1300, 2015. https: / / doi.org / 10.1517 / 17460441.2015.1083006 Li C, Qin X, Zhang Z. Structure-activity collective properties underlying selfassembled superstructures. Nano Today, 42(353): 101354, 2022. DOI:10. 1016 / j .nantod.2021.101354. htps: / / doi.Org / 10.1016 / i.nantod.2021.101354 Djeradi H, Rahmouni A, Cheriti A. Antioxidant activity of flavonoids: a QSAR modeling using Fukui indices descriptors. Journal of Molecular Modeling, 20(2476): 1-10, 2014. htps: / / doi.org / 10.1007 / sQ0894-014-2476-lChan K, Leung HCM, Tsoi JKH. Predictive QSAR model confirms flavonoids in Chinese medicine can activate voltage-gated calcium (CaV) channel in osteogenesis. Chinese Medicine, 15(31): 1-14, 2020. htps: / / doi.org / 10.1186%2Fsl3020-020-0Q313-l Alam S, Khan F. 3D-QSAR, Docking, ADME / Tox studies on Flavone analogs reveal anticancer activity through Tankyrase inhibition. Scientific Reports, 9(5414): 1-15, 2019. https: / / doi.org / 10.1038 / s41598-019-41984-7 Zuvela P, David J, Yang X, Huang D, Wong MW. Non-linear quantitative structure-activity relationships modelling, mechanistic study and in-silico design of flavonoids as potent antioxidants. International Journal of Molecular Sciences, 20(9):2328, 2019. https: / / doi.org / 10.3390%2Fiims20092328 Polishchuk P. CReM: chemically reasonable mutations framework for structure generation. Journal of Cheminformatics, 12(1):28, 2020. htps: / / doi.Org / 10.21203 / rs.2.23402 / yl Koirala N, Thuan NH, Ghimire GP, Thang DV, Sohng JK. Methylation of flavonoids: Chemical structures, bioactivities, progress and perspectives for biotechnological production. htps: / / doi.Org / 10.1016 / i.enzmictec.2016.02.003 Kim BG, Sung SH, Chong YC, et al. Plant flavonoid o -methyltransferases: substrate specificity and application. Journal of Plant Biology, 53:321-329,2010.https: / / doi.org / 10.1007 / sl2374-01Q-9126-7 Kumar S, Pandey AK. Chemistry and Biological Activities of Flavonoids: An Overview, The Scientific World Journal, vol. 2013, Article ID 162750, 16 pages, 2013. https : / / doi .org / 10, 1155 %2F2013 %2F 162750 Bailly C. The subgroup of 2'-hydroxy-flavonoids: Molecular diversity, mechanism of action, and anticancer properties. Bioorganic & Medicinal Chemistry, 32: 116001, 2021. https: / / doi.Org / 10.1016 / i.bmc.2021.116001Kim S, Chen J, Cheng T, Gindulyte A, He J, He S, Li Q, Shoemaker BA, et al. PubChem 2023 update, Nucleic Acids Research, Volume 51, Issue DI, 6 January 2023, Pages D1373-D1380. https : / / doi .org / 10,1093 / nar / gkac956 Sharma N, Ojha H, Raghav PK, Goyal RK. Chemoinformatics and Bioinformatics in the Pharmaceutical Sciences. Academic Press, 2021. ISBN: 978-0-12-821748-1. htps: / / doi.org / 10.1016 / C2019-0-04122-2 Feinberg EN, Sur D, Wu Z, Husic BE, Mai H, et al. PotentialNet for molecular property prediction. ACS Central Science, 4(11): 1520-1530, 2018. https: / / doi.org / 10.1021 / acscentsci.8b005Q7 Ponzoni I, Sebastian-Perez V, Requena-Triguero C, et al. Hybridizing feature selection and feature learning approaches in QSAR modeling for drug discovery. Sci Rep 7, 2403, 2017. htps: / / doi.org / 10.1038 / s41598-017-02114-3 Sheridan RP, Wang WM, Liaw A, Ma J, Gifford EM. Extreme gradient boosting as a method for quantitative structure-activity relationships. Journal of Chemical Information and Modeling, 56(12):2353-2360,2016.htps: / / doi.org / 10.1021 / acs.icim.6b00591 Tian H, Ketkar R, Tao P. Accurate ADMET Prediction with XGBoost. arXiv:2204.07532 q-bio.BM, 2022. https: / / doi.org / 10.48550 / arXiv.2204.07532 King -Diaz B, Granados-Pineda J, Bah M, Rivero-Cruz J F, Lotina-Hennsen B. Mexican propolis flavonoids affect photosynthesis and seedling growth. J Photochem Photobiol B: Biology 151, 213-220, 2015. htps: / / doi.Org / 10.1016 / i.iphotobiol.2015.08.019 Park HM, Hong JE, Park ES, Yoon HS, Seo DW, et al. Anticancer effect of tectochrysin in colon cancer cell via suppression of NF-kappaB activity and enhancement of death receptor expression. Molecular Cancer, 14(124): 1-15, 2015. htps: / / doi.org / 10.1186%2Fsl2943-015-Q377-2Ntakoulas DD, Pasias IN, Raptopoulou KG, Proestos C. Tectochrysin: Advances on resources, biosynthesis pathway, bioavailability, bioactivity, and pharmacology. In: J. Xiao (eds) Handbook of Dietary Flavonoids. Springer, Cham, 2023. htps: / / doi.org / 10.1007 / 978-3-03Q-94753-8 81-1 Zhang Q, Zheng Y, Hu X, Hu X, Lv W, et al., Ethnopharmacological uses, phytochemistry, biological activities, and therapeutic applications of Alpinia oxyphylla Miquel: A review. J Ethnopharmacol, 224: 149-168, 2018. htps: / / doi.Org / 10.1016 / i.jep.2018.05.002 Wu IT, Kuo CY, Su CH, Lan YH, Hung CC. Pinostrobin and tectochrysin conquer multidrug-resistant cancer cells via inhibiting p -glycoprotein ATPase. Pharmaceuticals, 16(2): 205, 2023. https: / / doi.org / 10.3390%2Fphl60202Q5 Hou R, Han Y, Fei Q, Gao Y, Qi R, Cai R, Qi Y. Dietary flavone tectochrysin exerts anti-inflammatory action by directly inhibiting MEK1 / 2 in LPS-primed macrophages. Molecular Nutrition & Food Research, (2): 1700288, 2017. https: / / doi.org / 10.1002 / mnfr.20170Q288 Shimada T, Nagayoshi H, MurayamaN, Sawai A, Kim V, Kim D, Yamazaki H, Guengerich FP, Takenaka S. Oxidation of 3 '-methoxyflavone, 4'-methoxyflavone, and 3',4'-dimethoxyflavone and their derivatives having with 5,7-dihydroxyl moieties by human cytochromes P450 1B1 and 2A13. Xenobiotica, 52(2), 134-145, 2022. htps: / / doi.org / 10.1080 / 00498254.2022.2Q62486
Claims
CLAIMS:
1. An apparatus for generating at least one novel molecule having predefined characteristics, comprising: processing circuitry; a memory device coupled to the processing circuitry, the memory device comprising instructions executable by the processing circuitry, wherein the processing circuitry is configured to: receive at least one input molecule, and randomize the at least one input molecule to generate a plurality of randomized molecules; generate a mutation of each of the plurality of randomized molecules iteratively until at least one candidate molecule having a predetermined substructure is obtained; standardize different representations or descriptions of the same molecule to a single, unique form to form a canonical Simplified Molecular Input Line Entry System (SMILES) string; store the at least one candidate molecule in the memory device to form a first training dataset for a neural network; compute feature vector representation of the at least one candidate molecule, wherein the feature vector representation comprises at least one of a descriptor, molecular structure and substructure; with a pretrained model, predict whether the at least one candidate molecule possesses at least one desired property; determine whether the at least one candidate molecule can be synthesized to produce a physical molecule.
2. The apparatus of claim 1, wherein each of the plurality of randomized molecules is represented in Self-Referencing Embedded Strings (SELFIES).
3. The apparatus of claim 1, wherein SMILES strings of each of the at least one candidate molecule is updated in a set of candidate molecules and a dictionary of similarity scores to the input molecule.
4. The apparatus of claim 1, wherein the processing circuitry executes the instructions to generate a chemical space of polyphenol derivatives with predetermined Absorption, Distribution, Metabolism, Excretion, and Toxicity (ADMET) properties.
5. The apparatus of claim 3, wherein the processing circuitry executes the instructions to train Quantitative Structure-Activity Relationship (QSAR) modeling to fdter the list of SMILES strings of modified unique molecules.
6. The apparatus of claim 5, wherein the processing circuitry executes the instructions to generate variant molecules from a given molecule scaffold.
7. The apparatus of claim 6, wherein the scaffold is a phytochemical.
8. The apparatus of claim 6, wherein the scaffold is a synthetic molecule.
9. The apparatus of claim 7, wherein the phytochemical is a polyphenol.
10. The apparatus of claim 9, wherein the polyphenol is at least one of a flavonoid and a stilbene.
11. The apparatus of claim 7, wherein the phytochemical at least one of a terpenoid and an alkaloid.
12. The apparatus of claim 1, wherein the plurality of randomized molecules are obtained using a computational algorithm.
13. The apparatus of claim 1, wherein the computational algorithm is a modified Superfast Traversal, Optimization, Novelty, Exploration, and Discovery (STONED) algorithm for generating a plurality of SMILES strings.
14. The apparatus of claim 13, wherein the modified STONED algorithm facilitates targeted modifications in molecular structures, and wherein the modifications comprise at least one of an addition or a substitution of chemical groups in a chemical space of a scaffold.
15. The apparatus of claim 13, wherein the chemical groups comprise: hydroxyl (- OH), methyl (-CH3), ethyl (-C2H5), propyl (-C3H7), butyl (-C4H9), or pentyl (- C5H11); prenyl groups (-C5H9, -C10H14); carbonyl (>C=O); carboxyl (-COOH), acetyl (CH3C0-), amino (-NH2), sulfhydryl (-SH), nitro (-NO2), or halogen (-C1, -Br, - I, -F) groups.
16. The apparatus of claim 4, wherein the processing circuitry executes the instructions to predict a molecular scaffold using scaffold splitting and stratified splitting to balance the training, test, and validation datasets.
17. The apparatus of claim 4, wherein the processing circuitry executes the instructions to predict interactions of the synthesized molecule with disease biomarker protein targets.
18. The apparatus of claim 4, wherein the processing circuitry executes the instructions to extract a particular set of features from extract a particular set of features from molecular objects and group the extracted features to generate the training data.
19. The apparatus of claim 4, wherein the processing circuitry executes the instructions to determine a performance of the synthesized molecule in an assay specific to a given protein target.
20. The apparatus of claim 4, wherein the processing circuitry executes the instructions to determine to rank and filter molecules based on a predefined criteria.
21. A method for generating at least one novel molecule, the method comprising the steps of: at processing circuitry executing instructions stored in a memory device, receiving at least one input molecule, and randomizing the at least one input molecule to generate a plurality of randomized molecules; generating a mutation of each of the plurality of randomized molecules iteratively until at least one candidate molecule having a predetermined substructure is obtained; standardizing different representations or descriptions of the same molecule to a single, unique form to form a canonical SMILES string; storing the at least one candidate molecule in the memory device to form a first training dataset for a neural network; computing feature vector representation of the at least one candidate molecule, wherein the feature vector representation comprises at least one of a descriptor, molecular structure and substructure; predicting whether the at least one candidate molecule possesses at least one desired property; and determining whether the at least one candidate molecule can be synthesized to produce a physical molecule.
22. The method of claim 21, wherein each of the plurality of randomized molecules is represented in SELFIES.
23. The method of claim 22, wherein SMILES strings of each of the at least one unique mutants is updated in a set of unique mutants and a dictionary of similarity scores to the input molecule.
24. The method of claim 21, comprising a further step of generating a chemical space of polyphenol derivatives with predetermined ADMET properties.
25. The method of claim 21, comprising a further step of predicting interactions of the synthesized molecule with disease biomarker protein targets.
26. The method of claim 21, comprising a further step of predicting unique molecules’ ADMET properties as an indication of drugability.
27. The method of claim 21, comprising a further step of executing instructions to rank and select a set of the at least one candidate molecules with promise as stable and feasible molecular entities.
28. The method of claim 21, comprising a further step of executing instructions for drug-target interaction modeling and for facilitating accurate and computationally efficient virtual screening.
29. The method of any one of claims 21 to 28, wherein a binding site of a candidate molecule is an active site for a target protein.
30. The method of any one of claims 21 to 29, further comprising synthesis of the physical molecule of at the least one of the candidate molecules generated and validating the predicted properties of the physical molecule.
31. The method of claim 30, further comprising performing in vitro or in vivo experiments with the physical candidate molecule.
32. A non-transitory computer readable medium having instructions stored thereon, the instructions being executable by the processing circuitry to configure the processing circuitry to at least: receive at least one input molecule, and randomize the at least one input molecule to generate a plurality of randomized molecules; generate a mutation of each of the plurality of randomized molecules iteratively until at least one candidate molecule having a predetermined substructure is obtained; standardize different representations or descriptions of the same molecule to a single, unique form to form a canonical SMILES string; store the at least one candidate molecule in the memory device to form a first training dataset for a neural network; compute feature vector representation of the at least one candidate molecule, wherein the feature vector representation comprises at least one of a descriptor, molecular structure and substructure; with a pretrained model, predict whether the at least one candidate molecule possesses at least one desired property; and determine whether the at least one candidate molecule can be synthesized to produce a physical molecule.
Citation Information
Patent Citations
Interpretable molecular generative models
US20220189578A1
Prediction device, trained model generation device, prediction method, and trained model generation method
US20220284987A1
Scaffold constrained molecular generation using memory networks
WO2022043362A1
Identification of molecular inhibitors using reinforcement learning and realtime docking of 3D structures
WO2023077124A1