Methods and systems for drug discovery and evaluation
AI integration into drug development addresses inefficiencies by improving target identification, synthesis planning, and clinical trial design, enhancing the efficiency and reducing costs in the drug development process.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- MACAU WO KING WAI GROUP LTD
- Filing Date
- 2026-01-20
- Publication Date
- 2026-07-23
AI Technical Summary
Drug development is a complex, time-consuming, and costly process with low success rates due to the complexity of diseases, vast chemical space to explore, and stringent regulatory requirements, which traditional methods struggle to address efficiently.
Integration of artificial intelligence (AI) technologies, particularly large language models and generative AI, into the drug development pipeline for tasks such as target identification, drug discovery, de novo design, synthesis planning, and clinical studies, utilizing multi-omics data, machine learning algorithms, and real-world data to enhance efficiency and accuracy.
AI accelerates drug development by improving target identification, predicting drug mechanisms of action, optimizing synthesis processes, and enhancing clinical trial design, thereby reducing time and costs while increasing the success rate of drug discovery.
Smart Images

Figure US2026011860_23072026_PF_FP_ABST
Abstract
Description
Attorney Docket No. 72194-701.602METHODS AND SYSTEMS FOR DRUG DISCOVERY AND EVALUATION CROSS-REFERENCE TO RELATED APPLICATION
[0001] This application claims the benefit and priority of International Patent Application PCT / CN2025 / 073210 entitled “AI-BASED DRUG MOA PREDICATION, DRUG EVALUATION METHOD AND SYSTEM,” filed January 20, 2025, the contents of which are incorporated herein by reference in the entirety for all purposes.FIELD
[0002] The present disclosure relates to the fields of biomedicine and artificial intelligence applications. The present disclosure provides a method for predicting the mechanism of action of drugs based on artificial intelligence, and based on this method, it has been discovered that ramelteon promotes estradiol. The present disclosure also provides a method for constructing an artificial intelligence drug evaluation system, as well as the corresponding evaluation method, computer system, and computer readable medium.BACKGROUND
[0003] Drug development is a complex and time-consuming endeavor that traditionally relies on the experience of drug developers and trial -and-error experimentation. The advent of artificial intelligence (Al) technologies, particularly emerging large language models and generative Al, is poised to redefine this paradigm. The integration of Al-driven methodologies into the drug development pipeline has already heralded subtle yet meaningful enhancements in both the efficiency and effectiveness of this process.
[0004] Drug development is a multifaceted process aimed at developing new medications to treat diseases. It involves several stages, including biomarker and target identification, drug discovery, preclinical studies, clinical trials, regulatory approval, and postmarket surveillance. Drug development currently faces numerous challenges, including high costs, long timeframes, and low success rates. On average, developing a new drug requires a substantial investment of approximately US$2.6 billion and may take 12 to 15 years to complete. Unfortunately, the success rate of new drugs, even at the clinical trial stage, is less than 10%. Several factors are responsible. Fundamentally, diseases are often complex and multifactorial, posing difficulties in identifying effective treatments. The drug development process itself is also complex, involving multiple stages where setbacks can doom the entire process. Additionally, the vast chemical space that needs to be explored in search of a potential drug candidate is estimated to be on the order of 10, making drug discovery comparable toAttorney Docket No. 72194-701.602finding a needle in a haystack. Lastly, regulatory requirements are stringent, and meeting the standards for safety, efficacy, and quality can be a time-consuming and costly endeavor. To overcome these challenges, scientists have been actively exploring new technologies and methods to improve the drug development process — with artificial intelligence (Al) being poised to radically alter this field.SUMMARY
[0005] Drug development is a complex and time-consuming endeavor that traditionally relies on the experience of drug developers and trial -and-error experimentation. The advent of artificial intelligence (Al) technologies, particularly emerging large language models and generative Al, is poised to redefine this paradigm. The integration of Al-driven methodologies into the drug development pipeline has already heralded subtle yet meaningful enhancements in both the efficiency and effectiveness of this process. In some aspects, provided herein are embodiments of Al applications across the drug development workflow, encompassing the identification, of disease targets, drug discovery and de novo design, synthesis planning, preclinical and clinical studies, and post-market surveillance.
[0006] In some embodiments, provided herein is a method for predicting drug mechanism of action based on artificial intelligence (Al), comprising the following steps: (1) a step of treatment of drugs with known MOAs and data acquisition: preparing a variety of drugs with known mechanisms of action (MOAs) or signaling pathways, and a genome-wide gene expression perturbation tool; applying the drugs with known MOAs to cells cultured in multi -well plates, and using the genome-wide gene expression perturbation tool to treat different wells to obtain initial data of the cells under drug action and information on gene expression changes; (2) a step of multi-omics detection and analysis comprising: conducting multi -omics detection on the cells treated in step 1; performing marker staining on the cells to detect the expression and localization of specific proteins or molecules within the cells through specific markers; performing morphological imaging on the cells to obtain morphological data; and integrating the multi-omics detection data, marker staining data, and cell morphological imaging data, and combining the existing drug MOA knowledge and gene signaling pathway knowledge for analysis; (3) a step of Al model training: inputting the data analyzed in step 2 into an Al system, and using the data to train the artificial intelligence model, enabling the artificial intelligence model to learn the relationships and regularities between drug characteristics and known MOAs, said drug characteristics including multi-omics characteristics and cell morphological characteristics; (4) a step of prediction of unknown drug MOAs: after the training of the AlAttorney Docket No. 72194-701.602model is completed, applying drugs with unknown MOAs to the cells, and obtaining the multi -omics data and cell morphological data of the cells after treatment with the drugs with unknown MOAs; inputting the obtained data after treatment with drugs with unknown MOAs into the trained Al model, and through the Al model, comparing and analyzing the input data characteristics with the previously learned patterns, thereby predicting the possible MOAs of the unknown drugs.
[0007] In some embodiments, provided herein is a method for constructing an artificial intelligence-based drug evaluation system, comprising the following steps: acquiring data from sources, for example, electronic health records (EHRs), insurance claims, and / or wearable devices; performing multimodal data embedding processing on the acquired multiple types of data to generate a unified embedded data representation; constructing a foundation model based on the embedded data representation; performing generative artificial intelligence processing based on the foundation model to obtain an artificial intelligence large language model for evaluating the effectiveness and safety of drugs; comparing the output of the artificial intelligence large language model with real clinical data to perform prediction quality control, fine-tuning the system based on the comparison results to further predict therapeutic response, survival intervals, and adverse events, and finally obtaining a trained artificial intelligence large language model.
[0008] In some embodiments, provided herein is an artificial intelligence-based drug evaluation method, comprising: inputting clinical trial data into an artificial intelligence large language model trained according to the method for constructing an artificial intelligence-based drug evaluation system; performing a prediction process for drug evaluation in the artificial intelligence large language model to predict therapeutic response, survival intervals, and adverse events according to the clinical trial data.
[0009] Another aspect of the present disclosure provides a non-transitory computer readable medium comprising machine executable code that, upon execution by one or more computer processors, implements any of the methods above or elsewhere herein.
[0010] Another aspect of the present disclosure provides a system comprising one or more computer processors and computer memory coupled thereto. The computer memory comprises machine executable code that, upon execution by the one or more computer processors, implements any of the methods above or elsewhere herein.
[0011] Additional aspects and advantages of the present disclosure will become readily apparent from the following detailed description, wherein only illustrative embodiments of the present disclosure are shown and described. As will be realized, the present disclosure isAttorney Docket No. 72194-701.602capable of other and different embodiments, and its several details are capable of modifications in various obvious respects, all without departing from the present disclosure. Accordingly, the drawings and description are to be regarded as illustrative in nature and not as restrictive.INCORPORATION BY REFERENCE
[0012] All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. To the extent publications and patents or patent applications incorporated by reference contradict the disclosure contained in the specification, the specification is intended to supersede and / or take precedence over any such contradictory material.BRIEF DESCRIPTION OF THE DRAWINGS
[0013] The novel features of the disclosure are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present disclosure will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the disclosure are utilized, and the accompanying drawings of which:
[0014] FIG. 1 shows an overview of Al applications in the drug development pipeline. The drug development pipeline encompasses several critical stages, including biomarker and target identification, drug discovery, preclinical studies, clinical trials, regulatory agency review and post-market surveillance. Al technologies have potential applications across nearly all these stages. CMC, chemical manufacturing and control; DMPK, drug metabolism and pharmacokinetics.
[0015] FIG. 2 shows a pipeline for Al-driven molecular generation in drug discovery. Molecular representations — one-dimensional (1D), two-dimensional (2D) and 3D structures — are derived from diverse compound, target and drug-target interaction databases and are used to train Al models such as generative adversarial networks (a type of neural network architecture consisting of two competing networks, a generator and a discriminator, working together to create realistic data samples), recurrent neural networks (used for processing sequential data), variational autoencoders (generative models that learn to encode input data into a latent space and then decode it back to reconstruct the original data), normalizing flows (a class of generative models that transform a simple probability distribution into a more complex one through a series of invertible transformations) and diffusion models (generative models thatAttorney Docket No. 72194-701.602create data by simulating a diffusion process). These models generate novel molecules, which are subsequently assessed for chemical validity, synthetic accessibility and drug-like properties, which ultimately enables the identification of new drug-like compounds.
[0016] FIG. 3 shows Al-driven synthesis planning and automation in drug discovery, a, Synthesis planning. The synthesis planning process commences with retrosynthetic analysis that breaks down target molecules into commercially available or known building blocks, followed by reaction prediction to anticipate the required chemical reactions and conditions for synthesizing the target molecule, b, Automated synthesis. The schematic illustrates the seamless integration of Al-driven software with experimental execution and result analysis within the automated synthesis process.
[0017] FIG. 4 shows Al-driven MOA prediction using high-content screening and multi -omics data, a, In high-content screening, cells are cultured in multi-well plates, treated with a diverse range of drugs with known MOAs or signaling pathways (upper) and genomewide gene expression perturbations (lower) are applied to different wells, b, Within each well, data on multi-omics features, marker staining patterns and cell morphology characteristics over time are combined with knowledge of the corresponding MOAs or gene signaling pathway changes, and used to train an Al model to understand the effect of each drug on cellular networks, c, As a result, this Al model is able to predict the MOA of a new compound with similar multi-omics and cell morphology features.
[0018] FIG. 5 shows an example of utilizing Al capabilities to enhance both clinical trial processes and real-world medical practice, a, Training process: training involves using diverse clinical and trial data (EHRs, wearables, genomics, imaging) to develop an AI-LLM via multimodal embedding and generative Al. It assesses drug efficacy, optimizes trial protocols and enables smart preclinical and clinical studies, b, Validation and prediction process: The AI-LLM is validated using real-world and clinical trial data, fine-tuned with therapeutic outcomes and adverse events. It predicts drug efficacy, assesses protocol feasibility and optimizes trials, enabling smart preclinical and clinical studies to accelerate drug development.
[0019] FIG. 6 shows ramelteon treatment alleviates female reproductive aging. (A) The workflow of the phenotypic screening on human granulosa cell KGN cells. N=4; (B) The treatment of ramelteon rescued E2 secretion in KGN cells. N=4; (C) The treatment of ramelteon upregulated the serum level of E2 in the 18-month-old female mice. N=9-11; (D) The treatment of ramelteon improved the fertility in the 18-month-old female mice. N=9-11. *, p<0.05; **, p<0.01; ***, p<0.001; ****, p<0.0001. Error bars stand for SEM of biological repeats.Attorney Docket No. 72194-701.602
[0020] FIG. 7 shows ramelteon rescued the ovarian dysfunction in the mouse model of POI. (A-C) The serum level of E2 (A), pregnant rate (B) and changes of ovarian follicles (C) in the young female mice upon dietary BCAA restriction induced POI, with or without ramelteon treatment. N=5-7. *, p<0.05; ***, p<0.001; ****, p<0.0001. SI, Primordial; S2, Primary; S3, Secondary; S4, Antral; S5, Atretic. Error bars stand for SEM of biological repeats.
[0021] FIG. 8A shows high-content analysis of various compounds on Huh7 with OA, In Huh7 cells treated with OA, Lipid droplet (LD) number and volume were assessed for ZK23.5 (10μM and 20μM), DMSO was used as a solvent control, DGAT1 / 2 inhibitor and BFA were negative controls, Aldometanib was positive control.
[0022] FIG. 8B shows high-content analysis of various compounds on Huh7 without OA. In Huh7 cells without OA treatment, Lipid droplet (LD) number and volume were assessed for ZK23.5 (10μM and 20μM), DMSO was used as a solvent control, DGAT1 / 2 inhibitor and BFA were negative controls, Aldometanib was positive control.
[0023] FIG. 8C shows decreased droplet area per cell in HepG2 cell treating with 10 pM ZK23.5 for 24 h or 36 h. BODIPY / DAPI staining for fixed HepG2 cells treating with 10μM ZK23.5 for 24 or 36 hours under normal, low glucose and MCDM conditions.
[0024] FIG. 8D shows lipid droplets in lysozyme increased significantly in HepG2 cell treating with ZK23.5. BODIPY / DAPI staining for live HepG2 cell which was treated with 10μM compound ZK23.5 for 4-6h and then treated with Lalistat 2 for 12 h.
[0025] FIG. 8E shows autophagic lipolysis assessment based on Lipodetector.HepG2 cells treated with 10 pM ZK23.5 for 24 hours in both normal and low -glucose media were analyzed by flow cytometry.
[0026] FIG. 9 shows F-actin fiber in cell treated with cytochalasin D and Y-27632. The fluorescence image of F-actin (green color) in Lifeact-GFP RPE1 cells were captured for treatment with DMSO (negative control), 0.3μM cytochalasin D, 1μM, 10μM Y-27632 and 0.3μM cytochalasin D and 1μM Y-27632 in combine at Oh, 8h, 16h and 24h.
[0027] FIG. 10 shows a computer system that is programmed or otherwise configured to implement methods provided herein.DETAILED DESCRIPTION
[0028] Recent advancements in Al, including image recognition, natural language processing (NLP), and computer vision, have shown particular promise for addressing key challenges in drug development. In particular, large language models (LLMs) like ChatGPT and Gemini, and generative Al such as Sora, have demonstrated capabilities that in some instances,Attorney Docket No. 72194-701.602already surpass human intelligence. Al’s ability to process vast amounts of data promises to greatly accelerate and improve the drug development process. Consequently, pharmaceutical companies, biotech firms, and research institutions are increasingly adopting Al-driven approaches to surmount the obstacles inherent in traditional methods. Al has proven valuable in analyzing complex biological systems, identifying disease biomarkers and potential drug targets, simulating drug-target interactions, predicting the safety and efficacy of drug candidates, and managing clinical trials (FIG. 1). However, it is important to recognize that Al-powered drug development still faces several unique challenges. There is an expectation for more successful applications of artificial intelligence in developing new drugs and uncovering new uses for existing drugs.
[0029] In some embodiments, provided herein are methods for Al-powered drug discovery, from target identification up to synthesis planning, and Al applications within clinical stages of drug development — including biomarker discovery, drug repurposing, prediction of pharmacokinetic properties and toxicity, and clinical trial conduct. In some embodiments, provided herein are methods for using Al on various aspects of drug discovery, including target identification, virtual screening, de novo design, ADMET (absorption, distribution, metabolism, excretion and toxicity) predictions, and synthesis planning and automating synthesis and drug discovery. In some embodiments, by leveraging advanced algorithms and techniques, provided herein are methods to accelerate the discovery of novel therapeutic agents, improve the accuracy of predictions and reduce the overall time and costs associated with drug development.Target identification
[0030] The identification of small-molecule targets, such as proteins or nucleic acids, is a critical process in drug discovery. Traditional methods such as affinity pull-down and whole-genome knockdown screening are widely used, but tend to be time consuming and labor intensive, with high failure rates.
[0031] Advances in Al technology are revolutionizing this field by enabling the analysis of large datasets within complex biological networks. In some embodiments, Al facilitates the identification of disease-related molecular patterns and causal relationships by constructing multi-omics data networks, thus facilitating the discovery of candidate drug targets. In some embodiments, NLP techniques (such as word2vec embeddings) are used to map gene functions into high-dimensional space, enhancing the sensitivity of target identification despite the sparsity of gene function overlap.Attorney Docket No. 72194-701.602
[0032] In some embodiments, the methods provided herein comprise integrating multi -omics data efficiently and ensuring the interpretability of Al models. In some embodiments, graph deep learning technology is used to address these by merging graph structures with deep learning, focusing on graph nodes related to key features (for example, atom type, charge) to effectively identify candidate targets. In some embodiments, the methods provided herein comprise an interpretable framework using multi-omics network graphs with graph attention mechanisms to predict cancer genes effectively. In some embodiments, the methods provided herein comprise integrating multi-omics data with scientific and medical literature into knowledge graphs allows Al to discern relationships between genes and disease pathways. In some embodiments, the methods provided herein comprise using biomedical LLMs which are deeply integrated with biological networks or knowledge graph functions, to provide efficient and precise methods for linking diseases, genes and biological processes. In some embodiments, the methods provided herein comprise using multi-omics data and biological network analysis to recognize TRAF2- and NCK-interacting kinase as a potential target for anti-fibrotic therapy to identify a specific TRAF2- and NCK-interacting kinase inhibitor. In some aspects, the potential publication biases in the literature suggest a need for supplementary methods to ensure the identification of novel and relevant targets.
[0033] Real-world data, such as medical records, self-reports, electronic health records (EHRs) and insurance claims, provide essential contextual information for understanding complex diseases and facilitating target discovery. However, real-world data often contain unstructured text, lack standardization and may include biases, limiting their application in this context. While high-quality, curated datasets are crucial for training models, real-world data are inherently noisy and complicated by the confluence of multiple diseases. Despite these issues, in some embodiments, the methods provided herein comprise using noisy real-world data to train effective models, advancing the potential for gene discovery and candidate drug targets in noisy medical record and non-expert disease labeling scenarios. In some embodiments, the methods provided herein comprise enhancing model generalizability across diverse populations, especially for diseases with low labeling or prevalence rates. As real-world and multi-omics data grow richer, in some embodiments, the methods provided herein comprise utilizing advanced data mining algorithms and expert knowledge to further enhance their integration, significantly improving the success rate of target discovery.Attorney Docket No. 72194-701.602Virtual screening
[0034] In some embodiments, the methods provided herein comprise virtual screening for efficiently identifying potential lead compounds or drug candidates. In some embodiments, the methods provided herein comprise Al technologies for ligand docking, in view of the rapid expansion of compound libraries leading to accelerated virtual screening of ultra-large libraries. In some embodiments, the methods provided herein comprise using AI-based receptor-ligand docking models to predict ligand spatial transformations, directly generate complex atomic coordinates using algorithms like equivariant neural networks and learn the probability density distribution of receptor-ligand distances to generate binding poses. In some embodiments, the methods provided herein comprise using receptor-ligand co-folding networks (e.g., based on AlphaFold2 or RosettaFold) in predicting complex structures directly from sequence information. In some embodiments, the methods provided herein comprise using postprocessing (for example, energy minimization) and / or geometric constraints to optimize docking pose validity, and to reduce or minimize unrealistic ligand conformations due to insufficient learning of physical constraints. In some embodiments, the methods provided herein comprise using deep learning-based binding pose prediction models in pocket-oriented docking tasks, which methods comprise considering receptor pocket flexibility. In some embodiments, the methods provided herein comprise using predicting precise receptor-ligand interaction, such as using deep learning models in affinity prediction. In some embodiments, the methods provided herein comprise handling both three-dimensional (3D) structural and nonstructural data, thereby outperforming traditional functions. In some embodiments, the methods provided herein comprise using deep learning models based on ligand pose accuracy and / or known receptor structures.
[0035] In some embodiments, the methods provided herein comprise docking-based virtual screening using known target structures. In some embodiments, the methods provided herein comprise docking-based virtual screening in cases where target structures are absent or incomplete. In some embodiments, the methods provided herein comprise using Al techniques in sequence-based prediction methods. In some embodiments, the methods provided herein comprise using Al techniques to capture the complexity of 3D protein-ligand interactions, providing accurate predictions of how binding pose changes affect interaction strength.
[0036] While targeted drug development is effective for defined targets, many diseases lack such targets. In some embodiments, the methods provided herein comprise using phenotype-based virtual screening for diseases with undefined targets (for example, rareAttorney Docket No. 72194-701.602diseases) and broadly phenotypic diseases (for example, aging). In some embodiments, the methods provided herein comprise using nuclear morphology and machine learning to identify compounds inducing senescence in cancer cells and / or for antibiotic discovery. In some embodiments, the methods provided herein comprise using case-specific phenotypic data. In some embodiments, the methods provided herein comprise using ligand chemical structures in Al-based activity prediction. To address challenges like data sparsity and imbalance and activity cliffs, in some embodiments, the methods provided herein comprise integrating related biological information like cell morphology and / or transcriptional profiles, to provide more accurate prediction of activity and / or MOA of drug candidates.
[0037] In some embodiments, the methods provided herein comprise using tasks such as scoring, pose optimization, and / or screening. In some embodiments, the methods provided herein comprise using universal models capable of handling multiple tasks. In some embodiments, the methods provided herein comprise incorporating inductive biases (which refer to the model’s inherent tendency to prioritize certain types of solutions over others) and / or data augmentation (which refers to techniques used to artificially expand the diversity of a training dataset without collecting new data) to improve model generalizability. In some embodiments, the methods provided herein comprise screening compound collections (e.g., to billions of molecules) and / or expanding molecular libraries beyond only a small portion of the druggable chemical space.
[0038] In some embodiments, the methods provided herein comprise active learning and / or Bayesian optimization, e.g., for addressing the chemical space search problem to enhance virtual screening efficiency. In some embodiments, the methods provided herein comprise the integration of quantum mechanics with Al for chemical space exploration, and in some aspects, molecular dynamics simulations can be used to add depth to protein-ligand interactions, addressing issues of binding affinity and selectivity to improve model accuracy. In some embodiments, the methods provided herein comprise using deep generative models to generate custom virtual libraries for specific targets or compound types, substantially narrow search spaces and enhance screening efficiency. In some embodiments, the methods provided herein comprise using a conditional recurrent neural network to generate a custom library to identify an inhibitor, e.g., an efficient and selective RIPK1 inhibitor in cell and animal models.De novo design
[0039] In some embodiments, the methods provided herein comprise de novo drug design involving autonomously creating new chemical structures to optimally satisfy desiredAttorney Docket No. 72194-701.602molecular features. Traditional methods, including structure-based, ligand-based and pharmacophore-based designs, are manual and rely on expert designers and explicit rules. In some embodiments, the methods provided herein comprise using Al, particularly deep learning, to enable the automated identification of novel structures that meet specific requirements, bypassing traditional expertise. In some embodiments, the methods provided herein comprise using this technology in developing small-molecule inhibitors, PROTACs, peptides, and / or functional proteins, which are then validated through wet-lab experiments.
[0040] FIG. 2 shows an example of a deep learning-driven de novo design, where the molecular generation component is central, normally using chemical language or graph-based models. In some aspects, chemical language models convert molecular generation tasks into sequence generation such as SMILES string (meaning “simplified molecular input line entry system,” a notation system that represents a chemical structure in a linear text format). In some cases, extensive pretraining is required and may produce invalid SMILES due to syntactic errors, and the methods provided herein comprise using these errors to aid model self-correction by filtering improbable samples. In some embodiments, the methods provided herein comprise using models like long short-term memory models (a type of deep learning model that analyzes sequential data). In some embodiments, the methods provided herein comprise using a model to address information compression bottlenecks. In some embodiments, the methods provided herein comprise using a model to learn global sequence attributes. In some embodiments, the methods provided herein comprise using a Transformer model to capture global properties. In some embodiments, the methods provided herein comprise integrating structured state-space sequences into chemical language models to reveal high chemical space similarity and alignment with key natural product design features, proving methods of using chemical language models in de novo design.
[0041] In some embodiments, the methods provided herein comprise using graph-based models to represent molecules as graphs, generating structures using autoregressive or non-autoregressive strategies. Autoregressive approaches construct molecules atom by atom, which can lead to chemically implausible intermediates and introduce bias. In contrast, non-autoregressive methods generate entire molecular graphs at once but need extra steps to ensure the graph’s validity, as these models’ limited perception of molecular topological structures can induce flawed structures.
[0042] Given the vastness of the drug-like chemical space, de novo generation often guides design toward target features, using optimization mechanisms such as scoring functions based on metrics including similarity to known active molecules and predicted bioactivity. InAttorney Docket No. 72194-701.602some embodiments, the methods provided herein comprise incorporating reinforcement learning for iterative optimization. In some embodiments, the methods provided herein comprise designing appropriate scoring functions and / or quantifying objectives like synthetic feasibility or drug likeness. In some embodiments, the methods provided herein comprise using active or curriculum learning strategies to address challenges in sample efficiency which can be associated with reinforcement learning.
[0043] Beyond introducing scoring functions, in some embodiments, the methods provided herein comprise incorporating constraints - such as disease-related gene expression features, pharmacophores, protein sequences or structures, binding affinity and protein-ligand interactions - in order to direct models toward generating desired molecules. In some embodiments, the methods provided herein comprise a PocketFlow model, conditioned on protein pockets, to effectively generate experimentally validated active compounds against targets. In some embodiments, the methods provided herein comprise models to refine leads by restricting outputs to specific scaffolds or fragments from desired candidates, albeit at the cost of limiting chemical diversity.ADMET
[0044] ADMET plays a critical role in determining drug efficacy and safety. While wet-lab evaluations are required for market approval and cannot be fully replaced by simulations, early-stage ADMET predictions can help reduce failures due to poor characteristics. Al has emerged as a valuable tool for predicting ADMET properties using predefined features like molecular fingerprints or descriptors. In some embodiments, methods provided herein comprise using machine learning techniques such as random forest and support vector machines, using descriptors like circular extended connectivity fingerprints to ensure accuracy and relevance. Over the past decades, various descriptors for ADMET predictions have been developed. However, feature engineering involved in these feature-based methods remains complex and limits generality and flexibility.
[0045] In some embodiments, methods provided herein comprise using deep learning to drive ADMET prediction, automatically extracting meaningful features from simple input data. In some embodiments, methods provided herein comprise using various neural network architectures, including transformers (designed to effectively handle sequential data), convolutional neural networks (a type of deep learning model commonly used for image and video recognition tasks) and, graph neural networks (deep learning models for processing graph-structured data, such as molecular structures), for modeling molecular properties from formatsAttorney Docket No. 72194-701.602such as SMILES strings and molecular graphs. Among them, SMILES strings offer compact molecular representation and can distinctly express substructures like branches, rings and chirality, but lack topological awareness — whereas graph neural networks (like the GeoGNN model) incorporate geometric information, providing superior performance in ADMET prediction. In some embodiments, methods provided herein comprise using transformer models using SMILES input for structure recognition. In some embodiments, methods provided herein do not comprise using transformer models using SMILES input for structure recognition, since for predictions involving properties like toxicity, the performance of representations generated by these models might saturate before training progresses, showing limited improvement after training.
[0046] Despite the advances propelled by novel deep learning algorithms, the field still faces challenges. High costs and considerable time investments lead to scarce labeled data in ADMET predictions, leading to potential overfitting. Unsupervised and self-supervised learning offers solutions, and while large transformer-based models show promise in other fields, their use in ADMET prediction remains underexplored. In some embodiments, methods provided herein comprise using carefully designed self-supervised training with contextual transformers equipped with linear attention mechanisms to effectively learn implicit structureproperty relationships, bolstering confidence in applying large-scale self-supervised models for ADMET predictions.
[0047] Furthermore, molecular representation is critical for Al performance. Highdimensional representations typically provide richer information than low-dimensional ones. In some embodiments, methods provided herein comprise integrating multiple levels of molecular representation to substantially enhance learning, leading to more comprehensive, generalizable and robust ADMET prediction. In some embodiments, methods provided herein comprise using multimodal ADMET models using multiple representations simultaneously.
[0048] In some embodiments, methods provided herein comprise understanding model parameters in ADMET predictions to help reveal the relationships between molecular substructures and properties. Attention mechanisms, which allow a model to focus on important parts of the input data, can enhance interpretability by identifying key atoms or groups.Integrating chemical knowledge can further enhance interpretability and expand models to achieve comprehensive chemical understanding.Attorney Docket No. 72194-701.602Synthesis planning and automating synthesis and drug discovery
[0049] Chemical synthesis, one of the bottlenecks in small-molecule drug discovery, is a highly technical and extremely laborious task. Computer-aided synthesis planning (CASP) and automatic synthesis of organic compounds can help alleviate the burden of repetitive laborious tasks for chemists, enabling them to engage in more innovative works. With the rapid development of Al, the pharmaceutical industry and academia are becoming increasingly interested in achieving intelligence and automation in this process.
[0050] CASP has been used as a tool to assist chemists in determining reaction routes via retrosynthesis analysis, a problem-solving technique in which target molecules are recursively transformed into increasingly simpler precursors (FIG.3). Early CASP programs were rule based (for example, logic and heuristics applied to synthetic analysis, simulation and evaluation of chemical synthesis and retrosynthesis-based assessment of synthetic accessibility programs). Since then, a range of machine learning techniques, particularly deep learning models, have been developed — yielding gradual improvements in the synthesis planning of artificial small molecules and natural products. In some embodiments, methods provided herein comprise applying the transformer model to retrosynthetic analysis, prediction of regioselectivity (the preference of a chemical reaction to occur at one particular location over another on a molecule with multiple possible reactive sites) and stereoselectivity (the preference of a reaction to produce one stereoisomer over another when multiple stereoisomeric products are possible and reaction fingerprint extraction. Concerns regarding the adequacy of purely data-driven Al methods for complex synthesis planning have spurred the development of hybrid expert-AI systems that incorporate chemical rules. Most current deep learning approaches, however, are unexplainable, showing as ‘black boxes’ that offer limited insights. In some embodiments, methods provided herein comprise using a retrosynthesis prediction model with an interpretable deep learning framework that reframes the retrosynthesis task as a molecular assembly process. In some embodiments, the retrosynthesis prediction model has superior performance compared to state-of-the-art retrosynthesis methods. Notably, its molecular assembly approach enhances interpretability, enabling transparent decision-making and quantitative attribution.
[0051] Automated synthesis of organic compounds represents a cutting-edge frontier in the field of chemistry-related fields (FIG. 3), including medicinal chemistry. An optimal automated synthesis platform would seamlessly integrate and streamline various components of the chemical development process, including CASP as well as automated experiment setup andAttorney Docket No. 72194-701.602optimization, and robotically executed chemical synthesis, separation and purification. In some embodiments, methods provided herein comprise using deep learning-powered automated flow chemistry and solid-phase synthesis techniques for pharmaceutical compound synthesis which have gained considerable attention. In particular, automated synthesis combined with designing, testing and analyzing technologies forms an automated central process of drug discovery called the design-make-test-analysis (DMTA) cycle. By leveraging deep learning, the efficiency of the DMTA cycle has been substantially improved, accelerating the discovery of hit and lead compounds for drug discovery. For example, by using an Al-powered DMTA platform with deep learning for molecular design and microfluidics for on-chip chemical synthesis, liver X receptor agonists were generated from scratch. In addition, LLMs are believed to ‘understand’ human natural language, enabling automation platforms to provide tailored solutions for specific challenges based on concise inputs from researchers. Although automated synthesis and the automated DMTA cycle have great prospects, their development is still in the infancy stage. Many technical challenges remain, including requirements to reduce solid formation to avoid blockage, predict solubility in nonaqueous solvents and at different temperatures, estimate optimal purification methods and optimize multistep reactions.
[0052] Following the planning and synthesis of new drug compounds, Al technology facilities the in vivo validation of the mechanism of action (MOA) of new drugs. In high-content screening, by monitoring the real-time changes in omics data, Al technologies would generalize these features and develop a model that is capable of deciphering the molecular and cellular MOA of a new compound and its associated pharmacokinetics, pharmacodynamics, toxicology and bioavailability properties (FIG. 4).Al in clinical trials and real-world practice
[0053] Al is increasingly guiding various aspects of clinical trials by analyzing patient data, including genetic information, clinical history and lifestyle factors. Applying Al methods to such data helps to identify biomarkers and patient characteristics that influence drug responses, enabling more efficient and informative trial designs. By optimizing parameters like patient selection, treatment regimens and outcome measurements, Al has the potential to boost trial success and accelerate the translation of candidate drugs into clinical practice. Real world data also offer a rich source of information from which Al applications can predict adverse events, drug-drug interactions and other outcomes. The sections below describe key applications of Al to the clinical stages of drug development.Attorney Docket No. 72194-701.602I. Biomarker discovery
[0054] Biomarkers serve as biological indices for objectively measuring and assessing normal versus pathological processes and responses to treatment, holding massive utility across medicine, biotechnology and biopharmaceuticals. However, traditional hypothesis-driven biomarker discovery methods are often inefficient while failing to address disease complexity comprehensively. These methods are time consuming and require substantial resources for hypothesis validation, while the constraints of limited sample sizes hinder broad validation across diverse populations.
[0055] Recent advancements in Al have dramatically enhanced biomarker discovery. Al models excel in identifying diagnostic biomarkers, offering predictive insights and diagnostic references for clinical pathology. A notable example is the ‘nuclei.io’ digital pathology framework, which merges active learning with real-time human-computer interaction. This helps to provide precise feedback to pathologists based on nuclear statistical data, by efficiently building datasets and Al models for various surgical pathology tasks, substantially boost-ing diagnostic accuracy and efficiency.
[0056] Al further excels in identifying prognostic biomarkers crucial for predicting disease progression and patient survival, thus enabling targeted and personalized treatments. For example, deep learning models can delineate CD8+T cell morphology in blood samples as effective sepsis prognostic indicators, differentiate nuclear features marking cellular senescence and identify proteomic biomarkers to accurately predict liver disease outcomes. Al also predicts prognostic biomarkers for various cancers, delivering precise risk scores for survival, recurrence and metastasis. Notably, survival analysis models using graph neural networks outperform existing models, effectively distinguishing risk groups beyond traditional clinical grading and staging - emphasizing Al’s potential in prognosis enhancement and the critical collaboration between pathologists and Al.
[0057] In drug development, identifying predictive biomarkers is crucial to enhancing research success by selecting patient populations most likely to benefit from treatments. These discoveries demand rigorous prospective clinical validation. Although AI-based predictive biomarkers have not yet been applied in the clinic, proof-of-concept studies indicate that Al can forecast patient responses to therapies by predicting known biomarkers such as microsatellite instability. The complexity of biological systems necessitates the integration of multiple types of biological data, including protein-protein interactions, into Al models for more comprehensive predictions.Attorney Docket No. 72194-701.602
[0058] Faced with a scarcity of large, labeled datasets, researchers are deploying diverse strategies to optimize the use of Al in biomarker discovery. The integration of datasets from multiple sources shows great promise. Digital biomarkers from wearable sensors also expand the scope of discovery, by providing rich, longitudinal datasets. Identifying multimodal biomarkers through molecular diagnostics, radiomics and histopathological imaging provides new avenues for precision medicine. Moreover, swarm learning and automated dataset processing pipelines lay the groundwork for large-scale, secure data collection.
[0059] Nonetheless, Al models face challenges relating to heterogeneity that impede their translational efficiency to clinical trials. Some studies use deep learning to elucidate cellular and tissue-level heterogeneity and the diversity of tumor ecosystems, offering new avenues for disease subtype classification and patient stratification. Interpretability and trust are critical for clinical acceptance of Al models and can be enhanced by integrating prior medical knowledge or embedding bio-logical relationships into neural networks. Addressing bias in AI-driven biomarker discovery requires strategies such as validating models across geographically diverse patient cohorts and developing fair and transparent algorithms. Robust validation and responsible data curation will facilitate biomarker identification and application, supporting future drug development and disease treatment.II. Predicting pharmacometrics properties
[0060] Applying Al and big data tools can effectively address pharmacometric problems and offer a powerful tool for time-to-event analysis, particularly in handling highdimensional data and nonlinear relationships in the hazard function. Al supports personalized treatment by optimizing dose-response relationships, improving drug safety profiles and refining therapeutic windows, which are central to resolving pharmacometric problems in precision medicine. Machine learning-based analysis of small-molecule kinases and adverse events enabled the discovery of novel kinase-adverse event pairs, enabling risk mitigation and the development of safer small-molecule kinase inhibitors. The multi-omics variational autoencoders (MOVE) framework integrates multi-omics data to reveal drug interactions — such as the link between metformin and the gut microbiota — and compares drug responses across various omics modalities. PharmBERT, a domain-specific language model, enhances drug safety by extracting crucial pharmacokinetics information from prescription labels, helping to identify adverse reactions and drug interactions. Al also optimizes drug dosages by analyzing genetic and physiological data, leading to personalized treatment recommendations that improve outcomes. Furthermore, Al can analyze patients’ genetic information, physiologicalAttorney Docket No. 72194-701.602characteristics and past treatment responses to provide personalized dosage adjustment recommendations for doctors, thereby optimizing treatment outcomes.III. Drug repurposing
[0061] In addition to new drug discovery, Al contributes to the drug pool by repurposing existing, approved drugs using large-scale biomedical datasets, thus speeding up the development of optimal treatments for various diseases. By discovering previously unidentified therapeutic properties of approved drugs, Al reduces both the time and cost associated with drug discovery. For instance, Al has accelerated the repurposing of drugs for coronavirus disease 2019, high- lighting the value of Al in finding brand new applications for existing medications. Al can also simulate clinical trials using real-world data (including EHRs and insurance claims) to facilitate drug repurposing. As an example of this approach, a deep learning recurrent neural network used causal inference and deep learning to analyze medical claims databases, effectively identifying potential drug candidates. Applied to a cohort of millions with coronary artery disease, it pinpointed drugs and combinations that enhanced outcomes.
[0062] Another deep learning-based approach to drug repurposing involves the application of deep neural networks to omics data, to classify drugs into therapeutic categories based on the transcriptional perturbations that they induce in vitro. One study leveraged perturbation samples from the LINCS Project (https: / / lincsproject.org / ) and 12 therapeutic categories derived from MeSH, resulting in high classification accuracy — especially with pathway-level data across diverse biological systems and conditions — offering potential for drug repositioning. Feature attribution techniques, combined with ensembles of interpretable machine learning models, enhance the identification of gene expression signatures associated with synergistic drug responses. This strategy has been shown to improve feature interpretability and support the selection of optimal anticancer drug combinations informed by molecular insights.
[0063] Furthermore, Al-based high-content screening could also be applied in drug repurposing (FIG. 4). A deep learning model, MitoRelD, was developed to identify MOAs through mitochondrial phenotyping. It offers a cost-effective, high-throughput solution for drug discovery and repurposing, validated with unseen drugs (which were not part of the training set) and validated in vitro. By analyzing 570,096 cell images, MitoRelD achieved 76.32% accuracy in identifying MOAs of US Food and Drug Administration-approved drugs and successfully validated the cyclooxygenase-2 inhibition of epicatechin, a natural compound in tea.Nevertheless, many challenges experienced in other stages of Al-driven drug development apply to drug repurposing, including issues with data quality, model interpretability, generalizability,Attorney Docket No. 72194-701.602validation costs, regulatory hurdles, integration with existing pipelines and high computational demands, which hinder widespread adoption and practical implementation.IV. Improving trial efficiency and predicting outcomes
[0064] Clinical trials are often expensive, time consuming and inefficient, with the majority facing delays in registration or struggling to find sufficient volunteers. Al has the potential to optimize trial design, streamline recruitment and predict patient responses, improving trial efficiency and success rates while reducing costs and timelines. An advanced pipeline has been created that integrates multimodal datasets, generates molecular leads using Al, ranks them by efficacy and safety, and uses deep reinforcement learning to create patentable analogs for testing. It also predicts phase I / I I clinical trial outcomes by estimating side effects and pathway activation, improving prediction accuracy and identifying potential risks in drug portfolios. In real-world studies, Al can analyze data from EHRs, insurance claims and wearable devices to assess drug effectiveness and safety (FIG. 5). For example, a study using real-world data and the Trial Pathfinder tool simulated trial outcomes from EHR data of 61,094 patients with advanced lung cancer, revealing that relaxing trial criteria could double eligible patients and improve survival outcomes. This approach, validated across various cancers, supports more inclusive, safer trials.
[0065] The challenge of finding suitable patients who meet inclusion criteria can be mitigated by using Digital Twins, as explored by Unlearn. ai. This technology creates virtual replicas of participants, allowing them to serve as the control group, thereby increasing the number of participants in the experimental group and improving trial efficiency. Unlearn. ai secured a US$12 mil- lion grant in April 2020 to advance this application (https: / / medcitynews.com / 2020 / 04 / funding-roundup-company- creating-digital-twins-for-clinical-trials-raises-12m / ), and other companies like Novadiscovery and Jinkō are conducting digital twin-based clinical trial simulations for diseases such as lung cancer (https: / / www.novainsilico.ai / new-demonstration-of-the-predictive-power-of-an-in-silico-clinical-trial-in-oncology / ). The proposed approach uses computer modeling grounded in gene expression and clinical data, incorporating deep learning and generative adversarial networks. By leveraging diverse health metrics, these Digital Twins offer quantitative insights into vital processes, deliver dynamic health guidance and optimize treatment strategies. This approach seeks to deepen the mathematical understanding of biological mechanisms, revolutionize clinical practices and fully personalize medical care, for example, by generating patient-specific models that predict survival probabilities based on drug inputs. These models can also enable the simulation of clinical trials and optimize trial parameters, enhancing the likelihood of success.Attorney Docket No. 72194-701.602But they are not without their challenges, including high computational costs, tricky workflow integration, ethical concerns and limited personalization. These issues impact patient simulation accuracy, trial designs and regulatory acceptance, thereby slowing innovation.
[0066] Beyond the clinical trial stage of drug development, Al can also analyze postmarket surveillance data to support the safety, efficacy and quality of drugs. The development and concurrent use of alternative methods to identify and address safety concerns early in the regulatory review process are essential for advancing regulatory science and optimizing drug development.
[0067] A key challenge is the lack of high-quality training data, due to high acquisition costs, privacy regulations and limited data sharing — especially for rare diseases or novel drug targets — which hinders the effectiveness of Al in identifying targets, biomarkers and other functions. Furthermore, available data often suffer from missing information, errors and biases, further reducing Al reliability. Drug discovery experiments can produce inconsistent results, and cost-saving measures may lead to incomplete data. Additionally, underrepresentation of ‘negative’ data (for example, unsuccessful experiments and negative trial outcomes) in the literature hinders a complete understand-ing of drug-target-disease interactions, efficacy and other clinical characteristics.
[0068] A key challenge in drug design is balancing multiple objectives for success. Current research often focuses too much on the chemical space, neglecting other key factors (such as druggability and synthesizability). While multi -objective design methods are improving, developing effective scoring functions (for example, for affinity pre-diction and bioactivity) remains complex and requires considerable experimentation. The absence of standardized evaluation processes further complicates model assessment, especially when conflicting objectives arise, such as maximizing similarity to known bioactive molecules while achieving structural novelty. Although benchmarking platforms like MOSES and Guacamol exist, a consensus on best practices has not yet been reached.
[0069] Appropriate molecular representation is key in generative models. Traditional methods like SMILES and graphs are common and are being complemented by new, emerging data-driven approaches like hierarchical molecular graph self-supervised learning. Nevertheless, capturing complexity and ensuring synthesizability are difficult. Current methods for assessing synthetic feasibility are often imprecise, leading to discovery of unsynthesizable molecules. The integration of reaction knowledge into molecular generation shows promise, but needs improvement. Issues such as model interpretability, uncertainty in generating new molecules andAttorney Docket No. 72194-701.602bias have become focal points of academic interest. Effectively integrating bias control with uncertainty estimation is essential for improving the quality of generated molecules.
[0070] Al faces challenges with so-called ‘undruggable’ targets that lack suitable binding sites, including certain disordered proteins, transcription factors (such as MYC and IRF4) and protein-protein interactions. New Al methods and high-content screening (FIG. 4) to explore their conformational space and identify ligand binding sites could help overcome these obstacles.
[0071] Finally, technical challenges with algorithms and computing power limit Al’s use in drug development. Many Al algorithms used in drug development were designed for other fields and may not be fully suitable; for example, new algorithms based on NLP are needed to capture 3D spatial interactions. Additionally, the high computational resources required by Al approaches pose barriers, especially for smaller research teams. Collaboration with cloud providers and developing more efficient algorithms can help address these challenges. Further, Al drug development faces talent shortages and investment risks due to long cycles, low success rates and uncertain returns, affecting investor confidence.
[0072] Al is revolutionizing the process of drug development by extracting critical insights from complex multi-omics biomedical data, identifying new biomarkers and detecting therapeutic targets and anomalies to facilitate the discovery of lead compounds and drug candidates. Additionally, Al accelerates drug discovery, repurposing and toxicity prediction, thus reducing time, cost and safety risks. However, the journey toward fully realizing AL enabled advancements in this area is ongoing, with many challenges to overcome and potential to be realized. Future efforts to address the challenges mentioned above should place particular emphasis on several key directions.
[0073] First, developing new strategies to address the data scarcity issue in AI-enabled drug development should be the top priority. Feasible strategies to enhance data sharing, establish data standards and develop new Al algorithms — such as ‘sparse’ Al methods that can produce accurate predictions from very limited data — are crucial. Multimodal pretrained models that integrate textual and chemical information offer promise in addressing the data scarcity, especially in zero-shot scenarios234. By integrating an array of data such as genomics, transcriptomics, disease-specific molecular pathways, protein interactions and clinical records, Al can also identify existing drugs with potential repurposing opportunities for neglected or orphan diseases
[0074] Current methods typically focus on single data types, thereby missing complex interrelations between various biological systems. Establishing effective multimodalAttorney Docket No. 72194-701.602fusion approaches can extract valuable insights from diverse sources and formats to advance drug development. With the rise of big data and GPU computing (based on a graphics processing unit, rather than a conventional central processing unit, or CPU), Al can now be applied to various data forms, including text, images and videos. Emerging models using omics data, including deep learning-based drug classification, show promise in drug efficacy prediction, mechanism identification and toxicity assessment, highlighting the future potential of multimodal Al in drug development.
[0075] Many current Al models are purely data driven, limiting their effectiveness in drug development due to the relative lack of sufficiently high-quality data. Because life systems all adhere to the principles of physics (also called the First Principle), drugs also follow the constraints of physical laws without exception. Incorporating physical laws into existing data-driven Al algorithms is one future research direction that could help reduce data dependency and improve both the accuracy and generalizability of these models.
[0076] Al, especially LLMs, can ensure compliance with drug regulations by analyzing extensive documentation and keeping up with the latest requirements. This boosts efficiency, reduces risks of non-compliance and prevents delays in drug approval. Developing not just accurate, but also interpretable Al models is essential for build-ing trust among drug developers, regulatory agencies, clinicians and patients by ensuring transparency and understanding in the decision-making process. These models can be incorporated early to optimize project funding and guide investments to accelerate drug development.
[0077] In the coming decades, Al’s role in medical modeling and simulation will be transformative. Advanced Al models will create ever more detailed virtual human simulations, further enhancing the understanding of disease mechanisms, drug actions and individual biological differences. Through simulations, Al can streamline clinical trial design and execution, testing different scenarios for best selection criteria to accelerate patient recruitment and enhance trial representativeness. Al will also provide personalized medical decision support by analyzing health data and genomics, enabling precise risk predictions, optimized treatments and improved surgical guidance. Medical education will benefit from Al-driven virtual reality, offering more realistic training scenarios and enhancing the quality of medical services.
[0078] Overall, the continuing advancements in Al technologies are substantially improving the efficiency and cost-effectiveness of drug development. However, it is essential to recognize that Al is not omnipotent. The strength of Al technologies lies in analyzing big and complex data and aiding quick decision-making to complement human functions and augment human capabilities, but Al is not designed to entirely replace human ingenuity or authority. TheAttorney Docket No. 72194-701.602drugs designed and properties predicted by Al still require validation through wet-lab experiments, and human input will still be needed to determine the direction of Al research and use.
[0079] According to a first aspect of the present disclosure, provided herein a method for predicting drug mechanism of action based on artificial intelligence (Al), comprising: (1) a step of treatment of drugs with known MOAs and data acquisition: preparing a variety of drugs with known mechanisms of action (MOAs) or signaling pathways, and a genome-wide gene expression perturbation tool; applying the drugs with known MOAs to cells cultured in multiwell plates, and optionally using the genome-wide gene expression perturbation tool to treat different wells to obtain initial data of the cells under drug action and information on gene expression changes; (2) a step of multi -omics detection and analysis comprising: conducting multi -omics detection on the cells treated in step 1; performing marker staining on the cells to detect the expression and localization of specific proteins or molecules within the cells through specific markers; performing morphological imaging on the cells to obtain morphological data; and integrating the multi-omics detection data, marker staining data, and cell morphological imaging data, and combining the existing drug MOA knowledge and gene signaling pathway knowledge for analysis; (3) a step of Al model training: inputting the data analyzed in step 2 into an Al system, and using the data to train the artificial intelligence model, enabling the artificial intelligence model to learn the relationships and regularities between drug characteristics and known MOAs, said drug characteristics including multi-omics characteristics and cell morphological characteristics; (4) a step of prediction of unknown drug MOAs after the training of the Al model is completed, applying drugs with unknown MOAs to the cells, and obtaining the multi-omics data and cell morphological data of the cells after treatment with the drugs with unknown MOAs; inputting the obtained data after treatment with drugs with unknown MOAs into the trained Al model, and through the Al model, comparing and analyzing the input data characteristics with the previously learned patterns, thereby predicting the possible MOAs of the unknown drugs.
[0080] In some embodiments, the Al model training comprises embedding the multi- omics data from each well via a generative machine learning model to generate one or more multi-omics data representation embeddings. In some embodiments, the one or more multi-omics data representation embeddings can comprise a vector representation of multi-omics data generated by a machine learning model.
[0081] In some embodiments, the Al model training comprises embedding the marker staining data from each well via a generative machine learning model to generate one orAttorney Docket No. 72194-701.602more marker staining data representation embeddings. In some embodiments, the one or more marker staining data representation embeddings can comprise a vector representation of marker staining data generated by a machine learning model.
[0082] In some embodiments, the Al model training comprises embedding the cell morphological imaging data from each well via a generative machine learning model to generate one or more cell morphological imaging data representation embeddings. In some embodiments, the one or more cell morphological imaging data representation embeddings can comprise a vector representation of cell morphological imaging data generated by a machine learning model.
[0083] In some embodiments, the Al model training comprises (i) embedding the multi -omics data from each well via a generative machine learning model to generate one or more multi -omics data representation embeddings, (ii) embedding the marker staining data from each well via the generative machine learning model to generate one or more marker staining data representation embeddings, and (iii) embedding the cell morphological imaging data from each well via the generative machine learning model to generate one or more cell morphological imaging data representation embeddings. In some embodiments, the Al model training comprises (i) generating one or more multi -omics data representation embeddings comprising a vector representation of multi -omics data generated by a machine learning model, (ii) generating one or more marker staining data representation embeddings comprising a vector representation of marker staining data generated by the machine learning model, and (iii) generating one or more cell morphological imaging data representation embeddings comprising a vector representation of cell morphological imaging data generated by a machine learning model.
[0084] In some embodiments, the Al model training comprises generating multi¬ modal representation embeddings in a feature space shared by the multi-modal representation embeddings. In some embodiments, the Al model training comprises generating multi-modal representation embeddings in a low dimensional feature space. In some embodiments, the feature space comprises a framework (or mapping) that represents one or more types of data or modalities in a common format or space. In some embodiments, the feature space comprises one or more multi-omics data representation embeddings, one or more marker staining data representation embeddings, and one or more cell morphological imaging data representation embeddings in a unified representation to enable comparisons, analysis, and / or learning across the different types of representation embeddings. In some embodiments, the feature space comprises a collection of vector representations (or other values) comprising a vector representation of multi-omics data, a vector representation of marker staining data, and a vectorAttorney Docket No. 72194-701.602representation of cell morphological imaging data in a unified representation to enable comparisons, analysis, and / or learning across the multi-modal representation embeddings.
[0085] In some embodiments, the Al model training further comprises annotating any one or more of (i) one or more multi-omics data representation embeddings, (ii) one or more marker staining data representation embeddings, and (iii)one or more cell morphological imaging data representation embeddings, with an MOA label corresponding to an MOA. In some embodiments, themodel training further comprises generating an MOA representation for the MOA by generating an embedding cluster within a feature space comprising (i) one or more multi -omics data representation embeddings, (ii) one or more marker staining data representation embeddings, and (iii)one or more cell morphological imaging data representation embeddings, each of which representation embeddings is annotated with an MOA label.
[0086] In some embodiments, the genome-wide gene expression perturbation tool is one or more selected from the group consisting of siRNA and CRISPR technology.
[0087] In some embodiments, the multi-omics detection includes one or more selected from the group consisting of genomics, transcriptomics, proteomics, and metabolomics detection.
[0088] In some embodiments, the morphological data includes shape, size, and / or structure of the cells.
[0089] In some embodiments, in step (2), the integrated data is analyzed to mine the potential patterns and regularities therein, thereby providing a data foundation for the training of Al model.
[0090] In some embodiments, in step (3), the drug characteristics includes multi- omics characteristics and / or cell morphological characteristics.
[0091] According to a second aspect of the present disclosure, provided rameiteon for use in treating an estradiol related di sease.
[0092] In some embodiments, the disease comprises premature ovarian insufficiency (POI) and diseases related to female reproductive system aging.
[0093] According to a third aspect of the present disclosure, provided use of rameiteon in preparation of a medicament for treating an estradiol related disease.
[0094] In some embodiments, the disease comprises premature ovarian insufficiency (POI) and diseases related to female reproductive system aging.
[0095] According to a fourth aspect of the present disclosure, provided a method for treating an estradiol related disease in a subject in need thereof, comprising administrating to the subject an effective amount of rameiteon.Attorney Docket No. 72194-701.602
[0096] In some embodiments, the disease comprises premature ovarian insufficiency (POI) and diseases related to female reproductive system aging.
[0097] According to a fifth aspect of the present disclosure, a method for constructing an artificial intelligence-based drug evaluation system is proposed. The method may comprise the following steps: acquiring data from sources, for example, electronic health records (EHRs), insurance claims, and / or wearable devices; performing multimodal data embedding processing on the acquired multiple types of data to generate a unified embedded data representation; constructing a foundation model based on the embedded data representation; performing generative artificial intelligence processing based on the foundation model to obtain an artificial intelligence large language model for evaluating the effectiveness and safety of drugs; comparing the output of the artificial intelligence large language model with real clinical data to perform prediction quality control, fine-tuning the system based on the comparison results to further predict therapeutic response, survival intervals, and adverse events, and finally obtaining a trained artificial intelligence large language model.
[0098] Preferably, the step of acquiring data may further comprise the step of cleaning, standardizing, and normalizing the acquired data.
[0099] Preferably, a deep learning algorithm may be adopted for the multimodal data embedding processing.
[0100] Preferably, the foundation model may include a neural network model for feature extraction and learning of the embedded data.
[0101] Preferably, the obtained artificial intelligence large language model in the step of performing generative artificial intelligence processing is configured to generate evaluation reports and optimization suggestions regarding drug efficacy and protocol feasibility.
[0102] According to a sixth aspect of the present disclosure, an artificial intelligencebased drug evaluation method is proposed. The method may comprise: inputting clinical trial data into an artificial intelligence large language model trained according to the method for constructing an artificial intelligence-based drug evaluation system of the first aspect of the present disclosure; performing a prediction process for drug evaluation in the artificial intelligence large language model to predict therapeutic response, survival intervals, and adverse events according to the clinical trial data.
[0103] According to a seventh aspect of the present disclosure, a computer system for constructing an artificial intelligence-based drug evaluation system is proposed. The computer system may comprise: a processor for executing a program, and a memory for storing the program. When the program is loaded into the processor, the processor implements theAttorney Docket No. 72194-701.602operations of the method for constructing an artificial intelligence-based drug evaluation system according to the fifth aspect of the present disclosure.
[0104] According to an eighth aspect of the present disclosure, a computer readable medium is proposed. The computer readable medium is used for recording instructions executable by a processor. The instructions, when executed by the processor, may cause the processor to perform the method for constructing an artificial intelligence-based drug evaluation system according to the fifth aspect of the present disclosure.
[0105] According to a ninth aspect of the present disclosure, a computer system for artificial intelligence-based drug evaluation is proposed. The computer system may comprise: a processor for executing a program, and a memory for storing the program. When the program is loaded into the processor, the processor implements the operations of the artificial intelligencebased drug evaluation method according to the sixth aspect of the present disclosure.
[0106] According to a tenth aspect of the present disclosure, a computer readable medium is proposed. The computer readable medium is used for recording instructions executable by a processor. The instructions, when executed by the processor, may cause the processor to perform the artificial intelligence-based drug evaluation method according to the sixth aspect of the present disclosure.Computing systems
[0107] Referring to FIG. 10, a block diagram is shown depicting an exemplary machine that includes a computer system 1000 (e.g., a processing or computing system) within which a set of instructions can execute for causing a device to perform or execute any one or more of the aspects and / or methodologies for static code scheduling of the present disclosure. The components in FIG. 10 are examples only and do not limit the scope of use or functionality of any hardware, software, embedded logic component, or a combination of two or more such components implementing particular embodiments.
[0108] Computer system 1000 may include one or more processors 1001, a memory 1003, and a storage 1008 that communicate with each other, and with other components, via a bus 1040. The bus 1040 may also link a display 1032, one or more input devices 1033 (which may, for example, include a keypad, a keyboard, a mouse, a stylus, etc.), one or more output devices 1034, one or more storage devices 1035, and various tangible storage media 1036. All of these elements may interface directly or via one or more interfaces or adaptors to the bus 1040. For instance, the various tangible storage media 1036 can interface with the bus 1040 via storage medium interface 1026. Computer system 1000 may have any suitable physical form, includingAttorney Docket No. 72194-701.602but not limited to one or more integrated circuits (ICs), printed circuit boards (PCBs), mobile handheld devices (such as mobile telephones or PDAs), laptop or notebook computers, distributed computer systems, computing grids, or servers.
[0109] Computer system 1000 includes one or more processor(s) 1001 (e.g., central processing units (CPUs) or general purpose graphics processing units (GPGPUs)) that carry out functions. Processor(s) 1001 optionally contains a cache memory unit 1002 for temporary local storage of instructions, data, or computer addresses. Processor(s) 1001 are configured to assist in execution of computer readable instructions. Computer system 1000 may provide functionality for the components depicted in FIG. 10 as a result of the processor(s) 1001 executing non-transitory, processor-executable instructions embodied in one or more tangible computer-readable storage media, such as memory 1003, storage 1008, storage devices 1035, and / or storage medium 1036. The computer-readable media may store software that implements particular embodiments, and processor(s) 1001 may execute the software. Memory 1003 may read the software from one or more other computer-readable media (such as mass storage device(s) 1035, 1036) or from one or more other sources through a suitable interface, such as network interface 1020. The software may cause processor(s) 1001 to carry out one or more processes or one or more steps of one or more processes described or illustrated herein. Carrying out such processes or steps may include defining data structures stored in memory 1003 and modifying the data structures as directed by the software.
[0110] The memory 1003 may include various components (e.g., machine readable media) including, but not limited to, a random access memory component (e.g., RAM 1004) (e.g., static RAM (SRAM), dynamic RAM (DRAM), ferroelectric random access memory (FRAM), phase-change random access memory (PRAM), etc.), a read-only memory component (e.g., ROM 1005), and any combinations thereof. ROM 1005 may act to communicate data and instructions unidirectionally to processor(s) 1001, and RAM 1004 may act to communicate data and instructions bidirectionally with processor(s) 1001. ROM 1005 and RAM 1004 may include any suitable tangible computer-readable media described below. In one example, a basic input / output system 1006 (BIOS), including basic routines that help to transfer information between elements within computer system 1000, such as during start-up, may be stored in the memory 1003.[OHl] Fixed storage 1008 is connected bidirectionally to processor(s) 1001, optionally through storage control unit 1007. Fixed storage 1008 provides additional data storage capacity and may also include any suitable tangible computer-readable media described herein. Storage 1008 may be used to store operating system 1009, executable(s) 1010, dataAttorney Docket No. 72194-701.6021011, applications 1012 (application programs), and the like. Storage 1008 can also include an optical disk drive, a solid-state memory device (e.g., flash-based systems), or a combination of any of the above. Information in storage 1008 may, in appropriate cases, be incorporated as virtual memory in memory 1003.
[0112] In one example, storage device(s) 1035 may be removably interfaced with computer system 1000 (e.g., via an external port connector (not shown)) via a storage device interface 1025. Particularly, storage device(s) 1035 and an associated machine-readable medium may provide non-volatile and / or volatile storage of machine-readable instructions, data structures, program modules, and / or other data for the computer system 1000. In one example, software may reside, completely or partially, within a machine-readable medium on storage device(s) 1035. In another example, software may reside, completely or partially, within processor(s) 1001.
[0113] Bus 1040 connects a wide variety of subsystems. Herein, reference to a bus may encompass one or more digital signal lines serving a common function, where appropriate. Bus 1040 may be any of several types of bus structures including, but not limited to, a memory bus, a memory controller, a peripheral bus, a local bus, and any combinations thereof, using any of a variety of bus architectures. As an example and not by way of limitation, such architectures include an Industry Standard Architecture (ISA) bus, an Enhanced ISA (EISA) bus, a Micro Channel Architecture (MCA) bus, a Video Electronics Standards Association local bus (VLB), a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, an Accelerated Graphics Port (AGP) bus, HyperTransport (HTX) bus, serial advanced technology attachment (SATA) bus, and any combinations thereof.
[0114] Computer system 1000 may also include an input device 1033. In one example, a user of computer system 1000 may enter commands and / or other information into computer system 1000 via input device(s) 1033. Examples of an input device(s) 1033 include, but are not limited to, an alpha-numeric input device (e.g., a keyboard), a pointing device (e.g., a mouse or touchpad), a touchpad, a touch screen, a multi-touch screen, a joystick, a stylus, a gamepad, an audio input device (e.g., a microphone, a voice response system, etc.), an optical scanner, a video or still image capture device (e.g., a camera), and any combinations thereof. In some embodiments, the input device is a Kinect®, Leap Motion®, or the like. Input device(s) 1033 may be interfaced to bus 1040 via any of a variety of input interfaces 1023 (e.g., input interface 1023) including, but not limited to, serial, parallel, game port, USB, FIREWIRE, THUNDERBOLT, or any combination of the above.Attorney Docket No. 72194-701.602
[0115] In particular embodiments, when computer system 1000 is connected to network 1030, computer system 1000 may communicate with other devices, specifically mobile devices and enterprise systems, distributed computing systems, cloud storage systems, cloud computing systems, and the like, connected to network 1030. Communications to and from computer system 1000 may be sent through network interface 1020. For example, network interface 1020 may receive incoming communications (such as requests or responses from other devices) in the form of one or more packets (such as Internet Protocol (IP) packets) from network 1030, and computer system 1000 may store the incoming communications in memory 1003 for processing. Computer system 1000 may similarly store outgoing communications (such as requests or responses to other devices) in the form of one or more packets in memory 1003 and communicated to network 1030 from network interface 1020. Processor(s) 1001 may access these communication packets stored in memory 1003 for processing.
[0116] Examples of the network interface 1020 include, but are not limited to, a network interface card, a modem, and any combination thereof. Examples of a network 1030 or network segment 1030 include, but are not limited to, a distributed computing system, a cloud computing system, a wide area network (WAN) (e.g., the Internet, an enterprise network), a local area network (LAN) (e.g., a network associated with an office, a building, a campus or other relatively small geographic space), a telephone network, a direct connection between two computing devices, a peer-to-peer network, and any combinations thereof. A network, such as network 1030, may employ a wired and / or a wireless mode of communication. In general, any network topology may be used.
[0117] Information and data can be displayed through a display 1032. Examples of a display 1032 include, but are not limited to, a cathode ray tube (CRT), a liquid crystal display (LCD), a thin film transistor liquid crystal display (TFT-LCD), an organic liquid crystal display (OLED) such as a passive-matrix OLED (PMOLED) or active-matrix OLED (AMOLED) display, a plasma display, and any combinations thereof. The display 1032 can interface to the processor(s) 1001, memory 1003, and fixed storage 1008, as well as other devices, such as input device(s) 1033, via the bus 1040. The display 1032 is linked to the bus 1040 via a video interface 1022, and transport of data between the display 1032 and the bus 1040 can be controlled via the graphics control 1021. In some embodiments, the display is a video projector. In some embodiments, the display is a head-mounted display (HMD) such as a VR headset. In further embodiments, suitable VR headsets include, by way of non-limiting examples, HTC Vive®, Oculus Rift®, Samsung Gear VR®, Microsoft HoloLens®, Razer OSVR®, FOVE VR®,Attorney Docket No. 72194-701.602Zeiss VR One®, Avegant Glyph®, Freefly VR® headset, and the like. In still further embodiments, the display is a combination of devices such as those disclosed herein.
[0118] In addition to a display 1032, computer system 1000 may include one or more other peripheral output devices 1034 including, but not limited to, an audio speaker, a printer, a storage device, and any combinations thereof. Such peripheral output devices may be connected to the bus 1040 via an output interface 1024. Examples of an output interface 1024 include, but are not limited to, a serial port, a parallel connection, a USB port, a FIREWIRE port, a THUNDERBOLT port, and any combinations thereof.
[0119] In addition or as an alternative, computer system 1000 may provide functionality as a result of logic hardwired or otherwise embodied in a circuit, which may operate in place of or together with software to execute one or more processes or one or more steps of one or more processes described or illustrated herein. Reference to software in this present disclosure may encompass logic, and reference to logic may encompass software.Moreover, reference to a computer-readable medium may encompass a circuit (such as an IC) storing software for execution, a circuit embodying logic for execution, or both, where appropriate. The present disclosure encompasses any suitable combination of hardware, software, or both.
[0120] Various illustrative logical blocks, modules, circuits, and algorithm steps described in connection with the embodiments disclosed herein may be implemented as electronic hardware, computer software, or combinations of both. To clearly illustrate this interchangeability of hardware and software, various illustrative components, blocks, modules, circuits, and steps have been described above generally in terms of their functionality.
[0121] The various illustrative logical blocks, modules, and circuits described in connection with the embodiments disclosed herein may be implemented or performed with a general purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. A general purpose processor may be a microprocessor, but in the alternative, the processor may be any conventional processor, controller, microcontroller, or state machine. A processor may also be implemented as a combination of computing devices, e.g., a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration.Attorney Docket No. 72194-701.602
[0122] The steps of a method or algorithm described in connection with the embodiments disclosed herein may be embodied directly in hardware, in a software module executed by one or more processor(s), or in a combination of the two. A software module may reside in RAM memory, flash memory, ROM memory, EPROM memory, EEPROM memory, registers, hard disk, a removable disk, a CD-ROM, or any other form of storage medium. An exemplary storage medium is coupled to the processor such the processor can read information from, and write information to, the storage medium. In the alternative, the storage medium may be integral to the processor. The processor and the storage medium may reside in an ASIC. The ASIC may reside in a user terminal. In the alternative, the processor and the storage medium may reside as discrete components in a user terminal.
[0123] In accordance with the description herein, suitable computing devices include, by way of non-limiting examples, server computers, desktop computers, laptop computers, notebook computers, sub-notebook computers, netbook computers, netpad computers, set-top computers, media streaming devices, handheld computers, Internet appliances, mobile smartphones, tablet computers, personal digital assistants, video game consoles, and vehicles. Select televisions, video players, and digital music players with optional computer network connectivity are suitable for use in the system described herein. Suitable tablet computers, in various embodiments, include those with booklet, slate, and convertible configurations.
[0124] In some embodiments, the computing device includes an operating system configured to perform executable instructions. The operating system is, for example, software, including programs and data, which manages the device’s hardware and provides services for execution of applications. Suitable server operating systems include, by way of non-limiting examples, FreeBSD®, OpenBSD®, NetBSD®, Linux®, Apple® Mac OS X Server®, Oracle Solaris®, Windows Server®, and Novell NetWare®. Suitable personal computer operating systems include, by way of non-limiting examples, Microsoft Windows®, Apple Mac® OS X, UNIX®, and UNIX-like operating systems such as GNU / Linux®. In some embodiments, the operating system is provided by cloud computing. Suitable mobile smartphone operating systems include, by way of non-limiting examples, Nokia Symbian® OS, Apple® iOS, Research In Motion BlackBerry® OS, Google® Android®, Microsoft® Windows Phone® OS, Microsoft® Windows Mobile OS, Linux®, and Palm® WebOS. Suitable media streaming device operating systems include, by way of non-limiting examples, Apple TV®, Roku®, Boxee®, Google TV®, Google Chromecast®, Amazon Fire®, and Samsung® HomeSync®. Suitable video game console operating systems include, by way of non-limiting examples, Sony® PS3®, Sony® PS4®,Attorney Docket No. 72194-701.602Microsoft® Xbox 360®, Microsoft Xbox One®, Nintendo Wii®, Nintendo Wii U®, and Ouya®. Suitable virtual reality headset systems include, by way of non-limiting example, Meta Oculus®.Non-transitory computer readable storage mediums
[0125] In some embodiments, the platforms, systems, media, and methods disclosed herein include one or more non-transitory computer readable storage media encoded with a program including instructions executable by the operating system of an optionally networked computing device. In further embodiments, a computer readable storage medium is a tangible component of a computing device. In still further embodiments, a computer readable storage medium is optionally removable from a computing device. In some embodiments, a computer readable storage medium includes, by way of non-limiting examples, CD-ROMs, DVDs, flash memory devices, solid state memory, magnetic disk drives, magnetic tape drives, optical disk drives, distributed computing systems including cloud computing systems and services, and the like. In some cases, the program and instructions are permanently, substantially permanently, semi-permanently, or non-transitorily encoded on the media.Computer programs
[0126] In some embodiments, the platforms, systems, media, and methods disclosed herein include at least one computer program, or use of the same. A computer program includes a sequence of instructions, executable by one or more processor(s) of the computing device’s CPU, written to perform a specified task. Computer readable instructions may be implemented as program modules, such as functions, objects, Application Programming Interfaces (APIs), computing data structures, and the like, that perform particular tasks or implement particular abstract data types. In light of the present disclosure provided herein, a computer program may be written in various versions of various languages.
[0127] The functionality of the computer readable instructions may be combined or distributed as desired in various environments. In some embodiments, a computer program comprises one sequence of instructions. In some embodiments, a computer program comprises a plurality of sequences of instructions. In some embodiments, a computer program is provided from one location. In other embodiments, a computer program is provided from a plurality of locations. In various embodiments, a computer program includes one or more software modules. In various embodiments, a computer program includes, in part or in whole, one or more web applications, one or more mobile applications, one or more standalone applications, one or more web browser plug-ins, extensions, add-ins, or add-ons, or combinations thereof.Attorney Docket No. 72194-701.602Web applications
[0128] In some embodiments, a computer program includes a web application. In light of the present disclosure provided herein, a web application, in various embodiments, utilizes one or more software frameworks and one or more database systems. In some embodiments, a web application is created upon a software framework such as Microsoft®. NET or Ruby on Rails® (RoR). In some embodiments, a web application utilizes one or more database systems including, by way of non-limiting examples, relational, non-relational, object oriented, associative, and XML database systems. In further embodiments, suitable relational database systems include, by way of non-limiting examples, Microsoft® structured query language (SQL) Server, mySQL™, and Oracle®. A web application, in various embodiments, is written in one or more versions of one or more languages. A web application may be written in one or more markup languages, presentation definition languages, client-side scripting languages, server-side coding languages, database query languages, or combinations thereof. In some embodiments, a web application is written to some extent in a markup language such as Hypertext Markup Language (HTML), Extensible Hypertext Markup Language (XHTML), or extensible Markup Language (XML). In some embodiments, a web application is written to some extent in a presentation definition language such as Cascading Style Sheets (CSS). In some embodiments, a web application is written to some extent in a client-side scripting language such as Asynchronous Javascript and XML® (AJAX), Flash Actionscript, Javascript®, or Silverlight®. In some embodiments, a web application is written to some extent in a server-side coding language such as Active Server Pages® (ASP), ColdFusion®, Perl®, Java®, JavaServer Pages® (JSP), Hypertext Preprocessor® (PHP), Python®, Ruby®, Tel®, Smalltalk®, WebDNA®, or Groovy®. In some embodiments, a web application is written to some extent in a database query language such as Structured Query Language (SQL). In some embodiments, a web application integrates enterprise server products such as IBM Lotus Domino®. In some embodiments, a web application includes a media player element. In various further embodiments, a media player element utilizes one or more of many suitable multimedia technologies including, by way of non-limiting examples, Adobe® Flash®, HTML 5, Apple® QuickTime®, Microsoft Silverlight®, Java®, and Unity®.
[0129] In some embodiments, a computer program includes a mobile application provided to a mobile computing device. In some embodiments, the mobile application is provided to a mobile computing device at the time it is manufactured. In other embodiments, the mobile application is provided to a mobile computing device via the computer network described herein.Attorney Docket No. 72194-701.602
[0130] In view of the present disclosure provided herein, a mobile application is created by techniques using hardware, languages, and development environments. Mobile applications are written in several languages. Suitable programming languages include, by way of non-limiting examples, C, C++, C#, Objective-C, Java®, Javascript®, Pascal®, Object Pascal®, Python™, Ruby®, VB. NET®, WML®, and XHTML / HTML with or without CSS, or combinations thereof.
[0131] Suitable mobile application development environments are available from several sources. Commercially available development environments include, by way of nonlimiting examples, AirplaySDK®, alcheMo®, Appcelerator®, Celsius®, Bedrock®, Flash Lite®,. NET Compact Framework®, Rhomobile®, and WorkLight Mobile Platform®. Other development environments are available without cost including, by way of non-limiting examples, Lazarus®, MobiFlex®, MoSync®, and Phonegap®. Also, mobile device manufacturers distribute software developer kits including, by way of non-limiting examples, iPhone® and iPad® (iOS) SDK, Android® SDK, BlackBerry® SDK, BREW SDK, Palm® OS SDK, Symbian® SDK, webOS® SDK, and Windows® Mobile SDK.
[0132] Several commercial sources are available for distribution of mobile applications including, by way of non-limiting examples, Apple® App Store, Google® Play, Chrome® WebStore, BlackBerry® App World, App Store® for Palm devices, App Catalog® for webOS, Windows® Marketplace for Mobile, Ovi Store for Nokia® devices, Samsung® Apps, and Nintendo® DSi Shop.Standalone applications
[0133] In some embodiments, a computer program includes a standalone application, which is a program that is run as an independent computer process, not an add-on to an existing process, e.g., not a plug-in. Standalone applications are often compiled. A compiler is a computer program(s) that transforms source code written in a programming language into binary object code such as assembly language or machine code. Suitable compiled programming languages include, by way of non-limiting examples, C, C++, Objective-C®, COBOL®, Delphi®, Eiffel®, Java®, Lisp®, Python®, Visual Basic®, and VB. NET®, or combinations thereof.Compilation is often performed, at least in part, to create an executable program. In some embodiments, a computer program includes one or more executable compiled applications. Additionally, microservices related to Python® and JavaScript® may be used.Web browser plug-ins
[0134] In some embodiments, the computer program includes a web browser plug-in (e.g., web extension, etc.). In computing, a plug-in is one or more software components that addAttorney Docket No. 72194-701.602specific functionality to a larger software application. Makers of software applications support plug-ins to enable third-party developers to create abilities which extend an application, to support easily adding new features, and to reduce the size of an application. When supported, plug-ins enable customizing the functionality of a software application. For example, plug-ins are commonly used in web browsers to play video, generate interactivity, scan for viruses, and display particular file types. Several web browser plug-ins may include Adobe Flash Player®, Microsoft Silverlight®, and Apple QuickTime®. In some embodiments, the toolbar comprises one or more web browser extensions, add-ins, or add-ons. In some embodiments, the toolbar comprises one or more explorer bars, tool bands, or desk band.
[0135] In view of the present disclosure provided herein, several plug-in frameworks are available that enable development of plug-ins in various programming languages, including, by way of non-limiting examples, C++, Delphi®, Java®, PHP®, Python®, and VB. NET®, or combinations thereof.
[0136] Web browsers (also called Internet browsers) are software applications, designed for use with network-connected computing devices, for retrieving, presenting, and traversing information resources on the World Wide Web. Suitable web browsers include, by way of non-limiting examples, Microsoft Internet Explorer®, Mozilla Firefox®, Google Chrome®, Apple Safari®, Opera Software Opera®, and KDE Konqueror®. In some embodiments, the web browser is a mobile web browser. Mobile web browsers (also called microbrowsers, mini-browsers, and wireless browsers) are designed for use on mobile computing devices including, by way of non-limiting examples, handheld computers, tablet computers, netbook computers, subnotebook computers, smartphones, music players, personal digital assistants (PDAs), and handheld video game systems. Suitable mobile web browsers include, by way of non-limiting examples, Google Android® browser, RIM BlackBerry® Browser, Apple Safari®, Palm Blazer®, Palm WebOS® Browser, Mozilla Firefox® for mobile, Microsoft Internet Explorer Mobile®, Amazon Kindle Basic Web®, Nokia Browser®, Opera Software Opera Mobile®, and Sony PSP® browser.Software modules
[0137] In some embodiments, the platforms, systems, media, and methods disclosed herein include software, server, and / or database modules, or use of the same. In view of the present disclosure provided herein, software modules are created by techniques using machines, software, and languages. The software modules disclosed herein are implemented in a multitude of ways. In various embodiments, a software module comprises a file, a section of code, a programming object, a programming structure, or combinations thereof. In further variousAttorney Docket No. 72194-701.602embodiments, a software module comprises a plurality of files, a plurality of sections of code, a plurality of programming objects, a plurality of programming structures, or combinations thereof. In various embodiments, the one or more software modules comprise, by way of nonlimiting examples, a web application, a mobile application, and a standalone application. In some embodiments, software modules are in one computer program or application. In other embodiments, software modules are in more than one computer program or application. In some embodiments, software modules are hosted on one machine. In other embodiments, software modules are hosted on more than one machine. In further embodiments, software modules are hosted on a distributed computing platform such as a cloud computing platform. In some embodiments, software modules are hosted on one or more machines in one location. In other embodiments, software modules are hosted on one or more machines in more than one location.Databases
[0138] In some embodiments, the platforms, systems, media, and methods disclosed herein include one or more databases (DB), or use of the same. In view of the present disclosure provided herein, many databases are suitable for storage and retrieval data. In various embodiments, suitable databases include, by way of non-limiting examples, relational databases, non-relational databases, object oriented databases, object databases, entity-relationship model databases, associative databases, XML databases, time-series databases, graph databases, and the like. Further non-limiting examples include SQL, PostgreSQL®, MySQL®, Oracle®, DB2®, and Sybase. In some embodiments, a database is internet-based. In further embodiments, a database is web-based. In still further embodiments, a database is cloud computing-based. In a particular embodiment, a database is a distributed database. In other embodiments, a database is based on one or more local computer storage devices.
[0139] Unless defined otherwise, all terms of art, notations and other technical and scientific terms or terminology used herein are intended to have the same meaning as is commonly understood by one of ordinary skill in the art to which the claimed subject matter pertains. In some cases, terms with commonly understood meanings are defined herein for clarity and / or for ready reference, and the inclusion of such definitions herein should not necessarily be construed to represent a substantial difference over what is generally understood in the art.
[0140] The term "about" as used herein refers to the usual error range for the respective value readily known to the skilled person in this technical field. Reference to "about" a value or parameter herein comprises (and describes) embodiments that are directed to that value or parameter per se. In some instances, the term “about” can refer to a value within 20% ofAttorney Docket No. 72194-701.602an indicated value. In some instances, the term “about” can refer to a value within 10% of an indicated value.
[0141] As used herein, the singular forms "a," "an," and "the" comprise plural referents unless the context clearly dictates otherwise. For example, "a" or "an" means "at least one" or "one or more."
[0142] Throughout this disclosure, various aspects of the claimed subject matter are presented in a range format. It should be understood that the description in range format is merely for convenience and brevity and should not be construed as an inflexible limitation on the scope of the claimed subject matter. Accordingly, the description of a range should be considered to have specifically disclosed all the possible sub-ranges as well as individual numerical values within that range. For example, where a range of values is provided, it is understood that each intervening value, between the upper and lower limit of that range and any other stated or intervening value in that stated range is encompassed within the claimed subject matter. The upper and lower limits of these smaller ranges may independently be comprised in the smaller ranges, and are also encompassed within the claimed subject matter, subject to any specifically excluded limit in the stated range. Where the stated range comprises one or both of the limits, ranges excluding either or both of those comprised limits are also comprised in the claimed subject matter. This applies regardless of the breadth of the range.
[0143] Use of ordinal terms such as “first”, “second”, “third”, etc., in the claims to modify a claim element does not by itself connote any priority, precedence, or order of one claim element over another or the temporal order in which acts of a method are performed, but are used merely as labels to distinguish one claim element having a certain name from another element having a same name (but for use of the ordinal term) to distinguish the claim elements. Similarly, use of a), b), etc., or i), ii), etc. does not by itself connote any priority, precedence, or order of operations in the claims. Similarly, the use of these terms in the specification does not by itself connote any required priority, precedence, or order.EXAMPLES
[0144] The following examples are provided to further illustrate some embodiments of the present disclosure, but are not intended to limit the scope of the disclosure; it will be understood by their exemplary nature that other procedures, methodologies, or techniques known to those skilled in the art may alternatively be used.Attorney Docket No. 72194-701.602High-content screen for drugs preventing premature ovarian insufficiency
[0145] KGN cells (Procell, STR tested) were maintained in DMEM / F12 medium (HyClone) with 10% FBS (Gibco) and 1% penicillin / streptomycin mixture (Gibco). For a phenotypic screen, 105cells were seeded in each well of a 24-well plate. A library of 3000 FDA approved compounds were screened using an AI method disclosed herein. The KGN cells were treated with ceramides (MCE), DMSO or approved drugs for 48 hours and stimulated with FSH (MCE) for 24 hours, before the medium was collected to measure the concentration of E2 by an ELISA kit (Cusabio). The details of measuring E2 secretion from KGN cells have been described in Guo et al., “BCAA insufficiency leads to premature ovarian insufficiency via ceramide-induced elevation of ROS,” EMBO Mol Med. 2023; 15:el7450. The chemical library for phenotypic screen was designed according to Licciardello et al., “A combinatorial screen of the CLOUD uncovers a synergy targeting the androgen receptor,” Nat Chem Biol. 2017; 13:771 -778. The E2 concentration was normalized to the protein concentration of the cells, for which the cells were lysed by NP40 buffer and measured with a BCA kit (Thermo Fisher Scientific) (FIG. 6, panel A).High-content screen for compounds, drugs and chemicals for lipid droplet markers that can efficiently direct lipid droplet degradations by lysosomesHigh-content analysis for ZK23.5
[0146] High-content microscopy is employed to observe the number and volume of lipid droplets of Huh7 cells that was treated with or without 200 pM oleic acid (OA). For cells with O A treatment, 2863 FDA-approved drugs, 1005 novel bioactive compounds including and 1028 private compounds with novel chemical perturbation space including 4-(4-Fluorobenzyl)-8-(lH-pyrrolo[2,3-b]pyridin-5-yl)-3,4-dihydrobenzo[f][l,4]oxazepin-5(2H)-one (namely compound ZK23.5) (10 pM and 20 pM) were added to the cell culture medium 6 hours after the addition of OA, DMSO was used as a solvent control, DGAT1 / 2 inhibitor (inhibits lipid droplet formation) and BFA (inhibits lipolysis under acidic conditions) were used as negative controls and Aldometanib was used as positive controls. High-content microscopy was conducted 24 hours after the addition of the 3000 compounds or control reagents. For cells without OA treatment, 3000 compounds and control reagents were added to the cell medium followed high-content microscopy.Lipid droplet assessment in fixed cells HepG2 cells
[0147] HepG2 cells under different nutritional conditions (normal culture medium, low-glucose culture medium, MCDM (methionine and choline-deficient medium) were treated with 10 pM ZK23.5 for 24 or 36 hours, then cells are fixed in 4% PF A for 20 minutes, washedAttorney Docket No. 72194-701.602three times with PBS buffer, and stained with BODIPY 493 / 503 (a specific lipid droplet dye, diluted 1:5000 from a 1 mg / mL DMSO stock solution) for 30 minutes. Cells are washed twice with PBS and mounted, followed by rapid fluorescence imaging.Examination of lipid droplets in lysozyme
[0148] HepG2 cells are treated with 10 pM ZK23.5 compound for 4-6 hours, followed by 12 hours of treatment with Lalistat 2. The state of lipid droplets and lysosomes is observed through BODIPY staining and lysosome tracking. In detail, BODIPY 493 / 503 is diluted from a 1 mg / mL stock solution in DMSO (1:2000) in culture medium and stained for 30 minutes in a cell culture incubator. Lyso-Tracker Red (1:15000, stained for 15 minutes) is used instead of the staining medium, diluted in pre-warmed culture medium, and stained in a cell culture incubator. The staining medium is then replaced with fresh culture medium without Lyso-Tracker. Imaging of cells is performed immediately after staining, followed by calculation of the proportion of LD entering lysosomes.Assessment of autophagic lipolysis based on Lipodetector
[0149] HepG2 cells cultured in normal and low-glucose culture media are treated with 10 pM ZK23.5 for 24 hours and subjected to cell flow analysis. Lipodetector are used to assess the effect of ZK23.5 on autophagic lipolysis, with fluorescence analysis conducted at 561 nm and 405 nm. In detail, stably lipodetector-expressing cells are cultured in glass-bottom dishes and imaged using an Olympus FV3000 confocal microscope equipped with a 63x objective lens. mKeima cells are alternately excited at 405 nm and 561 nm. When excited at 405 nm, the sample's emission spectrum is in the range of 600-620 nm, and when excited at 561 nm, the emission spectrum is in the range of 620-700 nm. The emitted light is separated into respective signal collectors using an SDM400-620 dichroic mirror.Flow Cytometry and Gating Analysis
[0150] Stably lipodetector-expressing HepG2 cells are processed as per the figure legend. Cells are collected and analyzed using a Beckman CytoFlex S. The fat detector is excited with dual lasers at 405 nm and 561 nm. Emissions from both excitations are collected through a 610 / 20bp filter. Data is analyzed using CytExpert software. Emission signals excited at 405 nm are recorded as lipodetector (405), and those excited at 561 nm are recorded as lipodetector (561). Approximately 10,000 cells are classified. Single cells are gated based on lipodetector (405) and lipodetector (561). The circled triangular gate represents the population of acidified fat detectors. The same gating parameters are applied to comparative samples within the same batch of experiments.Attorney Docket No. 72194-701.602Cell Staining with BODIPY 493 / 503 and Lyso-Tracker Red
[0151] For lipid droplet staining of fixed cells, cells are fixed in 4% PFA for 20 minutes, washed three times with PBS buffer, and stained with BODIPY 493 / 503 (diluted 1:5000 from a 1 mg / mL DMSO stock solution) for 30 minutes. Cells are washed twice with PBS and mounted. For live-cell staining with LD BODIPY 493 / 503 and lysosome Lyso-Tracker, calculate the proportion of LD entering lysosomes. BODIPY 493 / 503 is diluted from a 1 mg / mL stock solution in DMSO (1:2000) in culture medium and stained for 30 minutes in a cell culture incubator. Lyso-Tracker Red (1:15000, stained for 15 minutes) is used instead of the staining medium, diluted in pre-warmed culture medium, and stained in a cell culture incubator. The staining medium is then replaced with fresh culture medium without Lyso-Tracker. Imaging of cells is performed immediately after staining.High-content screen for compounds, drugs and chemicals for lipid droplet markers that can efficiently promote actin depolarizationCell Culture
[0152] RPE1 was used for high-content screen. The culture medium is DMEM / F-12 (11320033, Gibco), supplemented with 10% Fetal Bovine Serum (AD00004, Azaood) and 1% Penicillin-Streptomycin (15140122, Gibco). The method for cell digestion and passaging is as follows: Discard the cell culture medium and wash twice with HBSS solution (14170112, Gibco) to remove excess dye. Add 400 pl of 0.125% trypsin to each well, incubate in the cell culture incubator for 1 minute, gently pipette up and down 3-4 times with a pipette to detach the adherent cells, then add 1 ml of complete culture medium, and disperse the cells by pipetting. Transfer the cell suspension to a centrifuge tube, centrifuge at 1000 rpm for 5 minutes, discard the supernatant, and resuspend the cells in 2 ml of complete culture medium, mix well. Add another 10 ml of complete culture medium and mix again. Sequentially add the cell suspension to a 96-well plate.Construction of Lifeact-GFP Stable Cell Line
[0153] Use custom-packaged viruses carrying Lifeact-GFP, mixed with cell culture medium in equal proportions, and add the transfection reagent Polybrene (abs42025397, Absin) to a final concentration of 10 pg / ml. Prepare the virus -culture medium-Polybrene mixed solution as described above and preheat at 37°C. Culture cells in a six-well plate to 50% confluence, remove the culture medium, and add 2 ml of the prepared mixed solution to each well, gently mix, and incubate the cells at 37°C for 24 hours. Gently mix every 1.5 hours during the first 6 hours of cell culture. After 24 hours, select cells using Puromycin (HY-K1057, MCE) to eliminate cells that have not been successfully infected. InAttorney Docket No. 72194-701.602the early stages of the experiment, establish a kill curve by setting different concentration gradients to determine the optimal selection concentration of 10 pg / ml. Prepare Puromycin in complete culture medium to a final concentration of 10 pg / ml and preheat at 37°C.Remove the infection mixture and add 2 ml of the prepared selection mixture to each well, incubate the cells at 37°C for 48 hours.Drug, compounds and chemicals treatment
[0154] 2863 FDA-approved drugs, 1005 novel bioactive compounds including and 1028 private compounds with novel chemical perturbation space were screened using an Al method disclosed herein. The drugs, compounds and chemicals were added to the culture medium to achieve the following concentrations: 0.1 pM, 0.3pM, IpM, 3pM and lOpM. DMSO were used as negative control.High-content screen for cell fluorescence
[0155] Imagexpress micro 4 from MOLECULAR DEVICES system was used for the analysis of the high-content cell imaging. The plate size, thickness, and other related parameters were set according to the instrument's instructions. A 20X objective lens is applied and the fluorescence of FITC channel were captured for each well. During the imaging process, use the instrument's cell culture environmental control module to maintain a temperature of approximately 37°C and a carbon dioxide concentration of approximately 5%.High-content screen for compounds preventing premature ovarian insufficiency
[0156] Through a ceramide-induced premature ovarian insufficiency (POI) cell based phenotypic screening modem (FIG.6, panel A), ramelteon, a melatonin receptor agonist approved for sleeping disorder, was identified to increase the E2 production from human ovarian granulosa cell line KGN cells (FIG. 6, panel B). In mice with dietary BCAA restriction induced POI, ramelteon treatment not only prevented the downregulation of E2 levels (FIG.7, panel A) but also protected the fertility (FIG. 7, panel B), increased the number of primordial follicles and decreased the number of atretic follicles in the ovaries (FIG. 7, panel C). Importantly, the treatment of ramelteon affected the naturally aging process of female reproductive system by increasing the serum E2 levels and improving fertility in 18-month-old mice (FIG. 6, panel C and panel D).High-content screen for compounds, drugs and chemicals for lipid droplet markers that can efficiently direct lipid droplet degradations by lysosomes
[0157] By high-content analysis, 2863 FDA-approved drugs, 1005 novel bioactive compounds and 1028 private compounds were screened, using an Al method disclosed herein,Attorney Docket No. 72194-701.602with novel chemical perturbation space for lipid droplet markers that can efficiently direct lipid droplet degradations by lysosomes. DMSO was utilized as a solvent control, DGAT1 / 2 inhibitor and BFA were used as negative controls and Aldometanib was used as positive control on lipid droplet degradation in Huh7 cells under 200pM oleic acid (OA) treatment or without OA treatment. As a result, compound 4-(4-Fluorobenzyl)-8-(lH-pyrrolo[2,3-b]pyridin-5-yl)-3,4-dihydrobenzo[f][l,4]oxazepin-5(2H)-one (namely compound ZK23.5) (lOpM and 20pM) was the most efficient one on lipid droplet degradation. In Huh7 cells under 200pM OA treatment, distinct reductions in lipid droplet volume and number were observed within individual cells treated with compound ZK23.5 (FIG. 8A). Furthermore, for Huh7 cells that were treated with compound ZK23.5 (lOpM and 20pM) under OA-free conditions, notable reductions in both lipid droplet (LD) volume and number were observed (FIG. 8B), indicated an effective degradation of compound ZK23.5 on LD. Altogether, these results indicating ZK23.5's efficacy in modulating lipid metabolism and attenuating intracellular lipid accumulation, with consistent results across conditions with and without OA, suggesting its potential in both normal and dyslipidemic states.
[0158] To validate the effect of ZK23.5, HepG2 cells were employed to be treated with lOpM ZK23.5 for 24 or 36 hours under varying nutritional conditions and subsequent BODIPY / DAPI staining, results indicate that ZK23.5 may broadly modulate cellular mechanisms, positively influencing cells across diverse metabolic contexts and significantly enhancing lipid droplet degradation, as evidenced by reduced average lipid droplet area per cell (FIG. 8C). Furthermore, the lipid droplets in lysozyme were examined: HepG2 cell was treated with lOpM ZK23.5 compound for 4-6h and then treated with Lalistat 2 for 12 h. Through BODIPY staining and lysotracker, the status of lipid droplets and lysosomes could be seen. As a result, in the ZK23.5 compound group, lipid droplets in lysozyme increased significantly, indicating that the compound can promote the autophagy of lipid droplets (FIG. 8D). Moreover, KZ23.5 was assessed for its impact on autophagic lipolysis by using Lipodetector: HepG2 cells treated with 10 pM ZK23.5 for 24 hours in both normal and low-glucose media were analyzed by flow cytometry. Fluorescence analysis at 561 nm and 405 nm revealed a significant increase in the number of lysosomal lipid droplets in ZK23.5-treated cells, consistent across both media types, indicating that ZK23.5 promotes autophagic degradation of lipid droplets (FIG. 8E). High-content screen for compounds, drugs and chemicals for lipid droplet markers that can efficiently promote actin depolarization
[0159] High-content screen was applied to screen 2863 FDA-approved drugs, 1005 novel bioactive compounds including and 1028 private compounds with novelAttorney Docket No. 72194-701.602chemical perturbation space for drugs, compounds or chemicals that can efficiently promote actin depolarization. Lifeact-GFP RPE1 cell was firstly constructed in which the filamentous actin (F-actin, green color) structures were visualized by green fluorescent protein (GFP). The cells were treated by the screened drugs, compounds or chemicals at different concentration for 24 hours, their green fluorescence images were capture by a High-content screen instrument at Oh, 8h, 16h and 24h. The number and length of F-actin fibers were dramatically decreased by the treatment of 0.3pM cytochalasin D, IpM and lOpM Y-27632. Furthermore, a combined 0.3 pM cytochalasin D and IpM Y-27632 to treat the cells resulted in more depolarization of F-actin fibers than each compound alone, indicating their synergistic effect (FIG. 9).***
[0160] While preferred embodiments of the present disclosure have been shown and described herein, it will be obvious to those skilled in the art that such embodiments are provided by way of example only. Numerous variations, changes, and substitutions will now occur to those skilled in the art without departing from the disclosure. It should be understood that various alternatives to the embodiments of the present disclosure may be employed in practicing the present disclosure. It is intended that the following claims define the scope of the present disclosure and that methods and structures within the scope of these claims and their equivalents be covered thereby.
Claims
Attorney Docket No. 72194-701.602CLAIMS1. A method for predicting drug mechanism of action based on artificial intelligence (Al), comprising the following steps:(1) a step of treatment of drugs with known MOAs and data acquisition:preparing a variety of drugs with known mechanisms of action (MOAs) or signaling pathways, and a genome-wide gene expression perturbation tool;applying the drugs with known MOAs to cells cultured in multi -well plates, and using the genome-wide gene expression perturbation tool to treat different wells to obtain initial data of the cells under drug action and information on gene expression changes;(2) a step of multi-omics detection and analysis comprising:conducting multi-omics detection on the cells treated in step 1;performing marker staining on the cells to detect the expression and localization of specific proteins or molecules within the cells through specific markers;performing morphological imaging on the cells to obtain morphological data; and integrating the multi-omics detection data, marker staining data, and cell morphological imaging data, and combining the existing drug MOA knowledge and gene signaling pathway knowledge for analysis;(3) a step of Al model training:inputting the data analyzed in step 2 into an Al system, and using the data to train the artificial intelligence model, enabling the artificial intelligence model to learn the relationships and regularities between drug characteristics and known MOAs, said drug characteristics including multi-omics characteristics and cell morphological characteristics;(4) a step of prediction of unknown drug MOAs:after the training of the Al model is completed, applying drugs with unknown MOAs to the cells, and obtaining the multi-omics data and cell morphological data of the cells after treatment with the drugs with unknown MOAs;inputting the obtained data after treatment with drugs with unknown MOAs into the trained Al model, and through the Al model, comparing and analyzing the input data characteristics with the previously learned patterns, thereby predicting the possible MOAs of the unknown drugs.Attorney Docket No. 72194-701.6022. The method according to claim 1, wherein the genome-wide gene expression perturbation tool is one or more selected from the group consisting of siRNA and CRISPR technology.
3. The method according to claim 1 or 2, wherein the multi-omics detection includes one or more selected from the group consisting of genomics, transcriptomics, proteomics, and metabolomics detection.
4. The method according to any of claims 1 to 3, wherein the morphological data includes shape, size, and / or structure of the ceils.
5. The method according to any of claims 1 to 4, wherein in step (2), the integrated data is analyzed to mine the potential patterns and regularities therein, thereby providing a data foundation for the training of Al model.
6. The method according to any one of claims 1 to 5, wherein in step (3), the drug characteristics includes multi-omics characteristics and / or cell morphological characteristics.
7. Ramelteon for use in treating an estradiol related disease.
8. The ramelteon for use according to claim 7, wherein the disease comprises premature ovarian insufficiency (POI) and diseases related to female reproductive system aging.
9. Use of ramelteon in preparation of a medicament for treating an estradiol related disease.
10. The use according to claim 9, wherein the disease comprises premature ovarian insufficiency (POI) and diseases related to female reproductive system aging.
11. A method for treating an estradiol related disease in a subject in need thereof, comprising administrating to the subject an effective amount of ramelteon.
12. The method according to claim 11, wherein the disease comprises premature ovarian insufficiency (POI) and diseases related to female reproductive system aging.
13. A method for constructing an artificial intelligence-based drug evaluation system, comprising the following steps:Attorney Docket No. 72194-701.602acquiring data from sources, for example, electronic health records (EHRs), insurance claims, and / or wearable devices;performing multimodal data embedding processing on the acquired multiple types of data to generate a unified embedded data representation;constructing a foundation model based on the embedded data representation; performing generative artificial intelligence processing based on the foundation model to obtain an artificial intelligence large language model for evaluating the effectiveness and safety of drugs;comparing the output of the artificial intelligence large language model with real clinical data to perform prediction quality control, fine-tuning the system based on the comparison results to further predict therapeutic response, survival intervals, and adverse events, and finally obtaining a trained artificial intelligence large language model.
14. The method according to claim 13, wherein the step of acquiring data further comprises the step of cleaning, standardizing, and normalizing the acquired data.
15. The method according to claim 13, wherein a deep learning algorithm is adopted for the multimodal data embedding processing.
16. The method according to claim 13, wherein the foundation model includes a neural network model for feature extraction and learning of the embedded data.
17. The method according to claim 13, wherein the obtained artificial intelligence large language model in the step of performing generative artificial intelligence processing is configured to generate evaluation reports and optimization suggestions regarding drug efficacy and protocol feasibility.
18. An artificial intelligence-based drug evaluation method, comprising: inputting clinical trial data into an artificial intelligence large language model trained according to the method for constructing an artificial intelligence-based drug evaluation system of claim 13;performing a prediction process for drug evaluation in the artificial intelligence large language model to predict therapeutic response, survival intervals, and adverse events according to the clinical trial data.Attorney Docket No. 72194-701.60219. A computer system for constructing an artificial intelligence-based drug evaluation system, comprising:a processor for executing a program, anda memory for storing the program, wherein when the program is loaded into the processor, the processor implements the operations of the method for constructing an artificial intelligence-based drug evaluation system according to claim 13.
20. A computer readable medium used for recording instructions executable by a processor, the instructions, when executed by the processor, causing the processor to perform the method for constructing an artificial intelligence-based drug evaluation system according to claim 13.
21. A computer system for artificial intelligence-based drug evaluation, comprising: a processor for executing a program, anda memory for storing the program, wherein when the program is loaded into the processor, the processor implements the operations of the artificial intelligence-based drug evaluation method according to claim 18.
22. A computer readable medium used for recording instructions executable by a processor, the instructions, when executed by the processor, causing the processor to perform the artificial intelligence-based drug evaluation method according to claim 18.