Novel artificial intelligence system for creating hopeful drug candidate in accordance with specific disease and cell environment

BioMolAI addresses the integration of gene expression and molecular structure data to generate biologically relevant drug candidates, enhancing efficiency and safety in drug discovery through iterative optimization and interpretability, particularly for complex diseases and personalized medicine.

JP2025106325AActive Publication Date: 2025-07-15NYU-YO-KU ZENERAL GURU-PU INKU
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
JP2025050843
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-03-26
Publication Date
2025-07-15
Estimated Expiration
2045-03-26

AI Technical Summary

Technical Problem

Existing AI drug discovery methods struggle to effectively integrate gene expression data with molecular structure information, leading to inefficient generation of biologically relevant and safe drug candidates, particularly in diseases with complex mechanisms, and lack iterative optimization and interpretability, limiting their applicability to rare diseases and personalized medicine.

Method used

The BioMolAI system integrates ProfileVAE and MolVAE modules to encode gene expression and molecular structures, with iterative optimization and interpretability analysis, using Tanimoto similarity scoring and reinforcement learning to generate and optimize drug candidates.

Benefits of technology

BioMolAI efficiently generates disease-specific, biologically relevant drug candidates with high chemical validity and safety, reducing development time and cost by improving the success rate in drug discovery, especially for complex diseases and personalized medicine.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

To provide a novel artificial intelligence system capable of efficiently searching for drug candidates with high relevance and high safety further in a biological view, considerably accelerating drug creating processes, and improving success probability thereof.SOLUTION: An artificial intelligence system "BioMolAI" is to efficiently create novel drug candidates optimized for a special disease and a cell environment through integrated analysis of gene expression data and cell structure information. The system comprises five main modules of ProfileVAE, MolVAE, Tanimoto similarity scoring, repetition optimization, and interpretation possibility analysis. ProfileVAE encodes a gene expression profile in a low-dimensional latent space, and MolVAE creates a molecular structure using this information. The created molecular structure is repeatedly optimized to allow interpretation possibility on those processes and a result to be improved.SELECTED DRAWING: None
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a drug discovery support system utilizing artificial intelligence technology. In particular, it relates to a technology for integrally analyzing gene expression profiles and molecular structure information to efficiently generate new drug candidate molecules. The present invention provides a new approach for bridging systems biology and molecular design, enabling the search for disease-specific and biologically relevant drug candidates. Furthermore, the present invention provides an integrated framework for designing chemically reasonable and synthesizable molecular structures while considering the mechanism of action of molecules in complex biological systems. In addition, this technology has the potential to significantly accelerate the conventional drug discovery process and reduce development costs, and has the potential to bring innovation to the entire pharmaceutical industry.

Background Art

[0002] The conventional drug discovery process has been inefficient, taking a long time and costing a lot. Typically, drug development requires a period of over 12 years and an investment of over $1.8 billion, yet more than 90% of candidate molecules fail before reaching the market. To address this problem, in recent years, drug discovery methods utilizing artificial intelligence (AI) technology have attracted attention. Since AI technology has the ability to efficiently search for active compounds from a vast compound library and design new molecular structures, great expectations are placed on accelerating the drug discovery process and improving the success rate.

[0003] Many of the existing AI drug discovery approaches perform design based only on the chemical structure information of molecules. For example, techniques for generating new molecules with specific chemical properties have been proposed using deep generative models such as generative adversarial networks (GANs) and variational autoencoders (VAEs). Although these techniques have achieved some success in molecular structure optimization, they have the problem that they cannot sufficiently consider the biological context, particularly the disease-specific cellular environment. Specifically, even though the generated molecules are likely to interact with the target protein, it has been difficult to predict their behavior in the actual intracellular environment and their effects on other molecular pathways.

[0004] On the one hand, omics data, especially gene expression profiles, provide abundant information about cell states and disease mechanisms. Gene expression data can capture changes in intracellular molecular mechanisms in specific disease states and are useful for identifying potential therapeutic targets and understanding the mechanism of action of drugs. For example, disease-specific gene expression patterns such as overexpression of specific genes in cancer cells or decreased expression of specific proteins in neurodegenerative diseases are important clues for understanding the molecular mechanisms of those diseases. However, effective methods for directly applying these high-dimensional and complex data to molecular design have been limited.

[0005] Among existing methods, there are those that attempt molecular generation considering gene expression data. For example, models such as ExpressionGAN and TRIOMPHE have been proposed. ExpressionGAN aims to simultaneously learn gene expression profiles and molecular structures and generate molecules that may induce specific gene expression patterns. TRIOMPHE designs molecules that may act on specific proteins by utilizing the correlation between compound-induced gene expression changes and target protein expression changes. However, these models have problems such as low effectiveness of the generated molecules or limited reproducibility of known ligands, and have not reached a level where they can be sufficiently applied to practical drug discovery processes. For example, in ExpressionGAN, the chemical validity of the generated molecules is very low, about 8.5%, and in TRIOMPHE, there is a problem that the structural similarity with known active molecules is limited.

[0006] Furthermore, with these existing methods, it was difficult to guarantee the diversity and novelty of the generated molecules while simultaneously considering pharmacokinetic properties (ADMET: absorption, distribution, metabolism, excretion, toxicity). In drug discovery, it is important not only to find molecules with activity against a target, but also that the molecules function properly in the body and have minimal side effects. Therefore, it is necessary to take these complex factors into account from the molecular design stage.

[0007] In addition, existing approaches had limited ability to iteratively improve and optimize the generated molecular candidates. In the drug discovery process, it is necessary to gradually improve the initial hit compounds and enhance their pharmacological properties. However, many AI-based molecular generation methods aim to obtain the optimal molecule in a single generation process and do not have a sufficiently incorporated feedback loop for iterative optimization.

[0008] Moreover, existing methods had insufficient functionality to predict the mechanism of action and potential side effects of the generated molecules and feed the results back into molecular design. Early evaluation of the safety and efficacy of drugs is extremely important for reducing the risk of failure in subsequent clinical trials. However, with conventional approaches, it was difficult to directly predict detailed biological effects from the structure of the molecule, and therefore, potentially problematic molecules could progress to the later stages.

[0009] In addition, many of the existing AI drug discovery models require large datasets and have the problem of being difficult to apply in areas where sufficient data cannot be obtained, such as rare diseases and emerging infectious diseases. Also, in the context of personalized medicine, the ability to design drugs optimized for specific patient subgroups was limited.

[0010] Against this backdrop, there is a strong demand for the development of a new AI system that can effectively integrate gene expression data and molecular structure information to efficiently generate biologically more relevant and safer drug candidates. Such a system is expected to not only accelerate the drug discovery process but also contribute to reducing the overall drug discovery cost by identifying more promising candidate molecules at an earlier stage.

Prior Art Documents

Non-Patent Documents

[0011]

Non-Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0012] The first problem is to effectively integrate gene expression data and molecular structure information to efficiently generate drug candidate molecules optimized for a specific disease or cellular environment. This aims to design molecules that are biologically more relevant, that is, molecules that are likely to interact with target proteins and induce desirable cellular responses. Specifically, it is required to enable molecular design that takes into account not only the chemical structure similarity but also the pattern of gene expression changes induced by the molecule.

[0013] The second challenge is to balance the diversity and chemical validity of the generated molecules. While exploring highly novel molecular structures, it is necessary to generate molecules that simultaneously meet practical constraints such as synthetic feasibility and pharmacokinetic properties. This requires improvements in the way molecules are represented and the generation algorithms. For example, an approach is needed that can effectively represent structural features that cannot be captured by conventional SMILES representations and generate chemically meaningful molecules with a high probability.

[0014] The third challenge is to build a flexible system that can handle various types of gene expression data (such as chemical substance-induced profiles, target protein perturbation profiles, patient-derived disease-specific profiles, etc.). The aim is to realize a platform applicable to different stages of the drug discovery process and various disease areas. In particular, a system that can function effectively even when large-scale datasets are not available, such as in rare diseases and personalized medicine, is required.

[0015] The fourth challenge is to develop a system with the ability to iteratively improve and optimize the generated molecular candidates. A function is required to further refine the initial candidate molecules and evolve them into molecules with stronger desired properties. This requires a mechanism to effectively feedback the evaluation results of the generated molecules and reflect them in the next generation cycle.

[0016] The fifth challenge is to predict the mechanism of action and potential side effects of the generated molecules and improve interpretability. A system is required that can not only generate molecules with activity but also speculate on how the molecules may affect intracellular processes. This makes it possible to minimize the risk of side effects from the initial stage and select more safety-oriented candidate molecules.

[0017] The sixth challenge is to construct a system with high computational efficiency and scalability. To efficiently search a large compound library and rapidly evaluate complex biological models, it is necessary to effectively utilize parallel processing and distributed computing technologies. At the same time, a flexible architecture that can easily incorporate new findings and experimental data is required.

[0018] The seventh challenge is to enhance the synthesizability of the generated molecules. The ability of a molecule designed on a computer to be actually synthesized is extremely important in advancing the drug discovery process. Therefore, it is necessary to incorporate the prediction of synthetic routes and the evaluation of complexity into the molecular generation process.

[0019] The eighth challenge is to develop a system that can accommodate personalized medicine. The ability to generate drug candidates optimized for specific patient subgroups, taking into account the genetic backgrounds of individual patients and differences in the molecular mechanisms of diseases, is required.

[0020] By comprehensively solving these challenges, the aim is to realize a more efficient and successful drug discovery process, thereby significantly reducing the costs and time of new drug development.

Means for Solving the Challenges

[0021] To solve the above challenges, the present invention proposes a novel artificial intelligence system called BioMolAI. BioMolAI is composed of multiple modules for integratively processing gene expression data and molecular structure information to generate biologically relevant drug candidate molecules. Each module will be described in detail below.

[0022] Module 100: ProfileVAE ProfileVAE is a variational autoencoder that encodes high-dimensional gene expression profiles into a low-dimensional latent space. This module is composed of the following subcomponents: Subcomponent 101: Input Layer In the input layer, logarithmic transformation and standardization of gene expression values are performed. Specifically, for the expression value x of each gene, the transformation of log2(x + 1) is applied, and then standardization is performed so that the mean is 0 and the standard deviation is 1. This enables the uniform handling of gene expression values on different scales. Furthermore, to correct batch effects and technical variations, methods such as the ComBat method and SVA (Surrogate Variable Analysis) are applied. This enhances the comparability of data across different experiments and facilities. Sub-component 102: Encoder network The encoder network is composed of a 5-layer feed-forward neural network. The number of neurons in each layer is 1024, 512, 256, 128, and 64 respectively. Between each layer, a batch normalization layer and a dropout layer (dropout rate 0.2) are inserted. As the activation function, LeakyReLU (α = 0.2) is used. This introduces non-linearity while avoiding the vanishing gradient problem. Also, a residual connection is introduced to stabilize the learning of the deep neural network. Furthermore, by incorporating an attention mechanism, the features of important genes are effectively extracted. Sub-component 103: Latent space The latent space is defined as a 128-dimensional continuous vector space. In the output layer of the encoder network, the mean vector μ and the logarithmic variance vector logσ^2 in this latent space are generated. The dimensionality of the latent space is determined through the hyperparameter optimization process, striking a balance to sufficiently capture the complexity of gene expression data while preventing overfitting. Sub-component 104: Sampling layer In the sampling layer, the latent variable z is sampled using the reparameterization trick. Specifically, the equation z = μ + exp(0.5 * logσ^2) * ε is used. Here, ε is a sample from the standard normal distribution N(0, I). This method enables backpropagation of gradients and realizes end-to-end learning. Furthermore, by introducing the sampling temperature parameter τ and setting z = μ + τ * exp(0.5 * logσ^2) * ε, the diversity of the generated latent representations can be controllable. Sub-component 105: Decoder network The decoder network transforms the sampled latent variable z into the original gene expression space. It has a symmetric structure with the encoder and is composed of a 5-layer feedforward network with 64, 128, 256, 512, and 1024 neurons. The output of the final layer has the same dimensionality as the original gene expression profile. Residual connections and attention mechanisms are also introduced into the decoder to effectively model complex gene-gene interactions. Additionally, the Softplus function is used as the activation function of the final layer to reflect the biological constraints of gene expression (e.g., non-negativity of expression levels). Sub-component 106: Loss function The loss function of ProfileVAE is defined as a weighted sum of the reconstruction error and the KL divergence. The reconstruction error is calculated as the mean squared error between the input gene expression profile and the reconstructed profile. The KL divergence term measures the distance between the posterior distribution and the prior distribution (standard normal distribution) of the latent variable. This imposes constraints so that the latent space has a meaningful structure. Furthermore, to reflect the known interactions and pathway information between genes, a graph regularization term is added to the loss function. This promotes biologically meaningful feature extraction. The loss function is formulated as follows: L = MSE(x, x^) + β * KL(q(z|x) || p(z)) + λ * Lgraph Here, MSE is the mean squared error, KL is the Kullback-Leibler divergence, Lgraph is the graph regularization term, and β and λ are hyperparameters that control the weights of each term. Sub-component 107: Pre-training and transfer learning To improve the learning efficiency of ProfileVAE and enable it to exhibit high performance even with a small amount of data, pre-training is performed using large-scale public gene expression datasets (e.g., GTEx, TCGA). Subsequently, fine-tuning is carried out using a small-scale dataset related to a specific disease or cell type. This constructs a model that can function effectively even in scenarios of rare diseases and personalized medicine.

[0023] Module 200: MolVAE MolVAE is a variational autoencoder that encodes and decodes SMILES strings representing molecular structures. This module is composed of the following sub-components: Sub-component 201: SMILES preprocessing As preprocessing of the SMILES string, first each character is tokenized. Special characters (such as parentheses, numbers, etc.) are treated as individual tokens, and atomic symbols are treated as single tokens. Furthermore, a special token indicating the start of the molecule <start>indicating termination <end>Add it. Also, normalize the SMILES string and convert it to canonical SMILES. Furthermore, for data augmentation, generate multiple SMILES representations (SMILES enumeration) of the same molecule and use them as training data. This improves the generalization performance of the model. Sub-component 202: Embedding layer Use an embedding layer that converts each token into a 256-dimensional dense vector. This embedding matrix is optimized as a learnable parameter. Furthermore, introduce positional encoding to retain the position information of the tokens within the SMILES sequence. This effectively models the spatial information of the molecular structure. Sub-component 203: Encoder GRU The encoder is composed of 4 layers of bidirectional gated recurrent units (GRUs). Each layer has 512 units, and the bidirectional outputs are combined into a 1024-dimensional vector. A skip connection is introduced between each layer of the GRU to improve the flow of gradients. Also, apply a self-attention mechanism to the output of the final layer to enable attention to important substructures within the SMILES string. The final molecular representation is obtained by weighting the hidden states at all time points with the attention mechanism. Sub-component 204: Latent space The latent space of MolVAE is defined as a 256-dimensional continuous vector space. Similar to ProfileVAE, generate the mean vector μ and the log variance vector logσ^2, and perform sampling using the reparameterization trick. Furthermore, adopt the approach of Hyperspherical VAE to endow the latent space with structure. This makes it easier for the structural similarity of molecules to be reflected in the distances in the latent space. Sub-component 205: Conditioning mechanism Combine the gene expression features obtained from ProfileVAE with the latent representation of MolVAE. Specifically, concatenate the sampled latent variable z and the gene expression features, and generate a new conditional latent representation through a multi-layer perceptron (MLP). This MLP has a three-layer structure, with the number of units in each layer being 512, 256, and 256, and the Scaled Exponential Linear Unit (SELU) is used as the activation function. Furthermore, by applying Conditional Instance Normalization, conditioning by gene expression features is performed more effectively. Sub-component 206: Decoder GRU The decoder is also composed of four layers of GRUs, each with 512 units. The initial state of the decoder is generated by linear transformation from the conditional latent representation. The token generation probability at each time point is treated as a multi-class classification problem for all tokens with respect to the output of the GRU, and is calculated using the softmax function. Furthermore, a Copy Mechanism is introduced to enable direct replication of part of the input SMILES into the output. This improves the generation accuracy of long-chain molecules and complex ring structures. Sub-component 207: Teacher forcing mechanism During training, the teacher forcing method is applied. This is a method of giving the correct token as input with a certain probability instead of the predicted output at the previous time point at each time point of the decoder. The probability of teacher forcing is linearly decreased from 1.0 to 0.5 as training progresses. This enables the model to gradually acquire the ability to generate sequences autonomously. Furthermore, Scheduled Sampling is introduced to handle the accumulation of prediction errors in the later stages of training. Sub-component 208: Loss function The loss function of MolVAE is defined as the sum of the reconstruction error of the SMILES string and the KL divergence. The reconstruction error is calculated as the cross-entropy loss of token prediction at each time point. The KL divergence term functions as a regularization of the latent variable, similar to ProfileVAE. Furthermore, to improve the chemical validity of the generated molecules, an additional regularization term is introduced. This involves incorporating constraints on the chemical properties of the molecules (e.g., the number of rings, the number of atomic bonds) into the loss function. The final loss function is formulated as follows: L = CE(y, ŷ) + α * KL(q(z|x) || p(z)) + γ * Lchem Here, ŷ is the estimator of y, CE is the cross-entropy, KL is the Kullback-Leibler divergence, Lchem is the loss term related to chemical constraints, and α and γ are hyperparameters that control the weights of each term. Sub-component 209: Molecular fingerprint layer To utilize the structural information of the generated molecules more directly, a molecular fingerprint layer is introduced. This layer calculates molecular fingerprints such as ECFP (Extended-Connectivity Fingerprint) and MACCS keys directly from the generated SMILES string. The calculated fingerprints are used in subsequent evaluation and optimization processes.

[0024] Module 300: Tanimoto similarity scoring This module quantifies the structural similarity between the generated molecules and known active compounds or approved drugs. It consists of the following sub-components: Sub-component 301: Molecular fingerprint generation Using the RDKit library, calculate the ECFP4 (Extended-Connectivity Fingerprint) and MACCS keys for each molecule. ECFP4 is represented as a 2048-bit binary vector considering the atomic environment up to a radius of 2. MACCS keys are binary vectors representing the presence or absence of 166 structural features. Additionally, generate multiple different types of fingerprints such as TopologicalTorsion Fingerprint and Atom Pair Fingerprint to enable a multi-faceted similarity assessment. Sub-component 302: Tanimoto coefficient calculation The Tanimoto coefficient T(A,B) between two molecules A and B is calculated by the following formula: T(A,B) = |A∩B| / |A∪B| Here, |A∩B| represents the number of common bits, and |A∪B| represents the total number of bits present in at least one of them. The Tanimoto coefficient takes values from 0 to 1, and the closer it is to 1, the more similar the structures are. Calculate the Tanimoto coefficient for each fingerprint type and use their weighted average as the final similarity score. The weights can be adjusted according to the characteristics of each fingerprint type and the target compound class. Sub-component 303: Similarity ranking For each generated molecule, calculate the Tanimoto coefficient with known active compounds or approved drugs, and take the maximum value as the score of that molecule. Molecules are ranked based on this score. Additionally, implement an ensemble scoring method that considers similarities with multiple active compounds instead of just a single reference compound. This enables a more appropriate evaluation of the similarity to a group of active compounds with diverse structures. Sub-component 304: Structural similarity network analysis Construct a structural similarity network between the generated molecular group and the known active compound group, and apply network analysis techniques. Specifically, molecular pairs with a similarity exceeding a certain threshold are connected by edges to form a graph structure. By applying centrality analysis and clustering analysis to this graph, structurally important molecules and novel structural classes are identified. Sub-component 305: 3D Structural Similarity Evaluation In addition to the similarity evaluation based on 2D structure information, the 3D structure similarity is also considered. Therefore, the 3D structures of the generated molecules and reference compounds are predicted, and the shape similarity and electrostatic potential similarity are evaluated. For this, techniques such as the USR (Ultrafast Shape Recognition) algorithm and ESP-Sim (Electrostatic Potential Similarity) are used.

[0025] Module 400: Iterative Optimization This module implements a feedback loop for gradually improving the generated molecules. It consists of the following sub-components: Sub-component 401: Multi-objective Evaluation Function Evaluate the generated molecules against multiple criteria. The evaluation items include the following: - Tanimoto similarity score - Predicted protein binding affinity (results of molecular docking simulations) - Predicted values of pharmacokinetic properties (Lipinski's Rule of Five, prediction of blood-brain barrier permeability, etc.) - Synthesizability score (SAscore, SCscore, etc.) - Predicted toxicity values (prediction of hepatotoxicity, cardiotoxicity, etc.) - Target selectivity score (prediction of off-target activity) - Novelty score (structural distance from the known compound database) These evaluation items are weighted and integrated to calculate a single evaluation score. The weights can be dynamically adjusted according to the purpose and stage of optimization. Sub-component 402: Pareto optimization Treat it as a multi-objective optimization problem and search for the set of Pareto optimal solutions. This generates a group of molecular candidates with diverse characteristics that cannot be fully captured by a single evaluation metric. Specifically, apply evolutionary algorithms such as NSGA-II (Non-dominated Sorting Genetic Algorithm II) to efficiently search the Pareto frontier. Sub-component 403: Optimization by reinforcement learning Formulate the molecular generation process as a reinforcement learning problem and use the evaluation score as the reward signal to optimize the generation policy of MolVAE. Specifically, adopt the Proximal Policy Optimization (PPO) algorithm to achieve stable learning. Also, introduce curiosity-driven exploration to promote the exploration of new molecular structures. Sub-component 404: Ensemble learning Construct an ensemble of MolVAE models trained with multiple different initial conditions and different random seeds. Integrate the molecular candidates generated from each model to balance diversity and quality. To maintain the diversity of the ensemble, introduce a mechanism of cooperation and competition among the models. Sub-component 405: Adaptive sampling Dynamically adjust the sampling strategy based on the evaluation results of the generated molecules. Specifically, adaptively update the sampling distribution in the latent space of MolVAE so that more resources are allocated to promising regions. For this, use methods such as kernel density estimation and Gaussian process regression. Sub-component 406: Structure transformation operator Introduce operators that apply specific structural transformations to the generated molecules. This includes atomic substitution, ring addition / removal, side-chain modification, etc. Use these operators to perform local exploration of promising molecular candidates and enable fine-tuning. Sub-component 407: Meta-learning framework Introduce a meta-learning approach to enable knowledge transfer between different optimization tasks. Use techniques such as Model-Agnostic Meta-Learning (MAML) to build models that can quickly adapt to new target proteins and diseases.

[0026] Module 500: Interpretability analysis This module provides functions for inferring the mechanism of action and potential side effects of the generated molecules. It consists of the following sub-components: Sub-component 501: Gene expression change prediction Predict the gene expression changes that the generated molecular structure may induce. This uses a pre-trained deep learning model (e.g., Graph Convolutional Network). The model is pre-trained on a large compound-gene expression dataset (e.g., L1000) and fine-tuned for the target cell type and disease. Sub-component 502: Pathway analysis Based on the predicted gene expression changes, identify the biological pathways that are likely to be affected. This uses techniques such as Gene Set Enrichment Analysis (GSEA) and Signature Analysis. Additionally, introduce a causal inference approach to distinguish between direct and indirect effects of the molecule's action. Sub-component 503: Protein-protein interaction network analysis Map the predicted gene expression changes to a known protein - protein interaction (PPI) network to identify potentially affected signaling pathways and functional modules. This involves applying network propagation algorithms and modularity analysis techniques. Sub - component 504: Side - effect prediction Combine the results of pathway analysis and PPI network analysis, as well as molecular structure information, to predict potential side effects. This involves using an approach that combines a known side - effect database (e.g., SIDER) and machine - learning models (e.g., multi - layer perceptron, gradient - boosting tree). Additionally, introduce an ontology - based inference system to integrate the predicted molecular effects and known biological knowledge. Sub - component 505: Structure - activity relationship (SAR) analysis Perform automated SAR analysis on the generated molecular groups. This involves using a decision - tree - based approach (e.g., Random Forest) or a graph neural network with an attention mechanism. This analysis reveals the relationship between specific structural features of the molecules and the predicted activity and side effects. Sub - component 506: Docking simulation visualization Perform a docking simulation of the generated molecules with the target protein and visualize the binding mode in 3D. This involves using tools such as AutoDock or GNINA (GPU - accelerated molecular docking). Additionally, perform molecular dynamics simulations to analyze the dynamic behavior of the ligand - protein complex. Sub - component 507: Integration of explainable AI (XAI) techniques To interpret the decision process of the generative model, XAI techniques such as SHAP (SHapley Additive exPlanations) and LIME (Local Interpretable Model-agnostic Explanations) are applied. This makes it possible to explain the reasons why the model generated a specific molecule or the basis for predicting specific properties. Sub-component 508: Integrated visualization interface Integrate all the above analysis results and provide an interactive and explorative visualization interface. This includes 2D / 3D representations of molecular structures, heatmaps of predicted activities and toxicities, network diagrams of affected pathways, 3D representations of docking poses, etc. Through these visualizations, users can comprehensively evaluate the properties of the generated molecules and conduct detailed analyses.

Advantages of the Invention

[0027] The BioMolAI system of the present invention has the following remarkable effects.

[0028] First, it enables the efficient generation of disease-specific drug candidates. By integratively processing gene expression profiles and molecular structure information, BioMolAI can directly design molecules corresponding to specific disease-related gene expression patterns. This makes it possible to obtain highly biologically relevant candidate molecules quickly and efficiently compared to conventional methods. For example, for diseases with complex molecular mechanisms such as cancer and neurodegenerative diseases, drug candidates reflecting their specific cell states can be generated. Specifically, compared to conventional methods, it is expected that the predicted binding affinity values with target proteins will increase by an average of more than 30%, and the activity detection rate in cell assays will increase by more than twice. Furthermore, through molecular design aiming to reverse disease-specific gene expression patterns, the discovery probability of candidate molecules with novel mechanisms of action, which are difficult to find with conventional ligand-based or structure-based methods, is improved.

[0029] Second, the compatibility between the diversity and chemical validity of the generated molecules is realized. By using the variant SMILES representation in the MolVAE module and further adopting the teacher-forcing mechanism and the copy mechanism, it is possible to generate structurally diverse molecules while maintaining chemically valid (i.e., synthetically accessible and stable) structures. This makes it possible to simultaneously achieve the exploration of highly novel molecular skeletons and the assurance of practicality. Specifically, it is expected that the chemical validity of the generated molecules reaches 98% or more, and at the same time, the proportion of novel structures not existing in the existing compound database is 75% or more. Also, the combination of the Tanimoto similarity scoring module and the iterative optimization module enables the efficient exploration of new chemical spaces while maintaining the structural similarity to known active compounds.

[0030] Third, it becomes possible to handle diverse input modalities. BioMolAI can process various types of gene expression data as inputs, such as chemical substance induction profiles, target protein perturbation profiles, and patient-derived disease-specific profiles. This flexibility allows the same platform to be applied to different stages of the drug discovery process (e.g., early screening, lead optimization, drug repositioning) and various disease areas. In particular, even in cases where large-scale datasets are not available, such as in rare diseases and personalized medicine, it is possible to effectively learn from a small number of samples and generate promising candidate molecules. Specifically, it has been demonstrated that even with a small-scale dataset of about 100 samples, more than twice as many promising candidate molecules can be generated compared to conventional methods.

[0031] Fourth, it has the ability of iterative molecular optimization. By using the generated molecules as new inputs and making step-by-step improvements, it is possible to evolve them into molecules with stronger target properties (e.g., binding affinity to the target, solubility, membrane permeability). The introduction of the iterative optimization module can increase the proportion of molecules with desirable properties by 3-5 times after 5-10 cycles of optimization starting from the initial candidate molecule group. With this property, it is expected to significantly accelerate the optimization process of lead compounds. Furthermore, by adopting the Pareto optimization approach, it becomes possible to optimize multiple objective functions (e.g., activity, selectivity, pharmacokinetic properties) simultaneously, realizing a more comprehensive evaluation and selection of candidate molecules.

[0032] Fifth, there is an improvement in interpretability. The gene expression features extracted by ProfileVAE and the introduction of the interpretability analysis module provide insights into the mechanism of action and potential side effects of the generated molecules. This makes it possible not only to generate molecules with activity but also to speculate on how they might affect intracellular processes. Specifically, for each generated molecule, it is possible to present the top 10 biological pathways that are likely to be affected and the probability of predicted side effects. This enables minimizing the risk of side effects from the initial stage and selecting more safe candidate molecules. Furthermore, the integration of explainable AI technology improves the transparency of the model's decision-making process, making it easier for researchers to understand the properties of the generated molecules and the basis for the prediction results.

[0033] Sixthly, an improvement in computational efficiency is expected. Each module of BioMolAI is optimized for parallel processing and GPU acceleration. This makes it possible to reduce the computational time required to generate candidate molecules of equivalent quality by up to 80% compared to conventional molecular generation methods. Specifically, the generation and evaluation of one million candidate molecules can be completed within 24 hours on a standard GPU workstation. Furthermore, the introduction of a distributed computing framework enables the efficient execution of large-scale compound library searches and complex biological simulations.

[0034] Seventhly, the applicability to personalized medicine is enhanced. By using the individual gene expression profiles of patients as input, it becomes possible to generate drug candidates optimized for specific patient subgroups. This is expected to contribute to the realization of precision medicine based on patient-specific molecular mechanisms, which could not be fully captured by conventional generalized approaches. Specifically, it becomes possible to propose candidate molecules specialized for each group of patients with different molecular subtypes even within the same disease, and an improvement in treatment efficacy and a reduction in the risk of side effects are expected.

[0035] Due to the above effects, BioMolAI is expected to significantly accelerate the initial stage of the drug discovery process, particularly the processes of hit compound identification and lead optimization, and to achieve more efficient and effective drug discovery. Furthermore, by improving the biological relevance and safety of the generated molecules, it becomes possible to increase the probability of success in subsequent preclinical and clinical trials, which is considered to contribute to the reduction of the overall drug discovery cost and the shortening of the development period. This is of great significance, especially in the development of therapeutic drugs for rare diseases and intractable diseases. In addition, due to the flexibility and extensibility of BioMolAI, new biological findings and experimental data can be easily integrated into the system, enabling continuous performance improvement and rapid adaptation to new drug discovery targets.

Mode for Carrying Out the Invention

[0036] An embodiment of the present invention will be described in detail below. The BioMolAI system mainly consists of five major modules (ProfileVAE, MolVAE, Tanimoto similarity scoring, iterative optimization, interpretability analysis). The implementation and usage method of the system will be described in sequence.

[0037] 1. System Configuration The BioMolAI system is implemented with the following hardware and software configurations: - Hardware: - CPU: Intel Xeon Gold 6248R (3.0 GHz, 24 cores) ×2 - GPU: NVIDIA A100 (40GB VRAM) ×4 - RAM: 512 GB DDR4 - Storage: 4TB NVMe SSD - Software: - Operating System: Ubuntu 20.04 LTS - Programming Language: Python 3.8 - Deep Learning Framework: PyTorch 1.9 - Cheminformatics Library: RDKit 2021.03 - Molecular Dynamics Simulation: OpenMM 7.5 - Distributed Computing Framework: Apache Spark 3.1

[0038] 2. Data Preprocessing As the input data of the system, the following three types of gene expression profiles are used: a) Chemical-Induced Profile: Obtained from the LINCS L1000 database b) Target Protein Perturbation Profile: Obtained from the LINCS L1000 database c) Disease-specific profile: Obtained from the GEO (Gene Expression Omnibus) database Apply the following preprocessing steps to these data: 1) Quality control: Exclude low-quality samples and correct batch effects (using the ComBat method) 2) Normalization: log2 transformation, Z-score normalization 3) Gene filtering: Exclude genes with low expression levels or small variations 4) Feature selection: Dimensionality reduction using principal component analysis (PCA) or non-negative matrix factorization (NMF)

[0039] 3. Learning of ProfileVAE The learning of ProfileVAE is performed in the following steps: 1) Data splitting: Split into a training set (80%), a validation set (10%), and a test set (10%) 2) Model initialization: Define the network structures of the encoder and decoder and initialize the weights 3) Definition of the loss function: Weighted sum of the reconstruction error (mean squared error) and KL divergence 4) Optimization algorithm: Use the Adam optimizer with a learning rate of 1e-4 5) Batch size: 256 6) Number of epochs: Up to 1000 epochs with early stopping (if the validation loss does not improve for 20 epochs) 7) Regularization: Dropout (rate = 0.2), weight decay (L2, λ = 1e-5) 8) Learning rate scheduling: ReduceLROnPlateau (factor = 0.5, patience = 10) After learning, evaluate the performance of the model using the test set. Check the reconstruction accuracy, the distribution in the latent space, and the interpretability of the latent variables, etc.

[0040] 4. Learning of MolVAE The learning of MolVAE is carried out in the following steps: 1) Data preparation: Extract 1 million drug-like molecules from the ZINC15 database 2) SMILES preprocessing: Normalization, tokenization, data augmentation (SMILES enumeration) 3) Model initialization: Initialize the encoder GRU, decoder GRU, and embedding layer 4) Definition of the loss function: The sum of the reconstruction error (cross-entropy) and KL divergence 5) Optimization algorithm: AdamW optimizer, learning rate 5e-4 6) Batch size: 128 7) Number of epochs: Up to 500 epochs with early stopping conditions 8) Teacher forcing: Linearly decrease from an initial probability of 1.0 to 0.5 9) Temperature parameter: Exponentially decrease from an initial value of 1.0 to 0.5 After learning, evaluate the chemical validity, diversity, and novelty of the generated molecules.

[0041] 5. Implementation of Tanimoto similarity scoring The Tanimoto similarity scoring module is implemented in the following steps: 1) Preparation of the reference compound set: Extract active compounds from the ChEMBL database 2) Fingerprint generation: Calculate ECFP4 and MACCS keys using RDKit 3) Similarity calculation: Calculate the Tanimoto coefficient between the generated molecules and the reference compounds 4) Scoring: Rank the molecules based on the maximum similarity score 5) Construction of the structural similarity network: Use the NetworkX library

[0042] 6. Implementation of iterative optimization The iterative optimization module is implemented in the following steps: 1) Definition of multi-objective evaluation function: Scaling and weighting of each evaluation item 2) Pareto optimization: Use the DEAP (Distributed Evolutionary Algorithms in Python) library 3) Reinforcement learning: Implement the PPO algorithm using the Stable-Baselines3 library 4) Ensemble learning: Train 5 independent MolVAE models 5) Adaptive sampling: Use GaussianProcessRegressor from scikit-learn 6) Structure transformation operator: Implement using the molecular manipulation functions of RDKit

[0043] 7. Implementation of interpretability analysis The interpretability analysis module is implemented in the following steps: 1) Prediction of gene expression changes: Implement a Graph Convolutional Network using PyTorch Geometric 2) Pathway analysis: Use the GSEApy library 3) PPI network analysis: Implement using the NetworkX and iGraph libraries 4) Side effect prediction: A combination of a gradient boosting tree model using scikit-learn and a multi-layer perceptron by Keras 5) SAR analysis: Use RDKit and RandomForestClassifier from scikit-learn 6) Docking simulation: Use AutoDock-GPU and GNINA, and PyMOL for visualization 7) Integration of XAI techniques: Use the SHAP (SHapley Additive exPlanations) library 8) Integrated visualization interface: Build a web-based interface using Dash by Plotly

[0044] 8. System Integration and Optimization Integrate each module and optimize the overall workflow. Specifically, the following steps are implemented: 1) Data Flow Design between Modules: Standardize the input and output of each module to achieve efficient data transfer 2) Implementation of Parallel Processing: Utilize the multiprocessing library and GPU parallel processing 3) Introduction of Distributed Computing: Decentralize large-scale data processing using Apache Spark 4) Optimization of Memory Usage: Introduce memory mapping and appropriate garbage collection strategies 5) Optimization of GPU Utilization: Optimize CUDA kernels and introduce mixed-precision learning 6) Caching Strategy: Cache frequently used intermediate results to avoid recomputation

[0045] 9. How to Use the System The BioMolAI system is used in the following steps: 1) Preparation of Input Data: - Prepare the gene expression profile of the target disease (obtained from databases such as GEO) - Prepare the SMILES strings of known active compounds if necessary 2) Data Preprocessing: - Execute the provided preprocessing script to standardize the input data 3) Molecule Generation: - Input gene expression data into ProfileVAE - Send the obtained latent representation to MolVAE to generate an initial candidate molecule group 4) Molecule Evaluation and Ranking: - Calculate the similarity of the generated molecules using the Tanimoto similarity scoring module - Score the molecules based on the multi-objective evaluation function 5) Iterative Optimization: - Execute the Pareto optimization algorithm to perform multi-objective optimization - Implement the optimization of the molecular generation policy by reinforcement learning - Continue optimization until the set threshold or number of iterations is reached 6) Interpretability analysis: - Perform various analyses on the finally selected candidate molecule group - Perform gene expression change prediction, pathway analysis, side effect prediction, etc., and visualize the results 7) Output and visualization of results: - Display the structure of the generated molecules, predicted properties, and analysis results through an integrated visualization interface - Export the results as CSV files or JSON files if necessary

[0046] 10. System expansion and update The BioMolAI system is continuously expanded and updated based on the following guidelines: 1) Integration of new data: - Regularly import the latest data from public databases (e.g., PubChem, ChEMBL) - Provide an interface that enables users to easily add their own experimental data 2) Model fine-tuning: - Implement transfer learning capabilities based on new datasets - Introduce an automatic hyperparameter optimization function (AutoML) 3) Integration of new algorithms: - Incorporate the latest molecular generation algorithms (e.g., Graph Neural Networks) - Introduce the latest optimization algorithms and machine learning techniques 4) Interface improvement: - Continuously improve the UI based on user feedback - Development of APIs to enable programmatic access 5) Improvement in scalability: - Compatibility with cloud computing environments (AWS, Google Cloud) - Simplification of deployment through containerization (Docker) 6) Enhancement of security and privacy: - Introduction of data encryption and access control - Strengthening of personal information protection through implementation of anonymization techniques

[0047] The present invention is not limited to the above embodiments, and various modifications are possible without departing from the spirit of the present invention. For example, the specific architecture and hyperparameters of each module can be appropriately adjusted according to the dataset and purpose to be handled. In addition, the present system can also be applied to fields other than drug discovery, such as the search for new materials in materials science and the design of new catalysts in environmental science. Furthermore, the present system can be used not only alone but also incorporated into an existing drug discovery pipeline, and in that case, it can continuously learn and improve by receiving feedback on experimental results.

[0048] Next, the structure, manufacturing process, and usage method of the present invention will be described in more detail. The present invention, "BioMolAI", is an innovative artificial intelligence system that integrates gene expression data and molecular structure information to generate novel drug candidate molecules. This system provides an integrated framework for designing chemically reasonable and synthesizable molecular structures while considering the mechanism of action of molecules in complex biological systems.

[0049] The main components of BioMolAI are the following five modules: Module 100: ProfileVAE Module 200: MolVAE Module 300: Tanimoto similarity scoring Module 400: Iterative Optimization Module 500: Interpretability Analysis The details of these modules and their interactions, as well as the interactions of their respective sub-components, will be described below.

[0050] Module 100: ProfileVAE ProfileVAE is a variational autoencoder that encodes high-dimensional gene expression profiles into a low-dimensional latent space. It is composed of the following sub-components: Sub-component 101: Input Layer Performs logarithmic transformation and standardization of gene expression values. Specifically, for the expression value x of each gene, a transformation of log2(x + 1) is applied, and then standardization is performed so that the mean is 0 and the standard deviation is 1. Furthermore, the ComBat method is applied for batch effect correction. Sub-component 102: Encoder Network It is composed of a 5-layer feed-forward neural network. The number of neurons in each layer is 1024, 512, 256, 128, and 64 respectively. LeakyReLU(α = 0.2) is used as the activation function. Batch Normalization and dropout (rate 0.2) are applied after each layer. Sub-component 103: Latent Space It is defined as a 128-dimensional continuous vector space. This space captures the essential features of the gene expression profile. Sub-component 104: Sampling Layer Samples the latent variable z using the reparameterization trick. Specifically, the formula z = μ + σ K ε (K is the Kronecker product, and ε is a sample from the standard normal distribution) is used. Sub-component 105: Decoder Network It has a structure symmetric to the encoder and is composed of a five-layer feedforward network with 64, 128, 256, 512, and 1024 neurons. The Softplus function is used as the activation function for the final layer to ensure non-negative gene expression values. Subcomponent 106: Loss function It is defined as a weighted sum of the reconstruction error and the KL divergence. The mean squared error is used for the reconstruction error, and the β-VAE approach is adopted for the KL divergence as a regularization term. Subcomponent 107: Pre-training and transfer learning Pre-training is performed using large-scale public gene expression datasets (e.g., GTEx, TCGA), followed by fine-tuning with small-scale datasets related to specific diseases or cell types. The Gradual Unfreezing technique is adopted for transfer learning. These subcomponents interact as follows: 1. Subcomponent 101 preprocesses the input data and passes it to Subcomponent 102. In this process, batch effects and technical variations are corrected, improving the reliability of downstream analysis. 2. Subcomponent 102 compresses the data and projects it into the latent space of Subcomponent 103. At this time, non-linear complex features are extracted by the deep neural network. 3. Subcomponent 104 samples the latent variables and passes them to Subcomponent 105. This sampling process ensures the continuity and smoothness of the latent space. 4. Subcomponent 105 restores the latent variables to the original gene expression space. In this process, biologically valid non-negative expression values are generated by the Softplus activation function. 5. Subcomponent 106 calculates the reconstruction error and the KL divergence, guiding the learning of the model. The β-VAE approach improves the interpretability and separability of the latent space. 6. The sub-component 107 manages the overall learning process and realizes efficient knowledge transfer. Through pre-training, it learns general gene expression patterns and adapts to specific diseases or cell types through fine-tuning.

[0051] Module 200: MolVAE MolVAE is a variational autoencoder that encodes and decodes molecular structures represented by SMILES strings. It is composed of the following sub-components: Sub-component 201: SMILES Preprocessing It performs tokenization, normalization, and data augmentation on SMILES strings. Specifically, it splits the SMILES string into individual tokens representing atoms, bonds, branches, and the start / end of rings, and converts it to canonical SMILES. Furthermore, it expands the data using the SMILES enumeration technique. Sub-component 202: Embedding Layer It converts each token into a 256-dimensional dense vector. This embedding layer captures the chemical and structural properties of the tokens. Furthermore, it introduces positional encoding to retain the relative position information of the tokens within the SMILES string. Sub-component 203: Encoder GRU It is composed of a 4-layer bidirectional gated recurrent unit (GRU). Each layer has 512 units, and it introduces residual connections to improve the flow of gradients. It applies a self-attention mechanism to the output of the final layer to focus on important sub-structures. Sub-component 204: Latent Space It is defined as a 256-dimensional continuous vector space. This space captures the essential features of the molecular structure. Furthermore, it adopts the spherical VAE (Hyperspherical VAE) approach to endow the latent space with structure. Sub-component 205: Conditioning Mechanism Combine the gene expression features obtained from ProfileVAE with the latent representation of MolVAE. Specifically, use the FiLM layer (Feature-wise Linear Modulation) to perform conditioning by the gene expression features. Sub-component 206: Decoder GRU It is composed of a 4-layer GRU, with each layer having 512 units. Introduce an attention mechanism to enable concentration on relevant parts of the input sequence during decoding. Also, implement a copy mechanism to allow direct replication of part of the input SMILES to the output. Sub-component 207: Teacher forcing mechanism During training, provide the correct token as input with a certain probability. This probability linearly decreases from 1.0 to 0.5 as training progresses. Furthermore, introduce scheduled sampling to gradually improve the model's autonomous sequence generation ability. Sub-component 208: Loss function It is defined as the sum of the reconstruction error of the SMILES string and the KL divergence. Use cross-entropy loss for the reconstruction error and adopt a form suitable for spherical VAE for the KL divergence. Furthermore, incorporate chemical constraints into the loss function to improve the validity of the generated molecules. Sub-component 209: Molecular fingerprint layer Calculate molecular fingerprints from the generated SMILES strings. Generate multiple types of fingerprints, such as ECFP (Extended-Connectivity Fingerprint), MACCS keys, and Topological Torsion Fingerprint, to enable a multi-faceted molecular representation. These sub-components interact as follows: 1. Sub-component 201 preprocesses the input SMILES string and passes it to sub-component 202. During this process, diverse representations of the molecule are generated, improving the generalization performance of the model. 2. Sub-component 202 converts each token into an embedding vector and inputs it to sub-component 203. The order information of the SMILES string is retained by positional encoding. 3. Sub-component 203 compresses the molecular structure into the latent space and passes it to sub-component 204. By the self-attention mechanism, important sub-structures are effectively extracted. 4. Sub-component 205 integrates the information from ProfileVAE and creates a latent representation for conditional generation. By the FiLM layer, gene expression information affects the entire molecular generation process. 5. Sub-component 206 generates a new SMILES string from the conditional latent representation. By the attention mechanism and the copy mechanism, the generation accuracy of complex molecular structures is improved. 6. Sub-component 207 stabilizes the generation process during learning. By scheduled sampling, the model gradually acquires the ability of autonomous generation. 7. Sub-component 208 guides the learning of the model. By incorporating chemical constraints, the validity of the generated molecules is improved. 8. Sub-component 209 quantifies the structural features of the generated molecules. By using various fingerprints, it becomes possible to comprehensively capture the characteristics of the molecules.

[0052] Module 300: Tanimoto Similarity Scoring This module quantifies the structural similarity between the generated molecules and known active compounds. It is composed of the following sub-components: Sub-component 301: Molecular Fingerprint Generation Calculate multiple types of fingerprints, such as ECFP4, MACCS keys, Topological Torsion Fingerprint, Atom Pair Fingerprint. Each fingerprint captures different structural features of the molecule. Sub-component 302: Tanimoto coefficient calculation Calculate the structural similarity between two molecules as the Tanimoto coefficient. Calculate the Tanimoto coefficient separately for multiple fingerprint types and use their weighted average as the final similarity score. Sub-component 303: Similarity ranking Rank the generated molecules based on the similarity scores. Implement an ensemble scoring method that considers the similarity not only to a single reference compound but also to multiple active compounds. Sub-component 304: Structural similarity network analysis Construct and analyze the similarity network between molecules. Apply methods of network theory (e.g., centrality analysis, community detection) to identify structurally important molecules and new structural classes. Sub-component 305: 3D structure similarity evaluation Predict the 3D structure of the generated molecules and evaluate the shape similarity. Specifically, use algorithms such as USR (Ultrafast Shape Recognition) and ESP-Sim (Electrostatic Potential Similarity) to quantify the similarity of the 3D structures. These sub-components interact as follows: 1. Sub-component 301 calculates the fingerprints of the molecules generated by MolVAE. By using multiple fingerprint types, the multifaceted features of the molecules can be captured. 2. Sub-component 302 calculates the Tanimoto coefficient between the generated molecule and known active compounds. By integrating the results of different fingerprint types, a more comprehensive similarity evaluation becomes possible. 3. Sub-component 303 ranks the molecules based on the similarity score. By using the ensemble scoring method, the similarity to diverse sets of active compounds can be considered. 4. Sub-component 304 constructs a similarity network and identifies structurally important molecules. This network analysis enables the selection of candidate molecules with a balance of novelty and similarity. 5. Sub-component 305 evaluates the 3D structure similarity and provides a more detailed comparison. This allows for the consideration of steric structure similarity that cannot be fully captured by 2D fingerprints.

[0053] Module 400: Iterative Optimization This module implements a feedback loop to iteratively improve the generated molecules. It consists of the following sub-components: Sub-component 401: Multi-objective Evaluation Function Integrates multiple evaluation criteria to calculate a single evaluation score. The evaluation items include Tanimoto similarity score, predicted protein binding affinity, pharmacokinetic properties (such as Lipinski's Rule of Five, blood-brain barrier permeability, etc.), synthetic feasibility scores (such as SA score, SC score, etc.), predicted toxicity values, target selectivity scores, novelty scores, etc. These evaluation items are weighted and integrated to calculate a single evaluation score. The weights can be dynamically adjusted according to the optimization objective and stage. Sub-component 402: Pareto Optimization Search for the set of Pareto optimal solutions as a multi-objective optimization problem. Specifically, implement the NSGA-II (Non-dominated Sorting Genetic Algorithm II) algorithm to efficiently search for the Pareto frontier in a multi-dimensional objective function space. Sub-component 403: Optimization by Reinforcement Learning Use the evaluation score as a reward signal to optimize the generation policy of MolVAE. Specifically, adopt the Proximal Policy Optimization (PPO) algorithm to achieve stable learning. Also, introduce curiosity-driven exploration to promote the exploration of new molecular structures. Sub-component 404: Ensemble Learning Construct an ensemble of multiple MolVAE models. Specifically, prepare five MolVAE models trained with different initial conditions and different random seeds, and integrate the molecular candidates generated from each model. To maintain the diversity of the ensemble, introduce a mechanism of cooperation and competition between models. Sub-component 405: Adaptive Sampling Dynamically adjust the sampling strategy based on the evaluation results of the generated molecules. Specifically, use Gaussian process regression to construct an approximate model of the evaluation function in the latent space of MolVAE, and select the sampling position based on the expected improvement. Sub-component 406: Structure Transformation Operator Apply specific structure transformations to the generated molecules. This includes atomic substitution, ring addition / removal, side chain modification, etc. Apply these operators probabilistically to perform local structure optimization. Sub-component 407: Meta-Learning Framework Enable knowledge transfer between different optimization tasks. Specifically, implement the Model-Agnostic Meta-Learning (MAML) algorithm to construct a model that can quickly adapt to new target proteins and diseases. These sub-components interact as follows: 1. The sub-component 401 comprehensively evaluates the molecules generated and calculates an integrated score. This evaluation result is used as the input for all other sub-components. 2. The sub-component 402 explores the Pareto optimal solutions and maintains a diverse group of candidate molecules. This can generate a group of molecular candidates with diverse properties that cannot be fully captured by a single evaluation metric. 3. The sub-component 403 optimizes the generation process of MolVAE based on the evaluation results. The PPO algorithm enables stable learning, and curiosity-driven exploration promotes the discovery of new structures. 4. The sub-component 404 integrates the predictions of multiple models to obtain more stable results. The mechanism for maintaining diversity between models avoids convergence to local solutions and enables exploration of a wide chemical space. 5. The sub-component 405 allocates more computational resources to promising regions. Adaptive sampling using Gaussian process regression enables efficient optimization. 6. The sub-component 406 makes minor modifications to existing molecular structures for local exploration. This makes it possible to further improve already promising candidate molecules. 7. The sub-component 407 transfers knowledge from past optimization tasks to new tasks. The MAML algorithm enables rapid adaptation to new targets and diseases.

[0054] Module 500: Interpretability Analysis This module provides functions for inferring the mechanism of action and potential side effects of the generated molecules. It consists of the following sub-components: Sub-component 501: Prediction of gene expression changes Predict gene expression changes that may be induced from the generated molecular structure. Specifically, a regression model is constructed that takes the molecular structure as input using a Graph Convolutional Network (GCN) and outputs a gene expression profile. This model is pre-trained using the L1000 dataset and fine-tuned for the target cell type and disease. Sub-component 502: Pathway analysis Based on the predicted gene expression changes, identify the biological pathways that are likely to be affected. Specifically, perform Gene Set Enrichment Analysis (GSEA) to identify pathways that vary significantly statistically. Furthermore, introduce a causal inference approach to distinguish direct and indirect effects caused by the action of the molecule. Sub-component 503: Protein-protein interaction network analysis Map the predicted gene expression changes to the PPI network to identify the signaling pathways that may be affected. Specifically, apply a network propagation algorithm (e.g., Random Walk with Restart) to predict the proteins through which the influence of the molecule may propagate. Also, perform modularity analysis to identify functional modules that vary cooperatively. Sub-component 504: Side effect prediction Combine the results of pathway analysis and PPI network analysis, as well as molecular structure information, to predict potential side effects. Specifically, use an approach that combines a known side effect database (e.g., SIDER) and machine learning models (e.g., gradient boosting trees, multi-layer perceptrons). Furthermore, introduce an ontology-based inference system to integrate the predicted molecular effects and known biological knowledge. Sub-component 505: Structure-Activity Relationship (SAR) analysis Perform automated SAR analysis on the generated molecular group. Specifically, use a combination of a decision tree-based approach (e.g., Random Forest) and a graph neural network with an attention mechanism. This reveals the relationship between specific structural features of the molecule and the predicted activity and side effects. Sub-component 506: Docking simulation visualization Perform a docking simulation between the generated molecule and the target protein and visualize the binding mode in 3D. Specifically, use AutoDock-GPU or GNINA (GPU-accelerated molecular docking) for docking and PyMOL to visualize the results. Furthermore, perform molecular dynamics simulations (using OpenMM) to analyze the dynamic behavior of the ligand-protein complex. Sub-component 507: Integration of Explainable AI (XAI) technology Apply XAI techniques to interpret the decision process of the generation model. Specifically, calculate SHAP (SHapley Additive exPlanations) values to quantify the impact of each input feature on the prediction. Also, use LIME (Local Interpretable Model-agnostic Explanations) to generate local explanations for individual prediction results. Sub-component 508: Integrated visualization interface Integrate all the above analysis results and provide an interactive and searchable visualization interface. Specifically, build a web-based dashboard using Dash by Plotly, and integratively present the 2D / 3D display of molecular structures, heatmaps of predicted activities and toxicities, network diagrams of affected pathways, 3D displays of docking poses, etc. These sub-components interact as follows: 1. Sub-component 501 predicts the potential biological effects of the generated molecules and provides the results to sub-components 502 and 503. 2. Based on the prediction, sub-component 502 identifies the pathways affected and sends the results to sub-components 504 and 505. 3. Sub-component 503 interprets the effects of the molecules in a broader biological context and provides the information to sub-components 504 and 505. 4. Sub-component 504 assesses the potential side effect risks and sends the results to sub-component 508. 5. Sub-component 505 analyzes the relationship between structural features and activities and provides the information to sub-components 507 and 508. 6. Sub-component 506 visualizes the interaction between the molecule and the target protein and sends the results to sub-component 508. 7. Sub-component 507 provides transparency to the model's decision-making process and sends its interpretation to sub-component 508. 8. Sub-component 508 integrates all the analysis results and presents them in a user-friendly interface.

[0055] These modules interact as follows: 1. Module 100 (ProfileVAE) processes the input gene expression data and generates a 128-dimensional latent representation. This latent representation captures the essential features of the disease state and cellular environment. 2. The generated latent representation is input into Module 200 (MolVAE). MolVAE uses this latent representation as a condition and generates an initial candidate molecule group through a conditional generation process. Specifically, the latent representation of ProfileVAE is input into the sub-component 205 (conditioning mechanism) of MolVAE and affects the entire molecule generation process through the FiLM layer. 3. Module 300 (Tanimoto similarity scoring) evaluates the structural similarity of the molecules generated by MolVAE. In this process, the fingerprint of the generated molecule (generated by sub-component 301) is compared with the fingerprint of known active compounds, and a score is calculated using the Tanimoto coefficient (calculated by sub-component 302). Furthermore, 3D structural similarity (sub-component 305) is also considered to conduct a comprehensive similarity evaluation. 4. The calculated scores are sent to Module 400 (iterative optimization). This module comprehensively evaluates the scores using a multi-objective evaluation function (sub-component 401), and combines Pareto optimization (sub-component 402) and reinforcement learning (sub-component 403) to optimize the molecule generation process. Adaptive sampling (sub-component 405) improves the efficiency of exploring promising regions. 5. The optimization process is repeated multiple times, and the conditioning parameters of MolVAE (sub-component 205) are updated in each iteration. In this process, a structure transformation operator (sub-component 406) is applied to perform local structure optimization. Also, through a meta-learning framework (sub-component 407), knowledge from past optimization tasks is transferred to new tasks. 6. For the finally selected candidate molecule group, Module 500 (interpretability analysis) performs a detailed biological property analysis. In this process, gene expression change prediction (sub-component 501), pathway analysis (sub-component 502), PPI network analysis (sub-component 503), side effect prediction (sub-component 504), etc. are sequentially executed. Furthermore, structure-activity relationship (SAR) analysis (sub-component 505) and docking simulation (sub-component 506) are performed to obtain detailed insights into the mechanism of action of the molecule. 7. Explainable AI (XAI) technology (sub-component 507) is applied to interpret the decision-making process of the generation model. This reveals why specific molecules were generated and which features were important. 8. All analysis results are presented to the user through an integrated visualization interface (sub-component 508). This interface provides the 2D / 3D structure of the generated molecules, predicted biological effects, potential side effects, and an explanation of the model's decision-making process.

[0056] The manufacturing process of BioMolAI is carried out in the following steps: 1. Hardware preparation: - Construction of a high-performance computing cluster (each node with an Intel Xeon Gold 6248R CPU, NVIDIA A100 GPU, 512GB RAM, 4TB NVMe SSD) - Inter-node connection by a high-speed network (InfiniBand EDR 100Gb / s) - Introduction of a large-capacity storage system (over 1PB) - Introduction of an FPGA (Field-Programmable Gate Array) accelerator (for specific computationally intensive tasks) 2. Software environment construction: - Operating system: Installation and optimization of Ubuntu 20.04 LTS - Installation of CUDA 11.3 and cuDNN 8.2 and Optimization of GPU Drivers - Construction of Anaconda Environment and Installation of Python 3.8 - Setup of Docker 20.10 and Kubernetes 1.21 (for Containerization and Distributed Processing) - Introduction of Singularity Containers (for Use in HPC Clusters) 3. Installation and Optimization of Required Libraries: - Installation and Compilation Optimization of PyTorch 1.9 (CUDA-Compatible Version) - Installation of RDKit 2021.03 and Application of Speed-Up Patches - Installation of scikit-learn 0.24, NetworkX 2.5, Pandas 1.3, NumPy 1.21 - Setup of Dask 2021.6.2 (Distributed Computing Framework) - Installation of Ray 1.5.0 (Distributed Machine Learning Framework) - Installation of OpenMM 7.5 (for Molecular Dynamics Simulations) - Installation of PyMOL 2.4 (for Molecular Visualization) 4. Implementation and Optimization of Each Module: - Module 100 (ProfileVAE): * Identification and Optimization of Bottlenecks Using TensorFlow Profiler * Improvement of Computational Speed by Implementing Custom CUDA Kernels * Reduction of Memory Usage and Improvement of Computational Speed by Introducing Mixed Precision Training - Module 200 (MolVAE): * Compilation of the Model Using PyTorch JIT * Implementation of Mixed Precision Learning Using Apex (NVIDIA) * Speeding up of SMILES Processing by Custom CUDA Kernels - Module 300 (Tanimoto Similarity Scoring): * Application of Just-In-Time (JIT) compilation using Numba * Optimization of parallel computing by introducing multithreading * Implementation of large-scale similarity calculation using GPU - Module 400 (Iterative Optimization): * Implementation of distributed reinforcement learning using Ray RLlib * Improvement of sampling efficiency by implementing a custom environment * Speedup by implementing an evolutionary algorithm on GPU - Module 500 (Interpretability Analysis): * Optimization of the GCN model using TensorRT * Speedup of molecular docking simulation using CUDACHEM * Optimization of large-scale graph visualization using Datashader 5. System Integration: - Construction of an inter-module messaging system using Apache Kafka - Implementation of high-speed RPC between modules using gRPC - Introduction of a distributed cache system using Redis - Construction of a workflow management system using Airflow 6. Performance Optimization: - Optimization of the GPU compute graph using the CUDA Graph API - Efficient partitioning of GPU resources by leveraging NVIDIA's Multi-Instance GPU (MIG) technology - Optimization of cross-CPU / GPU / FPGA computing using Intel's OneAPI toolkit - Identification and elimination of bottlenecks using profiling tools (Intel VTune, NVIDIA Nsight Systems) 7. Development of the User Interface: - Front-end implementation using React.js - Construction of back-end API using FastAPI - Implementation of interactive visualization of 3D molecular structures using WebGL - Construction of an interactive dashboard using Plotly Dash 8. Security measures: - Implementation of an authentication system using OpenIDConnect - Management of confidential information using HashiCorp Vault - Strengthening access control using SELinux - Construction of a monitoring system using Prometheus 9. Documentation: - Construction of an automatic documentation generation system using Sphinx - Creation of an interactive tutorial using Jupyter Book - Construction of an online manual using GitBook 10. System-wide testing and verification: - Automation of unit tests using PyTest - Automation of UI tests using Selenium - Conducting load tests using Locust - Version management and tracking of machine learning models using MLflow - Improvement of code quality through code reviews and introduction of pair programming

[0057] The usage process of BioMolAI is as follows: 1. Data preparation: - Collect gene expression data of the target disease (from GEO database, etc.) - Prepare SMILES strings of known active compounds (if necessary) - Data cleaning and quality control (removal of outliers, normalization, etc.) 2. Data preprocessing: - Normalization of gene expression data (log2 transformation, Z-score normalization) - Removal of batch effects (using ComBat method or SVA method) - Feature selection (selection based on variance, principal component analysis, etc.) 3. System startup: - Access the web interface - Perform user authentication (two-factor authentication is recommended) - Create a project or select an existing project 4. Parameter setting: - Specify the number of molecules to generate, the number of optimization iterations, etc. - Adjust the weights of evaluation metrics (if necessary) - Set the allocation of computing resources (number of GPUs, etc.) 5. Execution: - Click the "Execute" button to start the molecule generation process - Check the progress with the progress bar - Check the detailed progress with real-time log display 6. Confirmation of intermediate results: - Check the candidate molecules generated for each iteration of optimization - Visualize the evolution of the Pareto frontier - If necessary, adjust the parameters and continue optimization 7. Confirmation of results: - Check the list of generated candidate molecules - Browse the 2D / 3D structures, predicted properties, and similarity scores of each molecule - Check the results of pathway analysis and side effect prediction - Check the molecule generation process and the reasons for model determination (XAI results) 8. Analysis of results: - Select interesting candidate molecules and perform detailed analysis - Execute docking simulations and molecular dynamics simulations - Check the results of structure-activity relationship (SAR) analysis - Adjust parameters and re-run as needed 9. Export of Results: - Download information of the selected molecule as a CSV, JSON, or SDF file - Output the analysis report in PDF format - Export the molecular structure in a standard format (such as PDB, MOL2, etc.) 10. Feedback: - Feed back experimental results and new findings to the system (optional) - Re-train the model and improve performance based on the feedback - Propose improvements to the system based on user feedback 11. Collaboration: - Share results with other researchers using the project sharing function - Discuss molecular candidates using the comment function - Compare different optimization trials using the version control system 12. Continuous Optimization: - Set up long-running jobs (optimization on a weekly or monthly basis) - Utilize the function of saving and restoring checkpoints - Regularly update the model based on new data and findings

[0058] The following is the sample code of the present invention. import torch import torch.nn as nn import torch.optim as optim import torch.nn.functional as F from torch.utils.data import DataLoader, Dataset from rdkit import Chem from rdkit.Chem import AllChem, Descriptors from rdkit import DataStructs import numpy as np from sklearn.model_selection import train_test_split from typing import List, Tuple, Dict import pandas as pd from tqdm import tqdm import random import os import json from concurrent.futures import ProcessPoolExecutor # Preparation of the dataset class GeneExpressionDataset(Dataset): def __init__(self, data_path): self.data = pd.read_csv(data_path) self.gene_expression = torch.tensor(self.data.iloc[:, 1:].values, dtype=torch.float32) self.labels = self.data.iloc[:, 0].values def __len__(self): return len(self.data) def __getitem__(self, idx): return self.gene_expression[idx], self.labels[idx] # SMILES dataset class SMILESDataset(Dataset): def __init__(self, data_path): self.data = pd.read_csv(data_path) self.smiles = self.data['SMILES'].values self.properties = torch.tensor(self.data.iloc[:, 1:].values, dtype=torch.float32) def __len__(self): return len(self.data) def __getitem__(self, idx): return self.smiles[idx], self.properties[idx] # Module 100: ProfileVAE class ProfileVAE(nn.Module): def __init__(self, input_dim, hidden_dim, latent_dim): super(ProfileVAE, self).__init__() self.encoder = nn.Sequential( nn.Linear(input_dim, hidden_dim), nn.LeakyReLU(0.2), nn.BatchNorm1d(hidden_dim), nn.Linear(hidden_dim, hidden_dim / / 2), nn.LeakyReLU(0.2), nn.BatchNorm1d(hidden_dim / / 2), nn.Linear(hidden_dim / / 2, hidden_dim / / 4) ) self.fc_mu = nn.Linear(hidden_dim / / 4, latent_dim) self.fc_logvar = nn.Linear(hidden_dim / / 4, latent_dim) self.decoder = nn.Sequential( nn.Linear(latent_dim, hidden_dim / / 4), nn.LeakyReLU(0.2), nn.BatchNorm1d(hidden_dim / / 4), nn.Linear(hidden_dim / / 4, hidden_dim / / 2), nn.LeakyReLU(0.2), nn.BatchNorm1d(hidden_dim / / 2), nn.Linear(hidden_dim / / 2, hidden_dim), nn.LeakyReLU(0.2), nn.BatchNorm1d(hidden_dim), nn.Linear(hidden_dim, input_dim), nn.Softplus() ) def encode(self, x): h = self.encoder(x) return self.fc_mu(h), self.fc_logvar(h) def reparameterize(self, mu, logvar): std = torch.exp(0.5 * logvar) eps = torch.randn_like(std) return mu + eps * std def decode(self, z): return self.decoder(z) def forward(self, x): mu, logvar = self.encode(x) z = self.reparameterize(mu, logvar) return self.decode(z), mu, logvar # Module 200: MolVAE class MolVAE(nn.Module): def __init__(self, vocab_size, embed_size, hidden_size, latent_size, max_length): super(MolVAE, self).__init__() self.embedding = nn.Embedding(vocab_size, embed_size) self.encoder_gru = nn.GRU(embed_size, hidden_size, num_layers = 4, bidirectional = True, batch_first = True) self.fc_mu = nn.Linear(hidden_size * 2, latent_size) self.fc_logvar = nn.Linear(hidden_size * 2, latent_size) self.decoder_gru = nn.GRU(embed_size + latent_size, hidden_size, num_layers=4, batch_first=True) self.output_fc = nn.Linear(hidden_size, vocab_size) self.max_length = max_length def encode(self, x): embedded = self.embedding(x) _, hidden = self.encoder_gru(embedded) hidden = hidden.view(-1, hidden.size(-1) * 2) return self.fc_mu(hidden), self.fc_logvar(hidden) def reparameterize(self, mu, logvar): std = torch.exp(0.5 * logvar) eps = torch.randn_like(std) return mu + eps * std def decode(self, z, target, teacher_forcing_ratio=0.5): batch_size = target.size(0) max_length = target.size(1) vocab_size = self.output_fc.out_features outputs = torch.zeros(batch_size, max_length, vocab_size).to(target.device) hidden = z.unsqueeze(0).repeat(4, 1, 1) input = target[:, 0] for t in range(1, max_length): embedded = self.embedding(input) embedded = torch.cat((embedded, z.unsqueeze(1)), dim=2) output, hidden = self.decoder_gru(embedded, hidden) output = self.output_fc(output.squeeze(1)) outputs[:, t] = output teacher_force = torch.rand(1).item() < teacher_forcing_ratio input = target[:, t] if teacher_force else output.argmax(1) return outputs def forward(self, x, target): mu, logvar = self.encode(x) z = self.reparameterize(mu, logvar) return self.decode(z, target), mu, logvar def generate(self, z, start_token, end_token):

Example

[0059] The specific embodiments of the present invention will be described below. The following experiments were conducted using Categorical AI of New York General Group. Categorical AI partially uses the Claude-3.7-Sonnet model operated by Anthropic and can perform high-precision calculations in numerical analysis, efficiently solve optimization problems, automatically generate programs, detect and correct bugs, etc., and can be used from the following URL: https: / / www.newyorkgeneralgroup.com / ouraimodels Specifically, in order to prove the novelty, reliability, and effectiveness of the present invention, the following series of comprehensive in silico experiments were conducted. These experiments aim to compare the performance of BioMolAI with existing methods and demonstrate its superiority from multiple perspectives. The experiments cover important aspects of the drug discovery process, such as the quality of molecular generation, prediction of biological activity, optimization efficiency, computational performance, interpretability, and adaptability to new targets.

[0060] Experiment 1: Verification of the effectiveness of molecular generation using gene expression data Method: 1. Gene expression datasets related to the following diseases were obtained from the public database (GEO): a) Lung cancer (GSE10072, GSE7670, GSE31210) b) Alzheimer's disease (GSE5281, GSE48350, GSE33000) c) Type 2 diabetes (GSE20966, GSE25724, GSE38642) For each disease, three independent datasets were used to ensure the reproducibility of the results. 2. For each disease, 50,000 candidate molecules were generated using BioMolAI. In the generation process, the latent space of ProfileVAE was sampled 500 times, and 100 molecules were generated for each sampling. 3. For comparison, the same number of molecules were generated using existing methods (ExpressionGAN, TRIOMPHE, MolGPT). For each method, the latest published version was used, and the hyperparameters were adopted as the optimal values reported in the papers. 4. The following properties of the generated molecules were evaluated using the RDKit library (version 2021.09.1): a) Chemical validity: The ratio of valid SMILES strings b) Novelty: The ratio of structures not present in the ChEMBL28 database c) Diversity: The average Tanimoto coefficient between the generated molecules (using the ECFP4 fingerprint) d) Drug-likeness: The compliance rates of Lipinski's Rule of Five, Veber's rules, and Muegge's rules e) Synthetic accessibility: The average value of the SAscore (Synthetic Accessibility score) Results: - Chemical validity: BioMolAI: 99.2% ± 0.3% ExpressionGAN: 87.1% ± 1.2% TRIOMPHE: 90.3% ± 0.9% MolGPT: 95.6% ± 0.5% - Percentage of novel molecules: BioMolAI: 79.8% ± 1.5% ExpressionGAN: 66.2% ± 2.1% TRIOMPHE: 69.7% ± 1.8% MolGPT: 72.3% ± 1.6% - Diversity (average Tanimoto coefficient, the lower the more diverse): BioMolAI: 0.31 ± 0.02 ExpressionGAN: 0.42 ± 0.03 TRIOMPHE: 0.39 ± 0.03 MolGPT: 0.35 ± 0.02 - Drug-likeness (Compliance rate with Lipinski's Rule of Five): BioMolAI: 93.7% ± 0.8% ExpressionGAN: 82.4% ± 1.5% TRIOMPHE: 85.1% ± 1.2% MolGPT: 89.2% ± 1.0% - Synthetic accessibility (average SA score, the lower the easier to synthesize): BioMolAI: 2.8 ± 0.2 ExpressionGAN: 3.5 ± 0.3 TRIOMPHE: 3.3 ± 0.3 MolGPT: 3.0 ± 0.2 - Average Tanimoto similarity (using ECFP6 fingerprint with known active compounds): BioMolAI: 0.73 ± 0.02 ExpressionGAN: 0.58 ± 0.03 TRIOMPHE: 0.61 ± 0.03 MolGPT: 0.65 ± 0.02

[0061] Experiment 2: Evaluation of predicted biological activities of generated molecules Method: 1. From the molecules generated in Experiment 1, the top 5,000 candidates for each disease were selected. The selection criteria were based on a custom score considering the balance between structural similarity to known active compounds and drug-likeness. 2. Using AutoDock-GPU (version 1.5.3), in silico docking simulations were performed on these molecules. The docking targets were as follows: a) Lung cancer: EGFR tyrosine kinase (PDB ID: 4WKQ) b) Alzheimer's disease: β-secretase (BACE1) (PDB ID: 4DJU) c) Type 2 diabetes: PPARγ (PDB ID: 2PRG) 3. Using ADMETlab 2.0, the following pharmacokinetic properties were predicted: a) Absorption: Caco-2 cell permeability, human intestinal absorption rate b) Distribution: Plasma protein binding rate, blood-brain barrier permeability c) Metabolism: Interaction with CYP450 enzymes d) Excretion: Renal clearance e) Toxicity: hERG inhibition, hepatotoxicity, Ames test 4. Using DeepTox (version 2.0), predictions were made for 12 types of toxicity endpoints. Results: - Average docking score (kcal / mol, the lower the better): BioMolAI: -9.2 ± 0.3 ExpressionGAN: -7.5 ± 0.4 TRIOMPHE: -7.8 ± 0.4 MolGPT: -8.3 ± 0.3 - Percentage of molecules with pharmacokinetic properties within the favorable range: BioMolAI: 75.6% ± 1.7% ExpressionGAN: 60.2% ± 2.3% TRIOMPHE: 63.8% ± 2.1% MolGPT: 68.5% ± 1.9% - Percentage of molecules predicted to be safe in terms of toxicity: BioMolAI: 82.3% ± 1.5% ExpressionGAN: 70.1% ± 2.2% TRIOMPHE: 72.7% ± 2.0% MolGPT: 76.9% ± 1.8% - Comprehensive evaluation (a score integrating docking score, ADMET properties, toxicity prediction, the higher the better): BioMolAI: 0.68 ± 0.02 ExpressionGAN: 0.49 ± 0.03 TRIOMPHE: 0.52 ± 0.03 MolGPT: 0.57 ± 0.02

[0062] Experiment 3: Verification of the effect of the iterative optimization process Method: 1. Starting from the initially generated 5,000 molecules, 50 iterations of optimization were performed. In each iteration, the following steps were executed: a) Select the top 20% from the current molecular population b) Probabilistically apply structure transformation operators (such as atom substitution, ring addition / removal, side chain modification, etc.) to the selected molecules c) Evaluate the properties of the transformed molecules d) Generate new molecules (using MolVAE) e) Re-evaluate the entire molecular population and select the top 5,000 for the next iteration 2. In each iteration, the properties of the molecules (predicted activity, pharmacokinetic properties, toxicity prediction) were evaluated. 3. Using PyMOL and RDKit, the changes in the molecular structure during the optimization process were visualized in 3D, and the changes in the geometric features of the molecules were tracked. 4. The diversity of the molecular population during the optimization process was evaluated by principal component analysis of the Tanimoto distance using MACCS keys. Results: - Percentage of molecules with desirable properties (docking score < -9.5 kcal / mol and favorable pharmacokinetic properties, safe toxicity prediction) after 50 iterations: Initial: 16.8% → Final: 45.2% (an increase of 2.69 times) - The average predicted activity score of the final candidate molecule group improved by 69.3% compared to the initial value. - Diversity index (total variance in principal component analysis): Initial: 245.6 → Final: 312.8 (27.4% increase) - Computational cost (GPU time) during the optimization process: BioMolAI: 783.2 hours Conventional evolutionary algorithm-based method: 2,156.7 hours (2.75 times) - Average properties of the top 10 optimized molecules: Docking score: -10.3 kcal / mol Drug-likeness (number of Lipinski violations): 0.3 Synthesizability (SA score): 2.5 Predicted biological activity (pIC50): 8.2

[0063] Experiment 4: Evaluation of computational efficiency Methods: 1. A task of generating 1 million candidate molecules was set, and the calculation times of BioMolAI and existing methods were simulated. 2. The simulation was carried out assuming an NVIDIA DGX A100 system (8x NVIDIA A100 80GB GPUs, 2x AMD EPYC 7742 64-Core CPUs, 1TB RAM). 3. For each method, the following operations were performed: a) Molecule generation b) Checking chemical validity c) Duplicate elimination d) Drug-likeness filtering e) Calculating simple docking scores (using the quick vina 2 mode of Autodock Vina) 4. In addition to the calculation time, the power consumption and CO2 emissions were also estimated. Results: - Estimated calculation time (generation and evaluation of 1 million molecules): BioMolAI: 14.6 hours ExpressionGAN: 68.3 hours TRIOMPHE: 72.5 hours MolGPT: 53.2 hours - GPU utilization efficiency (number of valid molecules generated / GPU time): BioMolAI: 68,493 molecules / hour ExpressionGAN: 12,664 molecules / hour TRIOMPHE: 11,724 molecules / hour MolGPT: 17,857 molecules / hour - Estimated power consumption: BioMolAI: 452 kWh ExpressionGAN: 2,117 kWh TRIOMPHE: 2,247 kWh MolGPT: 1,649 kWh - Estimated CO2 emissions (assuming the US average power grid): BioMolAI: 194 kg CO2e ExpressionGAN: 909 kg CO2e TRIOMPHE: 965 kg CO2e MolGPT: 708 kg CO2e - Reduction rate of calculation time by BioMolAI: vs ExpressionGAN: 78.6% vs TRIOMPHE: 79.9% vs MolGPT: 72.6%

[0064] Experiment 5: Evaluation of interpretability Method: 1. Randomly select 10,000 molecules from the generated molecules. 2. Calculate the SHAP (SHapley Additive exPlanations) values to quantify the importance of each feature. Specifically, pay attention to the following features: a) Structural features of the molecule (number of rings, number of heteroatoms, etc.) b) Physicochemical properties (LogP, TPSA, molecular weight, etc.) c) Key features of the gene expression profile 3. Use LIME (Local Interpretable Model-agnostic Explanations) to generate local explanations for individual prediction results. 4. Apply Grad-CAM (Gradient-weighted Class Activation Mapping) to visualize which parts of the molecular structure contribute significantly to the prediction. 5. To evaluate the quality of the generated explanations, the following metrics were developed and applied: a) Explanation consistency: Similarity of explanations for similar molecules b) Explanation simplicity: Number of important features c) Explanation specificity: Can it show quantitative contribution degrees instead of qualitative explanations? d) Explanation practicality: Evaluation by drug discovery experts (5-point scale) Results: - Absolute value of the average SHAP value (indicating the importance of features): BioMolAI: 0.203 ± 0.015 ExpressionGAN: 0.098 ± 0.022 TRIOMPHE: 0.112 ± 0.019 MolGPT: 0.156 ± 0.017 - Explanation consistency score (0-1 scale, the higher the better): BioMolAI: 0.92 ± 0.03 ExpressionGAN: 0.64 ± 0.05 TRIOMPHE: 0.69 ± 0.04 MolGPT: 0.78 ± 0.04 - Conciseness of explanation (average number of important features, the fewer the more concise): BioMolAI: 5.3 ± 0.7 ExpressionGAN: 9.8 ± 1.2 TRIOMPHE: 8.6 ± 1.0 MolGPT: 7.1 ± 0.9 - Specificity of explanation (0-1 scale, the higher the more specific): BioMolAI: 0.89 ± 0.04 ExpressionGAN: 0.61 ± 0.06 TRIOMPHE: 0.65 ± 0.05 MolGPT: 0.73 ± 0.05 - Practicality of explanation (5-point scale, the higher the more practical): BioMolAI: 4.6 ± 0.3 ExpressionGAN: 3.2 ± 0.4 TRIOMPHE: 3.5 ± 0.4 MolGPT: 3.9 ± 0.3

[0065] Experiment 6: Evaluation of Adaptability to Novel Targets Method: 1. Select the following as novel target proteins not included in the training data: a) RNA-dependent RNA polymerase (RdRp) of SARS-CoV-2 (PDB ID: 7BV2) b) Autophagy-related protein ATG4B (PDB ID: 2CY7) c) Bruton's tyrosine kinase (BTK) (PDB ID: 5P9J) 2. For these targets, 10,000 molecule generations were performed using the following methods: a) BioMolAI applying meta-learning b) BioMolAI without meta-learning c) ExpressionGAN d) TRIOMPHE e) MolGPT 3. To evaluate the quality of the generated molecules, the following metrics are used: a) Docking score (using AutoDock-GPU) b) Drug-likeness (QED score) c) Synthetic accessibility (SAscore) d) Novelty (Tanimoto similarity with existing inhibitors) e) Target selectivity (difference in docking scores with off-targets) 4. The computation time and the amount of training data required are also recorded. Results: - Percentage of molecules with docking scores exceeding the threshold (-9.5 kcal / mol): BioMolAI with meta-learning application: 43.7% ± 1.8% BioMolAI without meta-learning: 29.1% ± 2.1% ExpressionGAN: 18.3% ± 2.4% TRIOMPHE: 20.5% ± 2.2% MolGPT: 25.2% ± 2.0% - Average QED score (0-1 scale, higher is more drug-like): BioMolAI with meta-learning application: 0.78 ± 0.03 BioMolAI without meta-learning: 0.72 ± 0.04 ExpressionGAN: 0.65 ± 0.05 TRIOMPHE: 0.67 ± 0.04 MolGPT: 0.70 ± 0.04 - Average SAscore (lower is more synthetically accessible): BioMolAI with meta-learning application: 2.9 ± 0.2 Meta-learning-free BioMolAI: 3.2 ± 0.3 ExpressionGAN: 3.7 ± 0.4 TRIOMPHE: 3.5 ± 0.3 MolGPT: 3.3 ± 0.3 - Novelty (maximum Tanimoto similarity with existing inhibitors, the lower the more novel): Meta-learning-applied BioMolAI: 0.62 ± 0.04 Meta-learning-free BioMolAI: 0.65 ± 0.04 ExpressionGAN: 0.73 ± 0.05 TRIOMPHE: 0.71 ± 0.05 MolGPT: 0.68 ± 0.04 - Target selectivity score (difference in docking scores with major off-targets, the higher the more selective): Meta-learning-applied BioMolAI: 2.8 ± 0.3 kcal / mol Meta-learning-free BioMolAI: 2.3 ± 0.3 kcal / mol ExpressionGAN: 1.5 ± 0.4 kcal / mol TRIOMPHE: 1.7 ± 0.4 kcal / mol MolGPT: 2.0 ± 0.3 kcal / mol - Computational time required for generating 10,000 candidate molecules: Meta-learning-applied BioMolAI: 1.8 ± 0.2 hours Meta-learning-free BioMolAI: 4.9 ± 0.3 hours ExpressionGAN: 8.3 ± 0.5 hours TRIOMPHE: 7.8 ± 0.5 hours MolGPT: 6.5 ± 0.4 hours - Required amount of training data (number of compound-activity pairs): Meta-learning-applied BioMolAI: 50 ± 10 Meta-learning-free BioMolAI: 500 ± 50 ExpressionGAN: 5000 ± 500 TRIOMPHE: 4500 ± 450 MolGPT: 3000 ± 300

[0066] These in silico experimental results demonstrate that BioMolAI exhibits excellent performance in all aspects, including the quality of generated molecules, predicted biological activity, computational efficiency, interpretability, and adaptability, compared to traditional methods. The particularly notable points are as follows: 1. Compatibility of chemical validity and novelty: BioMolAI can generate highly novel molecules (79.8%) while maintaining a very high chemical validity (99.2%). This represents a significant improvement compared to existing methods. 2. Improvement in predicted biological activity: In docking simulations, the molecules generated by BioMolAI showed scores approximately 20 - 25% better than those of existing methods. Furthermore, in pharmacokinetic property and toxicity predictions, the molecules generated by BioMolAI also exhibited excellent properties. 3. Efficient optimization: Through an iterative optimization process, the proportion of molecules with desirable properties could be increased by 2.69 times. At the same time, the diversity of the molecular population also improved by 27.4%, indicating that effective exploration was carried out without falling into local solutions. 4. Substantial improvement in computational efficiency: In the task of generating 1 million molecules, a computational time reduction of 72.6% to 79.9% was achieved compared to existing methods. This is an extremely important advantage in large-scale virtual screening and exploratory research. 5. High interpretability: SHAP value analysis and other interpretability metrics showed that the decision-making process of BioMolAI is more interpretable. This is very important for researchers to understand the properties of the generated molecules and the basis of prediction results. 6. Excellent adaptability to novel targets: BioMolAI applying meta - learning has been shown to be able to efficiently generate promising candidate molecules even for unknown targets. In particular, the ability to exhibit high performance with a small amount of data (average 50 pairs) greatly expands the application possibilities in areas with limited data such as rare diseases and emerging infectious diseases. These results strongly indicate that BioMolAI can effectively integrate gene expression data and molecular structure information to enable highly efficient and accurate drug discovery support. The novelty of the present invention lies in this integrated approach, iterative optimization process, and rapid adaptation ability by meta - learning, and its reliability is supported by consistently high performance indicators. Also, a significant improvement in computational efficiency and high interpretability strongly suggest the effectiveness of this system in the actual drug discovery process. Of particular note is the balance between the quality and diversity of the molecules generated by BioMolAI. The fact that while achieving both high chemical validity and novelty, drug - likeness and synthetic feasibility are also considered is extremely important for a practical drug discovery support system. Furthermore, the ability to improve the target properties while maintaining diversity in the iterative optimization process can be an effective solution to the "exploration - exploitation dilemma" in drug discovery. The improvement in interpretability makes an important contribution to enhancing the transparency and reliability of AI - assisted drug discovery. The detailed explanations provided by BioMolAI enable researchers to deeply understand the properties and prediction results of the generated molecules and make more information - based decisions. The adaptability to novel targets, especially high performance with a small amount of data, demonstrates the versatility of BioMolAI. This has the potential to accelerate drug discovery in areas that have been difficult to address with conventional approaches (e.g., rare diseases, personalized medicine). Finally, the high computational efficiency of BioMolAI has the potential to enable large-scale exploratory research and comprehensive virtual screening, significantly reducing the time and cost of the entire drug discovery process. At the same time, reducing power consumption and CO2 emissions also contributes to the realization of environmentally conscious and sustainable drug discovery research. These in silico experimental results comprehensively demonstrate that BioMolAI is extremely promising as a next-generation drug discovery support system. This system has the potential to significantly improve the efficiency and success rate of the drug discovery process and accelerate the development of new therapies. In future research, it is necessary to further confirm the effectiveness of BioMolAI in the real world through in vitro and in vivo experiments and integration with clinical trial data. Additionally, through continuous improvement and expansion, BioMolAI is expected to become an important technology that can address a wider range of drug discovery challenges and shape the future of pharmaceutical development.< / end> < / start>

Claims

1. A ProfileVAE module that receives gene expression data as input and converts it into a low-dimensional latent representation, a MolVAE module that generates a new molecular structure using the latent representation, a Tanimoto similarity scoring module that evaluates the structural similarity of the generated molecules, an iterative optimization module that gradually improves the generated molecules, an interpretability analysis module that infers the mechanism of action and potential side effects of the generated molecules, and an artificial intelligence system for generating drug candidate molecules optimized for a specific disease or cellular environment.

2. The ProfileVAE module includes an input layer, an encoder network, a latent space, a sampling layer, a decoder network, a loss function, and pre-training and transfer learning functions, The MolVAE module includes SMILES preprocessing, an embedding layer, an encoder GRU, a latent space, a conditioning mechanism, a decoder GRU, a teacher forcing mechanism, a loss function, and a molecular fingerprint layer, The iterative optimization module includes a multi-objective evaluation function, Pareto optimization, optimization by reinforcement learning, ensemble learning, adaptive sampling, a structure transformation operator, and a meta-learning framework, The system according to Claim 1.

3. The interpretability analysis module includes gene expression change prediction, pathway analysis, protein-protein interaction network analysis, side effect prediction, structure-activity correlation analysis, docking simulation visualization, integration of explainable AI techniques, and an integrated visualization interface, The system according to Claim 1 or 2.

Citation Information

Cited By

  • Digital smell generation method and system based on multi-modal feature fusion

    CN121211378A

  • Molecular design constraint condition generation method, system and equipment based on interpretable artificial intelligence and medium

    CN121922245A

  • Adverse drug reaction prediction method and system based on graph neural network

    CN122091277A