Intelligent drug research and development system
By integrating a multi-unit collaborative intelligent drug development system, the problem of independent model operation in the drug development process has been solved, realizing the intelligentization and integration of the drug design process, improving R&D efficiency and accuracy, shortening the cycle and reducing costs.
Patent Information
- Application Number
- CN202511340364.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-19
- Publication Date
- 2025-12-23
AI Technical Summary
In the current drug development process, each model operates independently without integrated collaboration, resulting in low efficiency, poor accuracy, and high computational resource requirements, making it difficult to form a closed-loop system covering molecule generation, screening, optimization, evaluation, and transformation.
An intelligent drug development system is constructed by integrating a 3D conformation generation unit, a target specificity identification unit, a multi-task drug screening unit, a drug property prediction unit, a personalized drug design unit, and a dynamic feedback learning unit to form a multi-unit collaborative closed-loop system. By utilizing technologies such as deep learning and graph neural networks, the intelligent and integrated drug design process can be realized.
It significantly improves the efficiency and accuracy of drug development, shortens the cycle, reduces costs, enhances innovation and the degree of process integration support, and achieves efficient closed-loop operation from molecular generation to transformation.
Smart Images

Figure CN121191633A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of artificial intelligence and computer data processing technology, and in particular relates to an intelligent drug development system. Background Technology
[0002] With the rapid development of artificial intelligence (AI) technology, its application in the biomedical field has gradually become a key force driving the transformation of drug development paradigms. Relying on methods such as big data mining, deep learning, graph modeling, and predictive reasoning, AI has demonstrated significant acceleration and optimization effects in multiple stages, including drug target identification, candidate molecule screening, conformation prediction, and efficacy analysis. This has greatly improved R&D efficiency and innovation capabilities, making it an important supplementary technology in the new drug discovery process.
[0003] Currently, drug development still relies on traditional processes such as chemical synthesis, experimental screening, and clinical testing. Although computer-aided drug design (CADD) technology has made some progress, it still faces problems such as low efficiency, poor accuracy, the need for large amounts of computing resources, and reliance on a single model in many stages.
[0004] In existing technologies, each model operates independently. For example, the molecular property prediction model is used alone for property analysis, and the molecular generation-related model is not deeply linked with other links. It lacks the integrated collaboration like the patented platform and cannot form a closed-loop system covering "molecular generation-screening-optimization-evaluation-transformation".
[0005] To overcome the aforementioned problems in existing technologies, this invention proposes an artificial intelligence drug development platform – YLModels. By integrating multiple self-developed models and algorithm frameworks and employing cutting-edge artificial intelligence methods, it constructs an intelligent drug design closed-loop system covering "molecule generation-screening-optimization-evaluation-transformation," reshaping the functional boundaries and integration capabilities of traditional CADD tools, and improving the efficiency, accuracy, and integrated support of drug development. Summary of the Invention
[0006] The purpose of this invention is to provide an intelligent drug development system. This intelligent drug development system effectively integrates the intelligent models required at each stage of the drug design process through the mutual collaboration of multiple model units, and constructs a complete, closed-loop artificial intelligence-driven drug development platform.
[0007] To achieve the above-mentioned objectives, the present invention provides an intelligent drug development system, comprising:
[0008] Three-dimensional conformation generation unit: By combining dihedral subgraph structure representation and diffusion probability modeling, complex molecular three-dimensional conformations are modeled while preserving local molecular conformation information and overall topological structure.
[0009] Target-specific identification unit: used for intelligent dynamic analysis of drug-target interactions;
[0010] Multi-task drug screening unit: used to generate multiple molecular conformations and select the conformation that best meets the requirements of drug development through automated optimization algorithms; and to perform drug screening based on dynamic data to predict the changing trend of drug-target binding mode under different conditions.
[0011] Drug property prediction unit: at least using deep neural networks and graph neural networks, to predict multiple key properties of drug molecules screened by the multi-task drug screening unit;
[0012] Personalized drug design unit: It can reverse generate drug molecules from target properties or target structure to complete the design of personalized drugs;
[0013] Dynamic feedback learning unit: provides quantitative basis for the output of the unit; and provides synthetic route planning and feasibility assessment results.
[0014] As a further improvement of the present invention, the three-dimensional image generation unit includes a conversion module, a noise processing module, a model training module, and a model sampling module.
[0015] As a further improvement of the present invention, the conversion module is used to analyze and convert the input molecular graph structure to generate a set of dihedral subgraphs for diffusion modeling.
[0016] As a further improvement of the present invention, the dihedral subgraph set includes multiple dihedral subgraphs; each dihedral subgraph inherits the corresponding atomic feature information, bond feature information, and dihedral angle feature.
[0017] As a further improvement of the present invention, the conversion module is also used to calculate and obtain the atomic feature vector, bond feature vector and angle feature vector corresponding to each dihedral subgraph in the dihedral subgraph set.
[0018] As a further improvement of the present invention, the noise processing module is used to gradually add Gaussian noise to the three-dimensional spatial coordinates of atoms in each dihedral subgraph generated in the conversion module, so that the atomic coordinates of each atom approximate the standard Gaussian distribution and obtain a highly perturbed molecular noise conformation.
[0019] As a further improvement of the invention, the noise processing module is also used to perform inverse denoising to reconstruct the original molecular conformation from the highly perturbed molecular noise conformation through the inverse denoising.
[0020] As a further improvement of the present invention, the model training module is used to predict the direction and magnitude of the perturbation of each atom in the dihedral subgraph at the current time step based on the noise score of the atom in the current time step; the noise scores of the atoms in each dihedral subgraph are aggregated to form the atomic coordinate perturbation estimate of the entire molecule, and compared with the actual amount of noise added, the loss function is calculated, and the model parameters are optimized through backpropagation to complete the training of the model.
[0021] As a further improvement of the present invention, the model sampling module can initialize three-dimensional atomic coordinates based on a standard Gaussian distribution, and combine the transformation module with the neural network trained by the model training module to gradually restore the true conformation of the molecule through a time-reverse diffusion process.
[0022] The beneficial effects of this invention are:
[0023] The intelligent drug development system of this invention integrates a three-dimensional conformation generation unit, a target specificity identification unit, a multi-task drug screening unit, a drug property prediction unit, a personalized drug design unit, and a dynamic feedback learning unit. Through a multi-unit collaborative mechanism, it integrates intelligent models of each stage of the drug design process to form a complete, closed-loop artificial intelligence-driven drug development platform. Attached Figure Description
[0024] Figure 1 This is a structural block diagram of the intelligent drug development system of the present invention. Detailed Implementation
[0025] Embodiments of the present invention are described in detail below, examples of which are illustrated in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain the present invention, and should not be construed as limiting the present invention.
[0026] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be described in detail below with reference to the accompanying drawings and specific embodiments.
[0027] Please see Figure 1 As shown, this invention provides an intelligent drug development system 100. Through a multi-unit collaborative mechanism, it effectively integrates the intelligent models required at each stage of the drug design process, constructing a complete, closed-loop AI-driven drug development platform. The intelligent drug development system 100 includes:
[0028] 3D conformation generation unit 1: By combining dihedral subgraph structure representation and diffusion probability modeling, complex molecular 3D conformations are modeled while preserving local molecular conformation information and overall topological structure.
[0029] Target-specific recognition unit 2: used for intelligent dynamic analysis of drug-target interactions;
[0030] Multi-task drug screening unit 3: It is used to generate multiple molecular conformations and select the conformation that best meets the requirements of drug development through an automated optimization algorithm; it performs drug screening based on dynamic data to predict the changing trend of drug-target binding mode under different conditions.
[0031] Drug property prediction unit 4: at least using deep neural networks and graph neural networks, to predict multiple key properties of drug molecules screened by the multi-task drug screening unit;
[0032] Personalized drug design unit 5: It can reverse generate drug molecules from target attributes or target structure to complete the design of personalized drugs;
[0033] Dynamic Feedback Learning Unit 6: Provides quantitative basis for the output of the unit; and provides synthetic route planning and feasibility assessment results.
[0034] In this application, the 3D conformation generation unit 1 is the core unit of the intelligent drug discovery system 100, used for modeling complex molecular 3D conformations. In fact, traditional drug design often involves a lengthy process from molecular conformation generation and virtual screening to structural optimization, with each step requiring extensive computational simulations for verification, resulting in a long cycle. Especially in the molecular conformation generation and optimization process, traditional methods typically rely on complex simulation software, which is cumbersome, time-consuming, and lacks intelligent optimization strategies. In this invention, the 3D conformation generation unit 1 utilizes deep learning and machine learning to accelerate the molecular conformation generation and optimization process. The model can quickly generate multiple molecular conformations and, through automated optimization algorithms, select the conformation that best meets the requirements of drug discovery, reducing manual operations and computational costs in traditional methods and significantly shortening the drug design cycle.
[0035] Furthermore, traditional CADD techniques, such as molecular docking and virtual screening, typically focus on building static models and cannot effectively simulate the dynamic interaction processes between drugs and targets. Many drug candidate molecules may only bind effectively to targets under certain conditions, and traditional techniques cannot capture these complex dynamic processes, thus affecting the accuracy of screening results.
[0036] In addition, most existing CADD technologies are "evaluation-type", that is, they only support the prediction or screening of the activity of existing compounds and lack "generative" capabilities. They cannot reverse generate candidate molecules from the target, thus relying on human assumptions and experience design in the new drug structure design stage, which is inefficient and lacks innovation.
[0037] In this invention, the three-dimensional conformation generation unit 1 is based on the diffusion model architecture and uses a multi-step noise addition and denoising mechanism to generate high-quality molecular three-dimensional conformations from random noise. Unlike traditional conformation generation models that mainly use global graphs or atomic-level modeling, the three-dimensional conformation generation unit 1 uses dihedral angles as the core structural units, achieving a better balance between local structural expression and global generation, thus improving the rationality, diversity and efficiency of conformation generation.
[0038] The 3D model generation unit 1 includes a conversion module 11, a noise processing module 12, a model training module 13, and a model sampling module 14.
[0039] The conversion module 11 is the core component of the three-dimensional conformation generation unit 1. It is used to analyze and convert the input molecular graph structure to generate a dihedral subgraph set for diffusion modeling.
[0040] The conversion module 11 takes a standard molecular diagram as input (such as an adjacency matrix or a graph structure converted from SMILES) and identifies all possible segments that may constitute dihedral structural units by traversing the quaternary atomic sequences ABCD in the molecule that satisfy specific bond connection conditions.
[0041] Furthermore, the dihedral subgraph set includes multiple dihedral subgraphs; each dihedral subgraph inherits corresponding atomic feature information, bond feature information, and dihedral angle features. The atomic feature information includes at least the atom type; the bond feature information includes at least the type of bond between atoms (single bond, double bond, etc.) and the connection relationship; furthermore, the dihedral angle features are the dihedral angle features calculated based on the atomic feature information and bond feature information for the dihedral subgraph ABCD.
[0042] The conversion module 11 is also used to calculate and obtain the atomic feature vector, bond feature vector, and angle feature vector corresponding to each dihedral subgraph in the dihedral subgraph set. Furthermore, each dihedral subgraph output by the conversion module 11 can be used as a basic unit for independent modeling and input into the neural network for application in the model training module 13 and the model sampling module 14.
[0043] In fact, by localizing the expression of molecular structure, the transformation module 11 effectively enhances the model's ability to perceive key conformational changes, which helps to improve the overall generative model's ability to model the diversity of molecular conformations and physical rationality, as well as its generalization performance.
[0044] The noise processing module 12 is used to gradually add Gaussian noise to the three-dimensional spatial coordinates of atoms in each dihedral sub-graph generated in the conversion module 11, so that the atomic coordinates of each atom approximate the standard Gaussian distribution and obtain a highly perturbed molecular noise conformation.
[0045] Specifically, in this process, the noise processing module 12 follows a fixed variance sequence β = (β1…β2). T The Markov chain progressively adds Gaussian noise to the three-dimensional spatial coordinates of atoms in the dihedral subgraph:
[0046] The initial input is a set of real dihedral subgraphs (i.e., real three-dimensional atomic coordinates);
[0047]
[0048] The three-dimensional coordinates of the noise-injected atom gradually increase with the time step. The set of noise dihedral subgraphs at each time step t is obtained by adding noise to the set of dihedral subgraphs at the previous time step t-1.
[0049]
[0050] After T-step diffusion and noise addition, all atomic coordinates eventually approximate a standard Gaussian distribution, forming a highly perturbed molecular noise conformation.
[0051] In this process, the generated data serves as the input-prediction target pairs in the training samples for model construction and provides the initial state for subsequent noise processing.
[0052] The noise processing module 12 is also used to perform inverse denoising to reconstruct the original molecular conformation from the highly perturbed molecular noise conformation through the inverse denoising.
[0053] Specifically, the reverse denoising stage is the process of gradually "recovering" the original molecular conformation from the noise:
[0054] The noise processing module 12 uses the noise atom coordinates obtained by sampling with a standard Gaussian distribution as the initial input;
[0055]
[0056] At each step, the set of dihedral subgraphs at the current time step is input into the neural network to predict the "denoising score" ∈ [the value of the subgraph]. θ ;
[0057] Based on the prediction results, noise components are removed from the current atom coordinates t, generating a sub-graph set for the next time step t-1; where α t =1-β t ,
[0058]
[0059] After a complete denoising path from T to 0, the reasonable three-dimensional structure of the molecule is finally recovered. The model achieves effective conformational reconstruction by learning the noise score function at each step.
[0060] The model training module 13 is used to predict the direction and magnitude of the perturbation of each atom in the dihedral subgraph at the current time step based on the noise score of the atom in each dihedral subgraph. The noise scores of the atoms in each dihedral subgraph are aggregated to form the atomic coordinate perturbation estimate of the entire molecule, and compared with the actual amount of noise added. The loss function is calculated, and the model parameters are optimized through backpropagation to complete the training of the model.
[0061] In this application, the model training module 13 takes the molecular diagram as input and first maps it into a set of dihedral subgraphs through the transformation module 11. Each dihedral subgraph consists of four atoms constituting a specific dihedral and their local structural information, and is input as an independent modeling unit into the subsequent neural network. Subsequently, the noise processing module 12 randomly samples a time step t from the time interval of the diffusion process and adds Gaussian noise to the three-dimensional coordinates of the atoms in each subgraph according to the noise variance corresponding to the step size, generating a set of noise subgraphs at step t. For each dihedral subgraph, the noise processing module 12 extracts its atomic features, bond features, and dihedral angles as input. After processing by a neural network, it outputs the noise score for each atom at the current time step, used to predict the direction and magnitude of the perturbation it experiences. The noise scores of all atoms in the dihedral subgraphs are aggregated to form the estimated atomic coordinate perturbation value of the entire molecule, i.e., the noise score. This score is compared with the actual amount of noise added, and the loss function is calculated. The model parameters of the model training module 13 are then optimized through backpropagation. This training process enables the model training module 13 to gradually master the ability to recover the true conformation from noise of varying degrees, establishing a mapping relationship from random perturbations to a reasonable structure, thus laying the foundation for subsequent generation of the 3D molecular conformation.
[0062] The model sampling module 14 can initialize three-dimensional atomic coordinates based on a standard Gaussian distribution, and combine the neural network trained by the conversion module 11 and the model training module 13 to gradually restore the true conformation of the molecule through a time-reverse diffusion process.
[0063] During the sampling phase, the model sampling module 14 can start from the atomic three-dimensional coordinates generated by the standard Gaussian distribution, i.e. the molecular pure noise conformation, and gradually remove noise to recover the initial three-dimensional conformation of the molecule.
[0064] The input to this process is structural information representing the types of atoms, bond connections, and bond types, such as a two-dimensional molecular diagram or SMILES expressions. The conversion module 11 first randomly generates a set of three-dimensional atomic coordinates conforming to a standard Gaussian distribution based on the input structure, constructing an initial molecular noise conformation diagram. Subsequently, the noise processing module 12 maps this noise conformation to a corresponding set of dihedral noise subgraphs, and each subgraph is used as an independent modeling unit, progressively inputting into the trained neural network during the back-diffusion process from time step T to 0. At each time step t, the neural network predicts the noise score of the current atomic coordinates based on the subgraph structure, used to estimate the coordinate perturbation at the current step size. Based on this prediction result, the corresponding noise is removed from the current coordinates, thus obtaining the set of dihedral subgraphs for the next time step t-1. The above process iterates continuously until the denoising of all time steps is completed, and finally a three-dimensional conformation that is similar to the original molecular structure is generated. This setting enables the model sampling module 14 to effectively realize the generation process from random standard Gaussian noise conformation to reasonable three-dimensional molecular conformation, which is suitable for tasks such as diverse conformation sampling, virtual screening and drug conformation modeling.
[0065] The following description will provide a detailed explanation of the specific application of the three-dimensional modeling unit 1 through specific embodiments.
[0066] In this embodiment, taking the generation of a three-dimensional molecular conformation of caffeine as an example, the specific implementation process of the three-dimensional conformation generation unit 1 is as follows:
[0067] Input the caffeine SMILE formula or other molecular representations and convert it into a two-dimensional molecular diagram.
[0068] The DihedralsDiff model randomly generates standard Gaussian three-dimensional coordinates (N,3) based on the number of atoms N.
[0069] Molecular diagram Entering the DihedralsEncode module, it is mapped to a set of dihedral subgraphs.
[0070] Dihedral diagram set The noise score of the atom at time step T is obtained by inputting into the trained neural network. θ .
[0071] After removing the noise from the atomic three-dimensional coordinates at time step T, a set of dihedral subgraphs is obtained.
[0072] This process is used to gradually denoise the three-dimensional coordinates of atoms, ultimately yielding a conformation that approximates the initial molecular structure.
[0073] In summary, the conversion module 11 in the three-dimensional conformation generation unit 1 realizes the structural encoding of local conformations; that is, the conversion module 11 realizes the automatic identification and encoding of the dihedral subgraph set from the molecular diagram, extracts high-dimensional features such as atom type, bond type, connection structure and dihedral angle, realizes the optimized expression of local structure and global structure to neural network input, so that the conversion module 11 has structural generalization, adapts to a variety of molecular structures, and is convenient for modeling in high-dimensional chemical space.
[0074] Furthermore, the 3D conformation generation unit 1 introduces a "dihedral subgraph" as the basic input unit in the molecular conformation generation process for the first time. This transforms the traditional atomic diagram structure into a set of multiple dihedral subgraphs at four-atom units, capturing the core torsional degrees of freedom in molecular conformations and enhancing the model's ability to express local conformational changes. Compared to traditional methods based on full graphs or distance matrices, this approach significantly improves the rationality and diversity of conformation generation.
[0075] Meanwhile, the noise processing module 12 constructs a diffusion-denoising framework based on dihedral subgraph changes. By adding noise at the atomic coordinate level and combining it with angular features to predict noise scores, it guides the model to efficiently converge to a reasonable conformation in the structural perturbation space. This mechanism constructs forward and reverse diffusion paths using a Markov chain with fixed variance, and uses angular information at each time step to enhance the geometric recovery capability, ensuring the physical and chemical feasibility of the three-dimensional conformation.
[0076] During training, the model training module 13 independently inputs each subgraph into the neural network to predict the noise score of atoms in the subgraph, and then aggregates the prediction results of all subgraphs to recover the three-dimensional conformation of the whole molecule. This subgraph-level denoising training strategy improves the model's sensitivity to local perturbations and enhances the accuracy of overall conformation generation.
[0077] When the model sampling module 14 performs sampling, it initializes the three-dimensional atomic coordinates from the standard Gaussian distribution and combines DihedralsEncode with the trained neural network to gradually restore the real conformation through a time-reverse diffusion process, realizing an end-to-end generation process from molecular topological information to a complete three-dimensional conformation, which is suitable for practical applications such as virtual screening and molecular design.
[0078] Furthermore, the target-specific identification unit 2 of the intelligent drug discovery system 100 includes a generative molecular design module (such as one based on a variational autoencoder, graph generation network, or extended molecular graph model) capable of reverse-engineering molecules from target attributes or target structures. The system can synergistically optimize multiple objectives such as efficacy, synthetic feasibility, and ADMET properties, enhancing the innovation and controllability of new molecules and achieving intelligent molecule creation "from zero to one".
[0079] The intelligent drug development system 100 of this invention overcomes the limitations of traditional static models by setting up a multi-task drug screening unit 3, achieving intelligent dynamic analysis of drug-target interactions. Simultaneously, it efficiently simulates the complete dynamic process of drug-target interactions using AI-accelerated molecular dynamics simulations (such as machine learning force fields), and intelligently analyzes massive simulation trajectories using deep learning models (such as 3D-CNN / GNN), automatically and accurately identifying key binding states, transition states, intermediate conformations, and energy barriers, revealing complex dynamic binding pathways that traditional methods cannot capture. Furthermore, the AI-enabled platform predicts the changing trends of binding modes under different conditions based on dynamic data, providing deep mechanistic insights beyond static docking, thereby significantly improving the accuracy, efficiency, and scientific rigor of drug screening.
[0080] Meanwhile, in traditional drug development, the prediction and evaluation of key compound properties (such as efficacy, metabolism, and toxicity) face a dual challenge:
[0081] Firstly, the research and development process is highly dependent on experiments, which are costly and time-consuming. A high degree of reliance on numerous complex in vitro and in vivo experiments is the gold standard for obtaining reliable attribute data. However, these experiments have low throughput, are extremely expensive, and are time-consuming (usually taking months or even years). Furthermore, they are difficult to cover all potential risks in the early stages, leading to delays in the research and development process and persistently high costs.
[0082] Secondly, traditional CADD techniques have limited predictive capabilities: Although computer-aided drug design (CADD) techniques (such as molecular docking and pharmacophore models) aim to assist or partially replace experiments, their traditional methods are usually based on static, simplified molecular models and limited rule bases. They struggle to accurately simulate the dynamic behavior of compounds in real biological environments, complex metabolic pathways, and unexpected interactions with off-target proteins, resulting in insufficient accuracy and limited reliability in predicting multi-dimensional key properties (especially the toxicity and metabolic stability of complex molecules).
[0083] These two limitations together result in the difficulty of efficiently, accurately, and comprehensively predicting the overall characteristics of candidate molecules in the early stages of drug discovery, increasing the risk of failure and waste of resources in later research and development.
[0084] In this application, the drug property prediction unit 4 of the intelligent drug discovery system 100 utilizes advanced artificial intelligence technologies such as deep neural networks (DNNs) and graph neural networks (GNNs) to more accurately and comprehensively predict various key properties of drug molecules. The platform uses deep learning models to train large-scale, multi-dimensional drug datasets, effectively identifying deep-seated correlations between molecular structure and its complex properties such as biological activity, metabolic characteristics, and toxicity risks, providing a more scientific assessment. In this way, the intelligent drug discovery system 100 can comprehensively evaluate molecular characteristics at a very early stage of drug design, before initiating expensive experiments, effectively identifying potential molecules and eliminating high-risk candidates, thereby significantly reducing R&D costs, shortening the cycle, and increasing the success rate.
[0085] In fact, the candidate molecules screened by the traditional CADD method may theoretically have excellent binding ability or efficacy, but in actual synthesis, there may be problems such as difficulty in synthesis, complex steps, or unavailable raw materials. The lack of early consideration of the "manufacturability" of the molecules makes it difficult to convert the screening results into practical applications.
[0086] The personalized drug design unit 5 of the intelligent drug development system 100 introduces a synthesis route prediction model and a synthesis complexity assessment mechanism. Based on a reaction rule database and machine learning model, it automatically assesses the synthesis difficulty of candidate molecules and can predict possible synthesis routes, thereby controlling the synthesis complexity in the early stages of molecular design and ensuring the feasibility of the transformation of virtual screening results.
[0087] Furthermore, in existing technologies, the models and tools for each stage of drug development (molecule generation, screening, optimization, evaluation, and translation) are relatively independent, resulting in poor data flow and difficulty in forming a coherent integrated process. For example, the results output by the molecule generation model cannot be directly and efficiently integrated into the screening and optimization stages; the evaluation results are also difficult to provide timely feedback to guide the iteration of molecule generation, leading to "discontinuities" in the connection between different stages of development, redundancy and time consumption in the overall process, and difficulty in quickly responding to the high-efficiency needs of new drug development.
[0088] The Dynamic Feedback Learning Unit 6 of the Intelligent Drug Development System 100 constructs a closed-loop intelligent drug design system encompassing "molecule generation-screening-optimization-evaluation-transformation." Molecules generated by molecule generation models (such as target-guided design and conformation generation models) can be directly transferred to the screening module for initial screening using molecular attribute prediction and multi-dimensional screening rules. Screened molecules then enter the optimization phase, where their structures are iteratively optimized based on efficacy and synthetic feasibility data from the evaluation model. The evaluation phase is integrated throughout the entire process, providing quantitative data for each stage. The transformation phase, based on synthetic route planning and feasibility assessment results, propels the molecules from virtual design to actual production. Dynamic feedback and iterative iteration across all stages achieve a high degree of integration and intelligence in the R&D process, significantly improving the efficiency and success rate of new drug development.
[0089] In summary, the intelligent drug development system 100 of the present invention integrates a three-dimensional conformation generation unit 1, a target specificity identification unit 2, a multi-task drug screening unit 3, a drug attribute prediction unit 4, a personalized drug design unit 5, and a dynamic feedback learning unit 6. Through a multi-unit collaborative mechanism, it integrates intelligent models of each stage of the drug design process to form a complete, closed-loop artificial intelligence-driven drug development platform.
[0090] The above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention.
Claims
1. An intelligent drug development system, characterized in that, include: Three-dimensional conformation generation unit: By combining dihedral subgraph structure representation and diffusion probability modeling, complex molecular three-dimensional conformations are modeled while preserving local molecular conformation information and overall topological structure. Target-specific identification unit: used for intelligent dynamic analysis of drug-target interactions; Multi-task drug screening unit: used to generate multiple molecular conformations and select the conformation that best meets the requirements of drug development through automated optimization algorithms; and to perform drug screening based on dynamic data to predict the changing trend of drug-target binding mode under different conditions. Drug property prediction unit: at least using deep neural networks and graph neural networks, to predict multiple key properties of drug molecules screened by the multi-task drug screening unit; Personalized drug design unit: It can reverse generate drug molecules from target properties or target structure to complete the design of personalized drugs; Dynamic feedback learning unit: provides quantitative basis for the output of the unit; and provides synthetic route planning and feasibility assessment results.
2. The intelligent drug development system according to claim 1, characterized in that, The three-dimensional image generation unit includes a conversion module, a noise processing module, a model training module, and a model sampling module.
3. The intelligent drug development system according to claim 1, characterized in that, The conversion module is used to analyze and convert the input molecular graph structure to generate a set of dihedral subgraphs for diffusion modeling.
4. The intelligent drug development system according to claim 2, characterized in that: The set of dihedral subgraphs includes multiple dihedral subgraphs; each dihedral subgraph inherits the corresponding atomic feature information, bond feature information, and dihedral angle feature.
5. The intelligent drug development system according to claim 2, characterized in that: The conversion module is also used to calculate and obtain the atomic feature vector, bond feature vector and angle feature vector corresponding to each dihedral subgraph in the dihedral subgraph set.
6. The intelligent drug development system according to claim 1, characterized in that: The noise processing module is used to progressively add Gaussian noise to the three-dimensional spatial coordinates of atoms in each dihedral sub-graph generated in the conversion module, so that the atomic coordinates of each atom approximate a standard Gaussian distribution, thereby obtaining a highly perturbed molecular noise conformation.
7. The intelligent drug development system according to claim 6, characterized in that: The noise processing module is also used to perform inverse denoising to reconstruct the original molecular conformation from the highly perturbed molecular noise conformation through the inverse denoising.
8. The intelligent drug development system according to claim 1, characterized in that: The model training module is used to predict the direction and magnitude of the perturbation of each atom in the dihedral subgraph at the current time step based on the noise score of the atom in each dihedral subgraph. The noise scores of the atoms in each dihedral subgraph are aggregated to form the estimated atomic coordinate perturbation value of the entire molecule, which is compared with the actual amount of noise added, the loss function is calculated, and the model parameters are optimized through backpropagation to complete the training of the model.
9. The intelligent drug development system according to claim 1, characterized in that: The model sampling module can initialize three-dimensional atomic coordinates based on a standard Gaussian distribution, and, in conjunction with the transformation module and the neural network trained by the model training module, gradually restore the true conformation of the molecule through a time-reverse diffusion process.