An autonomous evolutionary drug discovery and delivery collaborative method and system based on a large language model
By employing a large language model-driven autonomous evolutionary collaborative approach to drug discovery and delivery, and through the collaborative optimization of the generation and delivery systems, the problems of low efficiency, high cost, and delayed drugability prediction in traditional drug development have been solved, thus achieving an efficient and accurate drug development process.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SICHUAN AGRI UNIV
- Filing Date
- 2026-02-25
- Publication Date
- 2026-05-01
AI Technical Summary
The current drug development process suffers from problems such as low efficiency, high cost, delayed prediction of drugability, lack of independent decision-making and self-evolution capabilities. In the traditional model, each link is independent and lacks collaborative optimization, resulting in long development cycles and low success rates.
We adopted a collaborative approach to drug discovery and delivery based on a large language model, generated a candidate molecule library through an SE(3) equivariant hybrid generation model, combined with a two-stage funnel-shaped high-throughput virtual screening process and a three-stage delivery scheme design, and optimized the delivery system using molecular dynamics simulation and reinforcement learning to build a full-chain intelligent R&D system.
It has improved the efficiency and precision of drug development, reduced reliance on physical trial and error, achieved pre-process synergistic optimization of molecular generation and delivery systems, improved the accuracy of drugability prediction and autonomous decision-making capabilities, and shortened the development cycle.
Smart Images

Figure CN121747689B_ABST
Abstract
Description
A collaborative method and system for autonomous evolutionary drug discovery and delivery based on a large language model Technical Field
[0001] This invention relates to the field of collaborative drug discovery and delivery, specifically to a method and system for collaborative drug discovery and delivery based on a large language model and an autonomous evolutionary model. Background Technology
[0002] The drug development field widely employs the traditional "empirical physical trial-and-error" development model. This model's development process is highly linear, primarily covering key stages such as target discovery and validation, lead compound discovery, lead compound optimization, drug delivery system design, multi-scale validation, and preclinical evaluation.
[0003] Target discovery and validation rely on literature mining and high-throughput biological experiments, which typically take several years and lack synergistic connection with subsequent molecular design. Lead compound discovery often relies on high-throughput wet screening or basic virtual screening to select a small number of active molecules from millions of compounds. Subsequent optimization depends on repeated chemical synthesis and activity testing, with frequent human intervention. Drug delivery system design lags behind molecular optimization, relying solely on the experience of researchers to select carriers and formulations without synergistic adaptation to the characteristics of drug molecules. Multi-scale validation and preclinical evaluation rely on physical validation methods such as in vitro cell experiments and animal models, which are time-consuming and costly.
[0004] While there are some distributed computational tools in the existing technology (such as single molecular generation tools, property prediction tools, molecular docking tools, etc.), each tool has independent functions and lacks a unified scheduling and data interaction mechanism, thus failing to build a fully intelligent R&D system.
[0005] In summary, the existing technology has the following drawbacks:
[0006] Low R&D efficiency, high cost and low success rate: The traditional model relies on a large number of physical trial and error experiments, which consumes huge resources and the R&D cycle is generally several years long. Due to the "double ten rule", the success rate of new drug launch is low.
[0007] The research and development stages are independent of each other: target discovery, molecular design, drugability assessment, and delivery system design are linearly linked and lack a synergistic optimization mechanism. For example, the design of the delivery system lags behind the molecular generation, resulting in the common situation of "molecular activity meeting the standard but delivery failure".
[0008] Delayed drugability prediction: Key drugability indicators such as ADMET properties (absorption, distribution, metabolism, excretion, toxicity) are only evaluated to a limited extent in the later stages of molecular screening, and cannot be prospectively optimized in the early stages of molecular design, resulting in a high proportion of products being eliminated in the later stages due to drugability defects.
[0009] Relying on human intervention and lacking autonomous decision-making ability: Key decisions such as setting screening criteria, judging the direction of molecular optimization, and selecting carrier types in each stage all rely on the experience of medicinal chemists, lacking a data-driven autonomous decision-making mechanism, and are highly subjective and inefficient.
[0010] Lack of self-evolution capability: Existing tools can only complete single calculation tasks and cannot learn and optimize from historical R&D data. They require manual summarization of experience to adjust parameters, making it difficult to continuously improve R&D efficiency and success rate. Summary of the Invention
[0011] The purpose of this invention is to overcome the shortcomings of the prior art and provide a collaborative method and system for autonomous evolutionary drug discovery and delivery based on a large language model, which greatly improves the efficiency and accuracy of drug development, while realizing the autonomous evolution of the system.
[0012] The present invention achieves the above objectives by adopting the following technical solution: Firstly, the present invention provides a collaborative method for autonomous evolutionary drug discovery and delivery based on a large language model, comprising:
[0013] S1. For a given three-dimensional structure of a biological target, a candidate molecular library that is complementary to the target pocket and also takes into account drug-like properties is generated by the SE(3) isovariant hybrid generation model.
[0014] S2. Based on the two-stage funnel-shaped high-throughput virtual screening process, the molecules with the highest comprehensive potential are screened from the candidate molecule library.
[0015] S3. Generate customized delivery solutions through a three-stage process;
[0016] S4. Simulate the targeted delivery process of drug-carrier complexes and evaluate their efficiency in a digital twin model that simulates the real in vivo tumor microenvironment.
[0017] S5. Through molecular dynamics simulations, the atomic-level performance of the ternary system of drug, carrier, and target is verified in multiple dimensions.
[0018] Furthermore, in step S1, the generation process combines continuous coordinate generation and discrete type generation modeling paradigms.
[0019] Continuous coordinate generation:
[0020] Flow matching is used to model the three-dimensional coordinates of ligand atoms. The modeling process is equivalent to a controlled denoising process, that is, starting from a random atomic coordinate distribution of Gaussian noise, the model gradually optimizes the spatial position of each atom in multiple iterations based on the geometric and chemical constraints of the protein pocket environment.
[0021] Discrete type generation:
[0022] The Markov bridge model is used to model the element type of atoms and the type of chemical bonds between atoms.
[0023] Furthermore, the generation process includes an adaptive molecular size adjustment strategy, a multimodal data fusion and transfer learning strategy, and a preference alignment-driven directional optimization strategy.
[0024] Adaptive molecular size adjustment strategy:
[0025] Adaptive molecular size adjustment is achieved by adding virtual nodes to the atom types. During the generation process, the model learns to mark the types of redundant nodes in the pocket as virtual nodes. After sampling, all virtual nodes and their associated chemical bonds are removed.
[0026] Multimodal data fusion and transfer learning strategies:
[0027] This approach integrates target structure data from PDB (Protein Data Bank), compound activity data from ChEMBL or CrossDocked2020, and pre-trained ADMET property prediction data. It also utilizes reduce to preprocess crystal structures to construct a chemically complete binding pocket. The model is then pre-trained and fine-tuned using small datasets related to the corresponding target families. (ChEMBL represents a specific database name; ADMET represents the five key stages of a drug molecule's journey in vivo: Absorption, Distribution, Metabolism, Excretion, and Toxicity).
[0028] Preference alignment-driven directional optimization strategy:
[0029] Using a pre-trained base model, a large number of molecules are generated, and they are labeled as winning and losing samples according to one or more target attributes, forming a synthetic preference dataset.
[0030] Starting with the base model, fine-tuning is performed under a new objective function that aims to maximize the probability of the model generating winning samples while minimizing the probability of generating failing samples.
[0031] Furthermore, in step S2, the two-stage funnel-shaped high-throughput virtual screening process specifically includes:
[0032] In the first stage, molecules with obvious defects are filtered based on physicochemical properties and rapid docking. The screening rules include geometric and spatial constraints, physicochemical properties, and preliminary binding efficiency. Geometric and spatial constraints include no collisions within the ligand or ligand-pocket, and conformation minimization (RMSD < 1.0 Å). Physicochemical properties include the number of rotatable bonds < 10, 1 < LogP (Partition Coefficient) < 3, and 1 < SA Score (Synthetic Accessibility Score) < 10. Preliminary binding efficiency includes Vina Score < -0.6, Vina Efficiency < -0.3, and CNN (Convolutional Neural Network) Score > 0.4. Vina is short for AutoDockVina software, which is essentially a calculated and predicted value of the binding free energy ΔG, one of the core evaluation indicators in computer-aided drug design. Efficiency is a key optimization metric in medicinal chemistry and computational chemistry, used to comprehensively evaluate the binding strength and molecular size of compounds in the early stages of drug design, and is expressed as efficiency based on the number of heavy atoms.
[0033] In the second stage, a multidimensional comprehensive evaluation was conducted based on the pre-trained model. Using ADMETlab 3.0 and GNINA (Grid-based Neural Interaction Analysis), a comprehensive potential scoring model was constructed through weighted summation. The weights were set as follows: GNINA affinity assessment 0.3, GNINA convolutional network prediction 0.2, hERG cardiotoxicity prediction 0.1, drug similarity 0.08, human intestinal absorption 0.08, water solubility 0.05, synthetic accessibility 0.05, lipid-water partition coefficient 0.05, molecular weight 0.05, and hydrogen bond acceptor / donor ratio 0.04.
[0034] ADMETlab 3.0 is a powerful, free online platform specifically designed for the comprehensive prediction and evaluation of ADMET (absorption, distribution, metabolism, excretion, toxicity) and related properties of chemical molecules.
[0035] Furthermore, step S3 specifically includes:
[0036] In the first stage, drug delivery route prediction was performed based on XGBoost (eXtreme Gradient Boosting). XGBoost was trained using a dataset containing over 2000 known drugs and their drug delivery routes. For each molecule in the dataset, over 200 physicochemical property descriptors were extracted, including molecular weight, lipid-water partition coefficient, and topological polar surface area. These were combined with molecular fingerprinting to capture structural information.
[0037] The second stage involves carrier strategy screening based on expert rules, with the core being an if-then rule tree, which selects carrier types based on molecular hydrophilicity / hydrophobicity and functional requirements.
[0038] The third stage involves optimizing carrier parameters based on reinforcement learning and molecular dynamics surrogate models.
[0039] The state space of the reinforcement learning framework is a vector of carrier parameters, including the core material composition, carrier particle size, type and density of surface-modified ligands, and adjustable parameters for drug loading. The action space is the parameter fine-tuning operation, and the reward function is:
[0040] ;
[0041] Where f() represents the performance term to be maximized, g() represents the penalty term to be minimized, w1 and w2 represent weight sparsity, Encapsulation refers to the ability and efficiency of a delivery carrier (such as lipid nanoparticles) to load and protect its internal therapeutic cargo (such as mRNA, small molecule drugs, siRNA), Targeting refers to the ability of a delivery carrier to actively or passively transport drugs to specific diseased tissues, cells, or even organelles, and Immunogenicity refers to the potential risk of the delivery carrier itself triggering an unexpected and harmful immune response in the body.
[0042] The molecular dynamics reward surrogate model is constructed as follows:
[0043] To build an offline database, firstly, an offline simulation system library containing dozens of known drug and carrier monomer combinations is constructed.
[0044] Perform molecular dynamics simulations for each system in the simulation system library, and conduct time-limited molecular dynamics simulations under simulated physiological conditions to generate physically realistic atomic motion trajectory data.
[0045] Physical observables extraction involves calculating and extracting a series of physical quantities characterizing intermolecular interactions and dynamic behavior from simulated trajectory data, including the interaction energy between drug molecules and carrier materials, radial distribution function, and diffusion coefficient.
[0046] The molecular dynamics reward surrogate model training uses physical quantities extracted from molecular dynamics simulations as feature inputs and known macroscopic performance indicators such as encapsulation rate as labels to train a small MLP (Multilayer Perceptron) neural network. The trained MLP neural network is the reward surrogate model.
[0047] Furthermore, step S4 specifically includes:
[0048] Construct a three-dimensional digital twin model of the tumor microenvironment that includes physical and chemical features;
[0049] The digital twin model of the tumor microenvironment includes an extracellular matrix physical barrier, a chemokine concentration field, and a dynamic tumor cell cluster. The extracellular matrix physical barrier is simulated by randomly distributed obstacles in three-dimensional space to simulate a dense extracellular matrix network composed of collagen. The chemokine concentration field simulates vascular endothelial growth factor secreted by tumor cells. A three-dimensional scalar field is constructed in the simulation environment with the tumor region as the center and the concentration gradually decreasing towards the periphery.
[0050] The hybrid navigation algorithm is divided into long-range chemotaxis navigation and short-range Q-Learning navigation. Long-range chemotaxis navigation uses a greedy algorithm to move towards the direction of maximum concentration gradient. The state space of short-range Q-Learning navigation is defined as the relative vector to the nearest target point, the local density of the surrounding ECM (Extracellular Matrix), and the local VEGF (Vascular Endothelial Growth Factor) gradient direction. The action space consists of discrete three-dimensional movement commands, and the reward function is... The final output includes average targeting time, path length, and successful targeting rate.
[0051] Furthermore, step S5 specifically includes:
[0052] The first verification dimension is the stability and affinity of the drug-target interaction;
[0053] Combining pattern and interaction analysis:
[0054] First, AutoDock Vina was used to refine the binding mode. Then, PLIP (Protein-Ligand Interaction Profiler) was used to identify and visualize the set interaction types to ensure that the binding mode conforms to the principles of medicinal chemistry. The set interaction types include hydrogen bonds, hydrophobic interactions, salt bridges, and π-π stacking.
[0055] Calculations based on free energy:
[0056] The binding free energy between the drug and the target was calculated using molecular mechanics or the generalized Born surface area method.
[0057] Combined with stability assessment:
[0058] An atomic-level model of the drug-target complex was constructed. Molecular dynamics simulations were performed at set times under an explicit water model and physiological ion concentrations. The stability of the binding posture was evaluated by analyzing the root mean square deviation of the ligand relative to the target binding pocket in the simulated trajectory.
[0059] The second verification dimension is the encapsulation stability and environmental responsiveness of the drug-carrier interaction;
[0060] Encapsulation stability simulation: A complete drug-carrier complex model is constructed, and time-limited molecular dynamics simulations are performed in a simulated blood environment. By monitoring the relative position and distance between the drug molecules and the carrier core, the risk of drug leakage is determined, thereby verifying the encapsulation stability of the carrier.
[0061] Environmentally responsive release simulation: Two parallel molecular dynamics simulations were conducted, with the control group in a physiological environment and the experimental group in a simulated tumor microenvironment. By comparing the changes in the carrier structure and the escape behavior of drug molecules from the carrier core in the two simulations, the environmental responsiveness characteristics of the carrier were verified intuitively and quantitatively.
[0062] The third verification dimension is the overall colloidal and cyclic stability of the delivery system;
[0063] Surface property simulation and evaluation: For surface-modified supports, molecular dynamics simulations were used to verify the conformation, coverage density, and hydration layer of polyethylene glycol chains on the support surface.
[0064] Colloidal stability assessment: By calculating the changes in solvent-accessible surface area and gyration radius of the entire nanoparticle over time during the simulation process, the aggregation tendency of the nanoparticle in solution is indirectly assessed.
[0065] Secondly, the present invention provides a collaborative system for autonomous evolutionary drug discovery and delivery based on a large language model, used to implement the aforementioned collaborative method for autonomous evolutionary drug discovery and delivery based on a large language model. The system includes:
[0066] The system consists of a central decision-making layer, a core execution layer, and a support and guarantee layer. The central decision-making layer is the decision-making brain of the LLM, responsible for scheduling and managing the entire drug development workflow. It dynamically allocates tasks to various modules of the core execution layer and receives feedback from various modules of the core execution layer to generate decision instructions and iterative optimization instructions. At the same time, it provides interpretability analysis throughout the entire process.
[0067] The core execution layer includes a drug molecule generation module, a virtual screening and property prediction module, an intelligent delivery system design module, a targeted pathway planning module, and a multi-dimensional comprehensive verification module;
[0068] The molecular generation module generates a candidate molecule library that is complementary to the target pocket and takes into account drug-like properties for a given three-dimensional structure of a biological target using the SE(3) isovariant hybrid generation model.
[0069] The virtual screening and property prediction module uses a two-stage funnel-shaped high-throughput virtual screening process to screen out molecules with the highest overall potential from the candidate molecule library.
[0070] The intelligent delivery system design module generates customized delivery solutions through a three-stage process;
[0071] The targeted pathway planning module simulates the targeted delivery process of drug-carrier complexes and evaluates efficiency within a digital twin model that simulates the real in vivo tumor microenvironment.
[0072] The multi-dimensional comprehensive verification module uses molecular dynamics simulations to confirm the atomic-level performance of the ternary system of drug, carrier, and target.
[0073] The support and guarantee layer includes a task queue management component, a data interaction interface component, and a closed-loop iteration engine component. The task queue management component is responsible for storing and distributing tasks, the data interaction interface component realizes data transmission and integration between modules, and the closed-loop iteration engine component stores and feeds back iteration optimization instructions.
[0074] The beneficial effects of this invention are as follows:
[0075] This invention adopts an asynchronous parallel architecture of "LLM decision brain + five core modules", which achieves high decoupling of modules through task ID and callback mechanism, avoids waiting loss in linear process and maximizes the utilization of computing resources.
[0076] This invention replaces traditional physical trial and error with full-chain computational simulation. From molecular generation and virtual screening to delivery system optimization and performance verification, all are accomplished with the help of SE(3) equivalent variation model, molecular dynamics simulation, digital twin simulation and other technologies, reducing the dependence on wet experiments.
[0077] This invention achieves efficient use of computational resources through a two-stage funnel-shaped screening process (rapid filtering + fine evaluation), quickly eliminating defective molecules and performing high-precision evaluation only on high-potential molecules, thereby improving screening efficiency.
[0078] This invention overturns the traditional concept of "molecule first, delivery later" and constructs a pre-delivery collaborative logic of "molecule generation - delivery adaptation". Through a three-stage delivery system design process (XGBoost predicts the drug delivery route → expert rule screening of the vector → reinforcement learning + molecular dynamics surrogate model to optimize parameters), it achieves integrated optimization of the two.
[0079] Multidimensional constraints are introduced during the molecular generation stage. ADMET property prediction data is integrated through a multimodal data fusion strategy. During the generation process, chemical spaces with better drug-like properties are explored first, and potential problems such as toxicity and solubility are eliminated in advance.
[0080] This invention constructs a closed loop of "execution-analysis-learning-optimization" with an LLM (Large Language Model) decision-making brain at its core. Through meta-analysis, it examines the data throughout the entire process to identify systemic bottlenecks (such as generally high molecular LogP values). Strategic iterative optimization instructions are generated and fed back to the corresponding modules by the closed-loop iterative engine component, adjusting molecular generation constraints, selecting model weights, and optimizing carrier strategies to achieve autonomous system evolution.
[0081] This invention utilizes transfer learning and preference alignment strategies in the molecular generation module to quickly adapt to new targets based on historical R&D data, thereby improving the accuracy of R&D in small sample scenarios. Attached Figure Description
[0082] Figure 1 is a flowchart of an intelligent tool management method provided by an embodiment of the present invention. Detailed Implementation
[0083] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings.
[0084] This invention provides a collaborative method for autonomous evolutionary drug discovery and delivery based on a large language model, as shown in Figure 1, comprising:
[0085] S1. For a given three-dimensional structure of a biological target, a candidate molecular library that is complementary to the target pocket and also takes into account drug-like properties is generated by the SE(3) isovariant hybrid generation model.
[0086] The generation process combines continuous coordinate generation and discrete type generation modeling paradigms.
[0087] Continuous coordinate generation:
[0088] Flow matching is used to model the three-dimensional coordinates of ligand atoms. The modeling process is equivalent to a controlled denoising process, that is, starting from a random atomic coordinate distribution of Gaussian noise, the model gradually optimizes the spatial position of each atom in multiple iterations based on the geometric and chemical constraints of the protein pocket environment.
[0089] Discrete type generation:
[0090] The Markov bridge model is used to model the element type of atoms and the type of chemical bonds between atoms.
[0091] The generation process includes an adaptive molecular size adjustment strategy, a multimodal data fusion and transfer learning strategy, and a preference alignment-driven directional optimization strategy.
[0092] Adaptive molecular size adjustment strategy:
[0093] Adaptive molecular size adjustment is achieved by adding virtual nodes to the atom types. During the generation process, the model learns to mark the types of redundant nodes in the pocket as virtual nodes. After sampling, all virtual nodes and their associated chemical bonds are removed.
[0094] Multimodal data fusion and transfer learning strategies:
[0095] We integrate PDB target structure data, ChEMBL or CrossDocked2020 compound activity data, and pre-trained ADMET property prediction data, and use reduce to preprocess the crystal structure to construct a chemically complete binding pocket. We then pre-train the model and fine-tune it using a small dataset related to the corresponding target family.
[0096] Preference alignment-driven directional optimization strategy:
[0097] Using a pre-trained base model, a large number of molecules are generated, and they are labeled as winning and losing samples according to one or more target attributes, forming a synthetic preference dataset.
[0098] Starting with the base model, fine-tuning is performed under a new objective function that aims to maximize the probability of the model generating winning samples while minimizing the probability of generating failing samples.
[0099] S2. Based on the two-stage funnel-shaped high-throughput virtual screening process, the molecules with the highest comprehensive potential are screened from the candidate molecule library.
[0100] In the first stage, molecules with obvious defects are filtered based on physicochemical properties and rapid docking. The screening rules include geometric and spatial constraints, physicochemical properties, and preliminary binding efficiency. Geometric and spatial constraints include no collisions within the ligand or ligand-pocket and conformation minimization (RMSD < 1.0 Å). Physicochemical properties include the number of rotatable bonds < 10, 1 < LogP < 3, and 1 < SA Score < 10. Preliminary binding efficiency includes Vina Score < -0.6, VinaEfficiency < -0.3, and CNN Score > 0.4.
[0101] In the second stage, a multidimensional comprehensive evaluation was conducted based on the pre-trained model. Using ADMETlab 3.0 and the GNINA tool, a comprehensive potential scoring model was constructed through weighted summation. The weights were set as follows: GNINA affinity assessment 0.3, GNINA convolutional network prediction 0.2, hERG cardiotoxicity prediction 0.1, drug similarity 0.08, human intestinal absorption 0.08, water solubility 0.05, synthetic accessibility 0.05, lipid-water partition coefficient 0.05, molecular weight 0.05, and hydrogen bond acceptor / donor ratio 0.04.
[0102] S3. Generate customized delivery solutions through a three-stage process;
[0103] In the first stage, XGBoost was used to predict drug delivery routes. XGBoost was trained using a dataset containing more than 2,000 known drugs and their drug delivery routes. For each molecule in the dataset, more than 200 physicochemical property descriptors were extracted, including molecular weight (MW), lipid-water partition coefficient (LogP), and topological polar surface area (TPSA), and molecular fingerprints were used to capture its structural information.
[0104] The second stage involves carrier strategy screening based on expert rules, with the core being an if-then rule tree, which selects carrier types based on molecular hydrophilicity / hydrophobicity and functional requirements.
[0105] The third stage involves optimizing carrier parameters based on reinforcement learning and molecular dynamics surrogate models.
[0106] The state space of the reinforcement learning framework is a vector of carrier parameters, including the core material composition, carrier particle size, type and density of surface-modified ligands, and adjustable parameters for drug loading. The action space is the parameter fine-tuning operation, and the reward function is:
[0107] ;
[0108] Where f() represents the performance term to be maximized, g() represents the penalty term to be minimized, and w1 and w2 represent weight sparsity.
[0109] The molecular dynamics reward surrogate model is constructed as follows:
[0110] To build an offline database, firstly, an offline simulation system library containing dozens of known drug and carrier monomer combinations is constructed.
[0111] Perform molecular dynamics simulations for each system in the simulation system library, and conduct time-limited molecular dynamics simulations under simulated physiological conditions to generate physically realistic atomic motion trajectory data.
[0112] Physical observables extraction involves calculating and extracting a series of physical quantities characterizing intermolecular interactions and dynamic behavior from simulated trajectory data, including the interaction energy between drug molecules and carrier materials, radial distribution function, and diffusion coefficient.
[0113] The molecular dynamics reward surrogate model training uses physical quantities extracted from molecular dynamics simulations as feature inputs and known macroscopic performance indicators such as encapsulation rate as labels to train a small MLP neural network. The trained MLP neural network is the reward surrogate model.
[0114] S4. Simulate the targeted delivery process of drug-carrier complexes and evaluate their efficiency in a digital twin model that simulates the real in vivo tumor microenvironment.
[0115] Construct a three-dimensional digital twin model of the tumor microenvironment that includes physical and chemical features;
[0116] The digital twin model of the tumor microenvironment includes an extracellular matrix physical barrier, a chemokine concentration field, and a dynamic tumor cell cluster. The extracellular matrix physical barrier is simulated by randomly distributed obstacles in three-dimensional space to simulate a dense extracellular matrix network composed of collagen. The chemokine concentration field simulates vascular endothelial growth factor secreted by tumor cells. A three-dimensional scalar field is constructed in the simulation environment with the tumor region as the center and the concentration gradually decreasing towards the periphery.
[0117] Hybrid navigation algorithms are divided into long-range chemotaxis navigation and short-range Q-Learning navigation. Long-range chemotaxis navigation uses a greedy algorithm to move towards the direction of maximum concentration gradient. The state space of short-range Q-Learning navigation is defined as the relative vector to the nearest target point, the local density of the surrounding ECM, and the local VEGF gradient direction. The action space is a discrete three-dimensional movement command, and the reward function is... The final output includes average targeting time, path length, and successful targeting rate.
[0118] S5. Through molecular dynamics simulations, the atomic-level performance of the ternary system of drug, carrier, and target is confirmed.
[0119] The first verification dimension is the stability and affinity of the drug-target interaction;
[0120] Combining pattern and interaction analysis:
[0121] First, AutoDock Vina was used to refine the binding mode. Then, PLIP was used to identify and visualize the set interaction types to ensure that the binding mode conforms to the principles of medicinal chemistry. The set interaction types include hydrogen bonds, hydrophobic interactions, salt bridges, and π-π stacking.
[0122] Calculations based on free energy:
[0123] The binding free energy between the drug and the target was calculated using molecular mechanics or the generalized Born surface area method.
[0124] Combined with stability assessment:
[0125] An atomic-level model of the drug-target complex was constructed. Molecular dynamics simulations were performed at set times under an explicit water model and physiological ion concentrations. The stability of the binding posture was evaluated by analyzing the root mean square deviation of the ligand relative to the target binding pocket in the simulated trajectory.
[0126] The second verification dimension is the encapsulation stability and environmental responsiveness of the drug-carrier interaction;
[0127] Encapsulation stability simulation: A complete drug-carrier complex model is constructed, and time-limited molecular dynamics simulations are performed in a simulated blood environment. By monitoring the relative position and distance between the drug molecules and the carrier core, the risk of drug leakage is determined, thereby verifying the encapsulation stability of the carrier.
[0128] Environmentally responsive release simulation: Two parallel molecular dynamics simulations were conducted, with the control group in a physiological environment and the experimental group in a simulated tumor microenvironment. By comparing the changes in the carrier structure and the escape behavior of drug molecules from the carrier core in the two simulations, the environmental responsiveness characteristics of the carrier were verified intuitively and quantitatively.
[0129] The third verification dimension is the overall colloidal and cyclic stability of the delivery system;
[0130] Surface property simulation and evaluation: For surface-modified supports, molecular dynamics simulations were used to verify the conformation, coverage density, and hydration layer of polyethylene glycol chains on the support surface.
[0131] Colloidal stability assessment: By calculating the changes in solvent-accessible surface area and gyration radius of the entire nanoparticle over time during the simulation process, the aggregation tendency of the nanoparticle in solution is indirectly assessed.
[0132] This invention also provides a collaborative system for autonomous evolutionary drug discovery and delivery based on a large language model, used to implement the collaborative method for autonomous evolutionary drug discovery and delivery based on a large language model described in this invention. The system includes:
[0133] The system consists of a central decision-making layer, a core execution layer, and a support and guarantee layer. The core component of the central decision-making layer is the LLM (Large Language Model) decision-making brain, which serves as the command center of the system and includes the following core functions.
[0134] Core Function 1: Dynamic orchestration and collaborative management of task flows;
[0135] This feature defines the LLM decision brain as the system's "central commander," responsible for the macro-scheduling and management of the entire drug development workflow.
[0136] Employing an event-driven asynchronous microservice architecture, once the workflow starts, the LLM decision-making brain places the initial task into the main task queue. Each functional module, acting as an independent work unit, listens for and retrieves tasks from the queue. When a module (such as "drug molecule generation") completes its computational task, it generates a "task completed" event containing the task ID and a result summary.
[0137] Upon receiving this event, the LLM decision-making brain will perform the following actions:
[0138] Status assessment: Parse the event content and update the global status of the task.
[0139] Dependency determination: Based on the preset workflow logic, determine the downstream dependent modules of the task.
[0140] Task distribution: Based on dependencies, new task instructions are generated and placed into the corresponding task queue.
[0141] This dynamic orchestration mechanism ensures high cohesion and low coupling between modules, greatly improving the system's robustness and scalability. Asynchronous parallel processing maximizes the utilization of computing resources and shortens the overall development cycle.
[0142] Core Function Two: Multi-dimensional data fusion analysis and professional insight generation;
[0143] This feature aims to address the interpretability issues of the massive amounts of heterogeneous data generated during drug development.
[0144] The LLM decision brain plays the role of an "AI expert," responsible for translating complex computational results into professional insights that humans can understand.
[0145] A complete R&D process produces highly complex, multi-dimensional data, including but not limited to:
[0146] Chemical information: SMILES molecular formula, molecular weight (MW), lipid-water partition coefficient (LogP), etc.
[0147] Bioactivity: docking affinity score, binding free energy (ΔG), ADMET property prediction score, etc.
[0148] Physical properties: root mean square deviation (RMSD), radius of gyration, etc. in molecular dynamics simulations.
[0149] Delivery system parameters: carrier material, particle size, ligand density, drug loading, etc.
[0150] The system integrates this raw data into a structured JSON object and sends it to the LLM via API call, along with a carefully crafted Prompt containing role descriptions and task descriptions. Upon receiving the input, the LLM performs in-depth data fusion analysis and automatically generates a logically clear and detailed comprehensive evaluation report. This report can:
[0151] Summary: Key performance metrics for each candidate molecule.
[0152] Comparison: Advantages, disadvantages, and trade-offs of different regimens in terms of efficacy, drug-likeness, and safety.
[0153] Recommendation: Select one or more best-performing candidate compounds as lead compounds and provide clear, data-driven reasons for the recommendation.
[0154] This feature greatly accelerates the transformation from raw data to effective decisions, providing R&D personnel with high-quality decision support.
[0155] Core Function Three: Data-driven closed-loop feedback and system iterative evolution;
[0156] This function is the core mechanism for the platform's "autonomous evolution." After evaluating the results of a single task, the LLM decision-making brain can be triggered to execute a higher-level "meta-analysis." At this point, the object of analysis is no longer a single molecule, but the performance of the entire R&D process. The LLM examines all data to identify potential systematic biases or areas for optimization in the process.
[0157] For example, the brain might analyze and discover that although screened molecules have high affinity, their synthetic accessibility scores (SA scores) are generally low, or they generally have high lipophilicity (LogP), which poses challenges for subsequent chemical synthesis and formulation development. Based on such findings, LLM will generate a macroscopic, strategic optimization directive. This directive is not...
[0158] This is not simply about fine-tuning parameters, but about providing strategic guidance for the next round of R&D. For example, the system might output the following instruction: "Comprehensive analysis shows that the current molecule generation module tends to produce highly lipophilic backbones. To improve the druggability of candidate molecules from the source and simplify delivery system design, it is recommended that in the next R&D cycle, 'LogP value below 3.0' be introduced as a strong constraint into the molecule generation or screening stage, and appropriately adjusted..."
[0159] The weights of the docking scores are adjusted to find a better balance between molecular affinity and drug-likeness. This instruction will be recorded by the system and used as an initial configuration reference for the next round of tasks, thus constructing a complete data-driven closed loop of "execution-analysis-learning-optimization" and truly realizing the platform's autonomous evolution.
[0160] Support and Guarantee Layer Components: This layer comprises three key components. The task queue management component is directly controlled by the LLM decision-making brain, storing initial and iterative tasks, supporting task priority sorting and status tracking, and ensuring orderly task distribution. The data interaction interface component is based on a standardized API design, using "task IDs" to associate data between modules, enabling automatic transmission, parsing, and integration of structured data, ensuring high decoupling between modules. The closed-loop iteration engine component stores iterative optimization instructions generated by the LLM decision-making brain, automatically matching corresponding functional modules and adjusting their parameters / strategies to trigger the next round of development cycle.
[0161] The core execution layer will be described in detail below.
[0162] The core execution layer comprises five core functional modules: drug molecule generation, virtual screening and property prediction, intelligent delivery system design, targeted pathway planning, and multi-dimensional comprehensive validation. These modules function as independent work units, listening to the main task queue to acquire tasks, executing their corresponding functions asynchronously and in parallel, and outputting structured data upon completion, triggering a "task completed" event to feed back to the LLM decision-making brain.
[0163] Drug molecule generation module:
[0164] Drug molecule generation is the starting point of this invention's system. Its core task is to perform efficient and high-precision de novo molecule design targeting a given three-dimensional structure of a biological target. The drug molecule generation module of this invention employs an advanced generative framework to directly construct candidate compounds with high binding potential, good druggability, and structural novelty within the target binding pocket.
[0165] The core technology of the drug molecule generation module of this invention is a hybrid generation framework that can simultaneously process continuous data (atomic coordinates) and discrete data (atomic and chemical bond types) in three-dimensional space. The SE(3) isovariant hybrid generation model strictly follows the SE(3) isovariance, ensuring that the generation results of the model remain physically consistent regardless of how the target protein rotates or translates in space, thereby ensuring the physical realism and geometric rationality of the generated molecule's three-dimensional conformation.
[0166] The specific generation process combines two advanced generative modeling paradigms:
[0167] Continuous coordinate generation: Flow matching is used to model the three-dimensional coordinates of ligand atoms. This process can be understood as a controlled "denoising" process: starting from a random distribution of atom coordinates with Gaussian noise, the model gradually optimizes the spatial position of each atom over dozens of iterations, based on the geometric and chemical constraints of the protein pocket environment.
[0168] Discrete type generation: The Markov bridge model is used to model the element type of atoms and the chemical bond type between atoms.
[0169] Equivariance guarantee: The entire neural network backbone is implemented through a geometric vector perceptron to strictly ensure that all operations on geometric quantities satisfy SE(3) equivariance.
[0170] The greatest advantage of this framework is that it achieves the integrated and synchronous generation of molecular structure (discrete map) and its three-dimensional binding conformation (continuous coordinates), fundamentally avoiding the problem that traditional methods need to rely on highly uncertain molecular docking to predict the binding mode.
[0171] The generation process includes an adaptive molecular size adjustment strategy, a multimodal data fusion and transfer learning strategy, and a preference alignment-driven directional optimization strategy.
[0172] Strategy 1: Adaptive molecular size adjustment;
[0173] To address the limitation of traditional generative models requiring pre-specified molecule sizes, this invention introduces an end-to-end adaptive size estimation mechanism in its drug molecule generation module. This is achieved by adding a virtual node to the atom type. During generation, the model learns to predict nodes in redundant or inappropriate spatial locations within the pocket as virtual types. After sampling, all nodes marked as virtual and their associated chemical bonds are removed. This strategy allows the model to autonomously learn and determine the optimal molecule size based on the pocket size and shape, avoiding atomic collisions caused by an excessive initial atom count or insufficient pocket coverage due to an insufficient initial atom count.
[0174] Strategy 2: Multimodal data fusion and transfer learning;
[0175] To ensure that the generated model follows the professional principles of medicinal chemistry, this invention designs and implements a multimodal data fusion and transfer learning strategy.
[0176] Modality 1: Biological target structure data;
[0177] This serves as the "mold" and core constraints for generating molecules using the model. This invention integrates publicly available structural databases such as PDB and utilizes reduce to preprocess the crystal structure (hydrogenation, residue repair) to construct a binding pocket with complete chemical information. This three-dimensional environmental information serves as the primary input condition for the model.
[0178] Modal 2: Compound bioactivity data;
[0179] To enable the model to learn effective drug design patterns, this invention incorporates the large-scale ChEMBL compound activity database and the CrossDocked2020 protein-ligand complex database. By training on these validated datasets, the model can implicitly learn the structural characteristics of highly active compounds and the universal intermolecular interaction patterns.
[0180] Modal 3: ADMET property prediction data;
[0181] To incorporate druggability considerations during the generation stage, this approach innovatively integrates a pre-trained ADMET property prediction model into the denoising process of the diffusion model. At each generation step, the system performs a rapid ADMET property assessment on the intermediate generated molecular fragments. This assessment result serves as an additional conditional signal, guiding the generation process to prioritize the exploration of chemical spaces with superior druggability potential.
[0182] Furthermore, to address the issue of data sparsity in studies of specific targets, this invention employs a transfer learning strategy. The model is first pre-trained on the CrossDocked2020 universal protein-ligand complex dataset to learn general intermolecular interaction rules. Subsequently, for specific tasks, it is fine-tuned using a small-scale dataset related to that target family, significantly improving the model's performance and specificity in small-sample scenarios.
[0183] Strategy 3: Preference-aligned driven targeted optimization;
[0184] In addition to learning the data distribution, the system also provides an expert-knowledge-driven proactive intervention mechanism—preference alignment—whose technical approach is inspired by cutting-edge alignment techniques such as direct preference optimization. The technical implementation process of this strategy is as follows:
[0185] Generate a preference dataset: Using a pre-trained base model, generate a large number of molecules and label them as "winning" and "losing" samples based on one or more target attributes (such as QED, SA score, Vina efficiency score, etc.) to form a synthetic preference dataset.
[0186] Model alignment fine-tuning: Starting with the base model, fine-tuning is performed under a new objective function. This objective function aims to maximize the probability of the model generating "winning" samples while minimizing the probability of generating "failed" samples.
[0187] This mechanism allows medicinal chemists' qualitative design preferences to be transformed into quantitative guidance for the generation process, pulling the molecular generation trajectory from the central region of the original data distribution to a higher-value chemical subspace that satisfies specific preferences, thus enabling rapid and targeted iteration of candidate molecules.
[0188] Virtual screening and property prediction module:
[0189] After the drug molecule generation module produces an initial candidate compound library, the virtual screening and property prediction module of this invention will execute a two-stage, funnel-shaped high-throughput virtual screening process. Its core objective is to identify a small number of candidates with the highest overall potential from a large group of molecules through a series of computational evaluations ranging from rapid to refined, while ensuring efficiency. This achieves intensive use of computational resources and provides high-quality input for downstream modules.
[0190] Phase 1: Filtration based on physicochemical properties and rapid docking;
[0191] This stage is the initial hurdle in the screening process. Its design philosophy is to employ a series of computationally inexpensive, hard filters based on medicinal chemistry expert rules to rapidly and on a large scale screen the initial molecular library. This stage aims to quickly eliminate molecules with significant defects in basic physicochemical properties, structural flexibility, or initial binding conformation. Specific filtering rules and their design rationale are detailed in Table 1.
[0192] Table 1. Physical Chemistry and Preliminary Connection Rules Filtering Table
[0193]
[0194] Phase Two: Multidimensional Comprehensive Evaluation Based on Pre-trained Models;
[0195] Molecules that pass the rapid filtering in the first stage will proceed to this stage for a more precise and comprehensive multi-dimensional evaluation. The core of this stage is to utilize a series of mature pre-trained models to comprehensively predict the binding affinity and ADMET (absorption, distribution, metabolism, excretion, and toxicity) properties of molecules. This invention employs the ADMET property prediction tool ADMETlab 3.0 and the molecular docking tool GNINA, which can perform high-precision quantitative predictions based on molecular structure.
[0196] To achieve an objective and comprehensive ranking of molecules, this invention designs and implements a comprehensive potential scoring model. This model integrates the predicted scores of multiple key performance indicators into a single, quantifiable comprehensive score through a weighted summation method. As shown in Table 2 below, the weighting coefficients reflect the comprehensive consideration of efficacy, safety, and drugability in modern drug development.
[0197] Table 2 Parameters and Weights of the Comprehensive Potential Scoring Model
[0198]
[0199] Through this two-stage screening process, the system can reduce the size of the candidate molecule library from hundreds to 3-5 "hero molecules" with the highest overall scores. These molecules are then submitted to subsequent modules for more in-depth delivery system design and performance verification.
[0200] Intelligent Delivery System Design Module:
[0201] In traditional drug development processes, the design of delivery strategies often lags behind the discovery of active compounds, which is one of the important reasons for the failure of candidate drugs in the preclinical or clinical stages. The core innovation of the intelligent delivery system design module of this invention lies in the fact that, through a three-stage, macro-to-micro intelligent decision-making process, the delivery system is matched and finely optimized for drug discovery in the early stages, thereby improving the druggability of candidate drugs in advance.
[0202] Phase 1: Predicting drug delivery routes based on machine learning;
[0203] Before designing specific delivery vectors, the primary task is to determine the optimal route of administration for the molecule. To this end, this invention constructs a classification model based on the XGBoost algorithm. This model is trained on a dataset containing over 2000 known drugs and their routes of administration, compiled from publicly available databases such as ChEMBL. For each molecule in the dataset, this invention extracts over 200 physicochemical property descriptors (such as molecular weight (MW), lipid-water partition coefficient (LogP), and topological polar surface area (TPSA), and combines them with molecular fingerprints to more comprehensively capture its structural information. The model is trained to perform high-precision classification of 10 common routes of administration, including oral, intravenous, and subcutaneous injection.
[0204] For candidate molecules screened by the upstream module, the model can quickly predict their most suitable route of administration, providing a clear strategic direction for subsequent vector design.
[0205] Phase Two: Carrier Strategy Selection Based on Expert Rules;
[0206] Once the macroscopic route of administration (e.g., "intravenous injection") is determined, the system invokes a built-in expert rule decision engine to quickly narrow down the selection of delivery vectors. This engine simulates the decision-making logic of pharmaceutical formulation experts, and its core is a structured if-then rule tree designed to filter out the few most promising vector types from a vast library of potential vectors.
[0207] Example of a decision-making process (taking intravenous drug administration as an example):
[0208] Assessing basic physicochemical properties: The engine first analyzes the molecule's hydrophilicity / hydrophobicity (e.g., LogP value) to determine its solubility in aqueous solvents. If the solubility is poor, a basic library containing solubilizing carrier options such as polymer micelles, liposomes, and solid lipid nanoparticles is matched.
[0209] Assessing Functional Requirements: Subsequently, the engine will consider the functional requirements of the drug, such as targeting and toxicity. If the candidate molecule is a highly active antitumor drug with potential systemic toxicity, the engine will determine targeted delivery and controlled release as the primary goals and recommend the use of "stimulus-responsive nanocarriers" (such as pH-sensitive ones) or "actively targeted long-circulating carriers" (such as PEGylated carriers with antibody-modified surfaces).
[0210] Through this stage, an open design problem is transformed into a well-defined, multi-parameter optimization problem targeting a few carrier categories.
[0211] Phase 3: Optimization of carrier parameters based on reinforcement learning and molecular dynamics surrogate models;
[0212] This stage is the core technology of the intelligent delivery system design module, aiming to perform refined quantitative parameter optimization on the advanced carrier categories selected in the previous stage. This invention constructs an intelligent agent based on reinforcement learning, whose goal is to autonomously find the optimal combination of carrier parameters in a high-dimensional design space.
[0213] Reinforcement learning framework definition:
[0214] State space: A parameter vector describing the current delivery carrier design scheme, including adjustable parameters such as core material composition, carrier particle size, type and density of surface-modified ligands, and drug loading.
[0215] Action space: Fine-tuning operations performed by the agent on the current state (design scheme), such as "reducing the particle size by 5nm" or "increasing the ligand density by 2%".
[0216] Reward function: An objective function used to quantify the merits of a design solution; its general form is:
[0217] ;
[0218] Where f() represents the performance term to be maximized, g() represents the penalty term to be minimized, and w1 and w2 represent weight sparsity.
[0219] Construction of a molecular dynamics reward surrogate model:
[0220] The bottleneck of reinforcement learning lies in how to provide immediate and accurate rewards for each "attempt" by the agent. To address this, we innovatively construct a reward proxy model based on molecular dynamics simulations to solve the problem of sparse reward signals.
[0221] The construction process of this proxy model is as follows:
[0222] Offline database construction: We first construct an offline simulation system library containing dozens of "known drug-carrier monomer" combinations.
[0223] Perform molecular dynamics simulations: For each system in the library, perform a 10 nanosecond (ns) molecular dynamics simulation in a simulated physiological environment (including water, ions, etc.) to generate physically realistic atomic motion trajectory data.
[0224] Physically observable quantities extraction: From the simulated trajectory, a series of physical quantities that can characterize intermolecular interactions and dynamic behavior are calculated and extracted, such as the interaction energy between drug molecules and carrier materials, radial distribution function, diffusion coefficient, etc.
[0225] Proxy Model Training: This invention uses physical quantities extracted from molecular dynamics simulations as feature inputs (X) and the known macroscopic performance index encapsulation rate as labels (Y) to train a small MLP neural network. This trained model is the reward proxy model.
[0226] Ultimately, this agent model can quickly predict the performance of any carrier based on the arbitrary carrier parameters (states) given by the RL agent and return a reward value, thereby guiding the agent to conduct efficient and physically realistic exploration and optimization in a high-dimensional parameter space, and finally find the optimal carrier design scheme.
[0227] To ensure that the surrogate model can faithfully reproduce the results of molecular dynamics simulations, we rigorously evaluated its performance using 5-fold cross-validation. The model performed exceptionally well in predicting the "encapsulation efficiency," achieving a coefficient of determination R² = 0.92 and a mean absolute error of 3.5%.
[0228] Targeted path planning module:
[0229] This invention's targeted pathway planning module aims to evaluate the system-level delivery efficiency of the "drug-carrier" complex designed in the upstream module through dynamic simulation. Its core task is to evaluate and plan the optimal path for the delivery agent to traverse physiological barriers from the dosing point and ultimately reach the lesion target precisely within a digital twin model simulating the real in vivo tumor microenvironment. This not only verifies the dynamic performance of the delivery system design but also comprehensively demonstrates the agent's autonomous decision-making capabilities.
[0230] Simulation Environment Construction: Digital Twin of the Tumor Microenvironment;
[0231] To ensure the biological relevance of the path planning simulation, this invention constructs a three-dimensional digital twin model of the tumor microenvironment that includes key physical and chemical features.
[0232] Physical barriers – extracellular matrix (ECM): These are simulated by randomly distributed obstacles in three-dimensional space, mimicking the dense extracellular matrix network composed of collagen and other proteins. This network forms complex, maze-like porous channels, constituting physical barriers that delivery agents must traverse, thus simulating the spatial steric hindrance effect within tissues.
[0233] Chemical Beacon - Chemokine Concentration Field: Simulating biomarkers such as vascular endothelial growth factor (VEGF) secreted by tumor cells, a three-dimensional scalar field is constructed in a simulation environment, centered on the tumor region and with concentrations gradually decreasing outwards. This concentration gradient field provides long-range, natural chemical navigation signals for the delivery agent.
[0234] Dynamic Targets - Tumor Cell Clusters: Targets (tumor cells) in the environment are designed to move randomly within a small range to simulate the dynamic changes of tissues in the body, which places higher demands on the tracking and locking capabilities of the intelligent agent.
[0235] Navigation strategy: A hybrid path planning algorithm based on Q-Learning;
[0236] Considering that the delivery agent needs to cope with two different scenarios in the body, namely long-distance macroscopic approach and short-distance complex obstacle avoidance, this invention designs a hybrid navigation strategy based on Q-Learning, which organically combines long-distance chemotactic motion with short-distance autonomous learning obstacle avoidance.
[0237] Phase 1: Long-range chemotactic navigation;
[0238] When the agent's distance from the target exceeds a preset threshold, its primary task is to quickly and efficiently approach the target area. At this point, the agent employs a greedy algorithm based on the chemokine concentration gradient: at each decision step, it senses the concentration information of the surrounding environment and chooses to move in the direction of the largest concentration gradient. This is a computationally inexpensive and robust macroscopic navigation method that ensures the agent can rapidly traverse vast areas of normal tissue.
[0239] Phase Two: Short-Range Q-Learning Navigation;
[0240] Once the agent enters a complex region near the tumor with high concentrations and high ECM density, a simple chemotaxis strategy can easily lead to collisions or getting stuck in local optima. At this point, the agent will seamlessly switch to a decision model pre-trained with the Q-Learning algorithm for fine-grained navigation.
[0241] State space: A feature vector describing the agent's current local environment, defined as: [relative vector to the nearest target point, local density of the surrounding ECM, gradient direction of the local VEGF].
[0242] Action space: a set of discretized three-dimensional movement commands, such as [forward, backward, turn left, turn right, ascend, descend].
[0243] Reward function: To guide the agent in learning the optimal policy, the reward function is designed as follows:
[0244] ;
[0245] in, This represents the current instance of the agent's relationship with the target, where k represents the distance reward coefficient. This represents a large negative penalty that is triggered only when the agent collides with the ECM. The function explicitly guides the agent to achieve the dual objectives of "reaching the goal as quickly as possible" and "avoiding collisions at all costs."
[0246] Training process: Thousands of parameterized, randomized virtual tumor microenvironments are generated offline, allowing the agent to explore and experiment extensively within them. The Q-Table is continuously updated using the Q-Learning algorithm, ultimately learning and converging to obtain an optimal navigation strategy in complex near-field environments.
[0247] The final output of the targeted path planning module is a series of key performance indicators for each candidate solution: average targeting time.
[0248] Interval, pathway length, and successful targeting rate. These dynamic performance indicators provide crucial, system-level decision-making basis for ultimately selecting the optimal drug-carrier regimen.
[0249] Multi-dimensional comprehensive verification module:
[0250] The multi-dimensional comprehensive verification module is the final verification stage of the technical solution. It aims to use molecular dynamics simulation methods to comprehensively validate the performance of the finally selected "drug-carrier-target" ternary system at the atomic scale. Molecular dynamics simulation can provide dynamic information that static docking or machine learning models cannot reach, thus providing the deepest computational evidence for the effectiveness and stability of candidate solutions.
[0251] Dimension 1 of verification: Stability and affinity of drug-target interaction;
[0252] This section aims to validate the interaction between candidate drug molecules and their biological targets with high precision.
[0253] Binding mode and interaction analysis: First, the binding modes were refined using AutoDock Vina. Then, PLIP was used to identify and visualize key interaction types (such as hydrogen bonding, hydrophobic interactions, salt bridges, π-π stacking, etc.) to ensure that the binding modes conform to the principles of medicinal chemistry.
[0254] Binding free energy calculation: To obtain a more accurate affinity assessment than docking fraction, the molecular mechanics / generalized Born surface area method (MM / GBSA) was used to calculate the binding free energy between the drug and the target. This value provides a more reliable, physicochemical-based quantitative measure of molecular affinity.
[0255] Stability assessment: An atomic-level model of the drug-target complex was constructed, and molecular dynamics simulations were performed for 100 nanoseconds under an explicit water model and physiological ion concentrations. The stability of the binding posture was assessed by analyzing the root mean square deviation (RMSD) of the ligand relative to the target binding pocket in the simulated trajectory. A rapid RMSD curve reaching equilibrium with minimal fluctuations demonstrated the structural stability of the complex under dynamic physiological conditions.
[0256] Dimension 2 of verification: Encapsulation stability and environmental responsiveness of drug-carrier interactions;
[0257] This section aims to verify the core performance of the nanodelivery system designed for it: whether it can stably load drugs and intelligently release them under specific conditions.
[0258] Encapsulation stability simulation: A complete drug-carrier complex model was constructed, and molecular dynamics simulations were performed for 100 nanoseconds in a simulated blood environment (pH 7.4). By monitoring the relative position and distance between the drug molecules and the carrier core, the risk of premature drug leakage was assessed, thereby verifying the encapsulation stability of the carrier.
[0259] Environmentally Responsive Release Simulation: To verify the "intelligent" release capability of the carrier, we conducted two parallel molecular dynamics simulations. The control group was simulated at physiological pH (7.4), while the experimental group was simulated in a tumor microenvironment (pH 6.5). By comparing the changes in carrier structure and the escape behavior of drug molecules from the carrier core in the two simulations, the environmentally responsive characteristics of the carrier can be verified intuitively and quantitatively.
[0260] Verification Dimension 3: Overall colloidal and cyclic stability of the delivery system;
[0261] This section assesses the physical stability of the entire drug-carrier complex in the bloodstream from a more macroscopic perspective.
[0262] Surface property simulation and evaluation: For surface-modified carriers, molecular dynamics simulations were used to verify the conformation, coverage density, and hydration layer of polyethylene glycol (PEG) chains on the carrier surface, ensuring that they can form an effective "polymer brush" structure to provide steric shielding, thereby potentially reducing the risk of being recognized and cleared by the immune system.
[0263] Colloidal stability assessment: The aggregation tendency of nanoparticles in solution is indirectly assessed by calculating the changes in solvent-accessible surface area (SASA) and radius of gyration (R) of the entire nanoparticle during the simulation process.
[0264] A stable R value indicates that the nanoparticles maintained a stable size and shape during the simulation time, demonstrating good colloidal stability.
[0265] The alternative solutions of the present invention will be described in detail below.
[0266] Alternative or modified solutions for the overall architecture:
[0267] A serial or workflow-based solution driven by a "central decision-making brain";
[0268] Core Idea: Retain the core decision-making mechanism and full-chain functional modules of the LLM (Limited Least Machine) system, replace the existing asynchronous parallel scheduling mechanism, and adopt a pre-set logic tree or linear workflow pattern to drive the orderly progress of the R&D process, achieving the same end-to-end automated R&D goals. Technical Implementation: The LLM system does not require real-time dynamic orchestration of the entire process. It only needs to perform single-node logical judgments (such as "progress to the next stage" or "adjust the screening threshold") at each R&D node (e.g., completion of molecule generation or screening). Based on pre-set process rules, it triggers the execution of the next module. This solution simplifies the scheduling logic, ensures the standardization of the R&D process through fixed workflow templates, and is suitable for R&D scenarios with lower real-time collaboration requirements, while still reducing manual intervention and improving R&D efficiency.
[0269] A rapid iteration scheme based on local small models;
[0270] Core Idea: Replace existing large-scale Language Models (LLMs) with miniaturized Transformer models or expert system decision engines specifically trained for the field of medicinal chemistry. This reduces computational overhead while maintaining core decision-making and analytical capabilities. Technical Implementation: The miniaturized model is specifically trained using specialized data from the drug development field (such as molecular structure data, ADMET property data, and delivery system parameter data), enabling targeted data analysis and decision-making. The expert system decision engine performs logical judgments based on a rule base in medicinal chemistry (such as molecular screening rules and vector adaptation rules), eliminating the need for large-scale parameter training. This approach reduces hardware deployment costs and computational latency, making it suitable for resource-constrained scenarios while still enabling data-driven R&D decision-making and process optimization.
[0271] A semi-automated "human-machine harmony" solution;
[0272] Core Concept: This approach retains the data analysis and expert insight generation capabilities of LLM (Limited Learning Model), while adding a human expert interaction element. Human experts select iterative paths or confirm optimization instructions, serving as a backup or supplement to the fully autonomous evolutionary solution, balancing the need for automation and human intervention. Technical Implementation: After completing multi-dimensional data fusion analysis, LLM generates a comprehensive evaluation report and multiple iterative optimization candidate instructions, presented to human experts through the platform's interactive interface. Experts, based on their experience, select the optimal iterative path or modify instruction parameters and feed them back to the closed-loop iterative engine, triggering the next round of the R&D cycle. This solution leverages the data analysis capabilities of LLM while retaining human control over key decisions, making it suitable for R&D scenarios involving high-risk, high-value targets, and still enabling iterative optimization of the R&D process.
[0273] Alternatives to the core execution layer:
[0274] Alternatives to the drug molecule generation module;
[0275] Non-equivariant generative models:
[0276] Core idea: Replace the SE(3) isovariant hybrid generation model with a diffusion model or variational autoencoder (VAE) based on non-isovariant properties, and ensure the geometric fit between molecules and target pockets through manual rotation and translation scanning. Technical implementation: The non-isovariant generation model focuses on the generation of molecular structures. By performing multi-angle rotation and translation scanning on the target protein, multiple target conformational copies are generated. After generating candidate molecules for each copy, the binding effect is evaluated through molecular docking, and the molecule with the best geometric fit is selected. This approach does not rely on the SE(3) isovariant constraint. The post-processing step ensures the physical authenticity of the molecules and the fit with the target, and can still generate a candidate molecule library with novel structures and druggability.
[0277] Generation based on fragment assembly:
[0278] Core Idea: Replacing the atom-level de novo generation model, this approach utilizes a predefined library of chemical fragments (such as drug-active fragments and scaffold fragments). Through fragment combination, linkage, and optimization, complete molecules adapted to target pockets are generated. Technical Implementation: A large library of druggable fragments is pre-constructed. Based on the geometric and chemical characteristics of the target pocket, potentially suitable fragments are screened. Fragment splicing algorithms (such as fragment linkage and cyclization reaction simulation) are used to combine fragments into complete molecules. Further steps, including energy minimization and property optimization, enhance the druggability and binding activity of the molecules. This approach leverages the strengths of existing active fragments, reduces the druggability risks of generated molecules, and simultaneously meets the requirements of target adaptation and structural innovation.
[0279] Non-adaptive sizing design:
[0280] Core concept: This approach replaces the "virtual node" adaptive size adjustment mechanism with molecules generated using a fixed number of atoms. This is then combined with a subsequent pruning algorithm to remove redundant atoms or chemical bonds, adapting the molecule to the target pocket size. Technical analysis: Based on the atomic number range of common drug molecules, several fixed atomic number thresholds (e.g., 20-50 atoms) are preset to generate candidate molecules of corresponding sizes. Using molecular docking and spatial collision detection, redundant parts of the molecule that do not match the target pocket are identified. A pruning algorithm is then used to remove redundant atoms and chemical bonds, or adjust local structures, allowing the molecule to fit the target pocket space. This scheme simplifies the size adjustment logic in the generation stage, achieves molecular size optimization through post-processing, and still prevents issues such as atomic collisions and insufficient pocket occupancy.
[0281] Alternative solutions for the delivery system design module:
[0282] Carrier selection driven by physical simulation;
[0283] Core Concept: This approach replaces the initial screening method of "machine learning + expert rules" with a large-scale blind screening of vector libraries using high-throughput coarse-grained molecular dynamics (MD) simulations to select suitable vector types. Technical Analysis: A vector library covering multiple vector types (such as polymer micelles, liposomes, and nanoparticles) is constructed. A coarse-grained model simplifies molecular structures and reduces simulation computation costs. High-throughput MD simulations are performed for each "drug molecule-carrier" combination to evaluate key indicators such as the carrier's encapsulation ability and compatibility with the drug, selecting the best-performing vector type before proceeding to subsequent parameter optimization. This scheme directly assesses carrier suitability through physical simulations, avoiding prediction errors from machine learning models, while still achieving synergistic adaptation between drugs and carriers.
[0284] End-to-end deep learning optimization;
[0285] Core Concept: This approach uses an end-to-end graph neural network (GNN) to directly predict the nonlinear mapping relationship between carrier parameters and performance indicators such as encapsulation efficiency and targeting efficiency, replacing the parameter optimization combination of "reinforcement learning + MD surrogate model". Technical Analysis: Carrier parameters (core material components, particle size, ligand density, etc.) are used as input features, and encapsulation efficiency and targeting efficiency obtained from experiments or MD simulations are used as labels to train the graph neural network model. After training, the optimal carrier parameter combination can be directly output based on the physicochemical properties of the drug molecule and R&D needs. This solution simplifies the optimization process, eliminating the need for iterative exploration using reinforcement learning, and directly achieving parameter prediction through deep learning, while still achieving the goal of quantitative optimization of carrier parameters.
[0286] Alternative solutions for the route planning module:
[0287] Navigation based on the potential field method;
[0288] Core Concept: This approach utilizes the potential field method to construct a virtual potential field. The combined force of attraction (target) and repulsion (obstacles) guides the movement of the drug-carrier complex, achieving targeted path planning and replacing the Q-Learning hybrid navigation algorithm. Technical Analysis: In the digital twin model of the tumor microenvironment, the target region is designated as the gravitational source, generating an attractive force towards the target; physical barriers such as the extracellular matrix (ECM) are designated as repulsive sources, generating a repulsive force away from the obstacles. The drug-carrier complex, acting as a moving intelligent agent, adjusts its direction of movement under the combined force of attraction and repulsion, avoiding obstacles and moving towards the target. This solution eliminates the need for offline Q-Table training, achieving path planning through real-time calculation of the potential field force, while still outputting key indicators such as average targeting time and successful targeting rate.
[0289] Path planning based on topology maps;
[0290] Core Concept: Pre-segment the tumor microenvironment geometrically, construct a topological connectivity graph, and use A.S. or Dijkstra's algorithm to find the optimal path, improving path planning efficiency and replacing dynamic navigation. Technical Analysis: The tumor microenvironment is divided into multiple non-overlapping regions (e.g., normal tissue area, ECM-dense area, tumor periphery area) using a geometric segmentation algorithm. A topological connectivity graph is constructed with regions as nodes and connecting channels between regions as edges. Edge weights are assigned based on the traversal difficulty of each channel (e.g., ECM density, chemokine concentration), and the optimal path from the drug delivery point to the target region is calculated using A.S. or Dijkstra's algorithm. This approach simplifies the path search process by pre-constructing a topological graph, is suitable for static or low-dynamic tumor microenvironment models, and still ensures the effectiveness of the targeting path.
[0291] The invention will now be described in further detail with reference to specific examples.
[0292] Task definition:
[0293] The task of this case study is to design a novel small molecule inhibitor with high affinity, good druggability, and a well-defined delivery method for KRAS G12C, a high-value but difficult-to-drug target.
[0294] Target information input: When the mission starts, the high-resolution crystal structure PDB data of the KRAS G12C mutant is first obtained from the PDB.
[0295] Structural pretreatment: The original structure is standardized and pretreated, including removing water molecules and non-interacting ions, repairing missing side chains, adding hydrogen atoms, and precisely defining the target binding pocket of the covalent inhibitor.
[0296] Design goal setting: The initial design goal was given to the LLM decision-making brain: "Generate an initial library containing 300 candidate molecules, requiring the molecules to have high complementarity in geometry and chemical properties with the KRAS G12C target pocket, and to consider the optimization of ADMET properties during the generation stage."
[0297] Process execution:
[0298] After receiving the task, the central decision-making brain automatically initiated the preliminary process of drug discovery.
[0299] Molecular generation: The drug molecule generation module was invoked, and based on the preprocessed KRAS G12C pocket information, 300 novel initial candidate molecules were successfully generated within the agreed computation time.
[0300] First-stage screening: The 300 generated molecules were automatically fed into the virtual screening and property prediction module. After the first-stage filtering based on physicochemical properties and rapid docking rules, approximately 75% of the molecules were eliminated, and the remaining 72 candidate molecules entered the next stage.
[0301] Second-stage evaluation: These 72 molecules then underwent a refined evaluation using a comprehensive potential scoring model based on a pre-trained model. The system provided a comprehensive, multi-dimensional score for each molecule.
[0302] Screening Decision: Ultimately, based on CPS scores, the system successfully screened out the five candidate molecules with the highest overall potential, named CX-A-01 to CX-A-05. These five molecules achieved the best balance between predicted affinity, safety, and druggability, marking the successful completion of the screening phase.
[0303] Solution generation:
[0304] The top 5 candidate molecules selected are automatically submitted to the delivery system design module, allowing for the creation of customized delivery protocols for each molecule. Taking CX-A-01, which has the highest overall score, as an example:
[0305] Route of administration prediction: Based on the physicochemical property descriptors of the molecules, the XGBoost classification model predicts with high confidence that the optimal route of administration is "oral".
[0306] Carrier strategy screening: Based on the "oral" route and the characteristics of CX-A-01 as a highly effective anticancer drug, the expert rule engine recommends using "actively targeted liposomes" as the carrier category to achieve tumor targeting and reduce systemic toxicity.
[0307] Vector parameter optimization: A reinforcement learning agent is initiated to optimize the specific parameters of the liposomes in a high-dimensional parameter space. This process is guided in real time by a reward agent model trained by MD simulation, ultimately determining an optimal set of vector formulation parameters.
[0308] The other four molecules also went through a similar process, resulting in their respective delivery system designs.
[0309] Result evaluation:
[0310] Five complete drug-carrier protocols were then incorporated into a comprehensive performance validation module for in-depth molecular dynamics (MD) simulations. The simulation results (including binding stability RMSD, encapsulation stability, pH-responsive release efficiency, etc.) were then combined with all data from the preceding modules.
[0311] The LLM decision-making system utilized multidimensional data fusion analysis to transform all structured data into a comprehensive evaluation report. The report stated: "Candidate CX-A-01 exhibited the lowest and most stable RMSD value (<2.1 Å) in 100 ns MD simulations, and its matched liposome carrier demonstrated the most significant drug release behavior at pH 6.5. Based on comprehensive evaluation, C-KRAS-001 is the highest priority compound for this mission."
[0312] Targeted path planning:
[0313] Based on the LLM brain's decision-making, the optimal solution CX-A-01 and its delivery system were instantiated as a delivery agent and entered the target path planning module for final dynamic delivery simulation. In the digital twin model of the colorectal tumor microenvironment, this agent, utilizing a hybrid navigation strategy, demonstrated highly efficient target approach and obstacle avoidance capabilities in multiple repeated simulations, achieving an average target success rate exceeding 99%. Its simulated delivery path trajectory was recorded and visualized, providing an intuitive theoretical basis for the in vivo delivery behavior of this solution.
[0314] System self-evolution:
[0315] Following the completion of the entire task for KRAS G12C, the LLM decision-making system activated its data-driven closed-loop feedback mechanism, performing a meta-analysis of the entire process. It identified a systematic trend: the high-affinity molecules screened in this study generally exhibited higher LogP values.
[0316] Based on this insight, the system automatically generated and archived the following strategic iteration instructions to guide future R&D tasks: "Iteration Instructions: Comprehensive analysis shows that the current molecule generation module tends to produce highly lipophilic backbones. To improve the druggability of candidate molecules from the source and simplify delivery system design, the core strategy for the next R&D cycle is: In the molecule generation stage, introduce 'LogP value < 3' as a strong preference into the preference alignment module, while appropriately relaxing the pursuit of the ultimate initial docking score. The aim is to find a better balance between affinity and druggability, rather than simply maximizing efficacy."
[0317] Comparison of Effects (Compared to Existing Technologies: Traditional Drug Development Paradigm): Regarding the preclinical candidate molecule discovery cycle, existing technologies take 2-3 years, while this invention takes several tens of weeks, a reduction of over 90%. In terms of candidate molecule screening efficiency, existing technologies achieve 100,000+ molecules / month (wet experiments), while this invention achieves 300+ molecules / 3 days (computational simulation), a 100-fold increase in screening efficiency per unit time. Regarding the druggability prediction stage, existing technologies target the later stages of molecule optimization, while this invention targets the early stages of molecule generation, proactively avoiding druggability defects. Regarding drug-delivery system compatibility, existing technologies have a <30% compatibility rate (disjointed design), while this invention achieves >90% (co-optimization), a 200% improvement. Regarding dynamic delivery evaluation capabilities, existing technologies lack this capability, while this invention supports tumor microenvironment simulation evaluation, adding a new dimension of system-level performance verification. Regarding self-evolution capabilities, existing technologies lack this capability, while this invention supports an "execution-analysis-optimization" closed loop, continuously improving the targeting and success rate of research and development. Note: Existing technology data is derived from publicly available statistics based on the "Double Ten Law," while data for this invention is derived from the actual operational results of the above embodiments.
[0318] In summary, the core advantages and technical implementation logic of this invention are as follows:
[0319] First, it significantly improves R&D efficiency and greatly reduces costs and time.
[0320] Key advantages: It shortens the time for discovering preclinical candidate molecules from the traditional 2-3 years to several weeks to several months, increases the screening efficiency per unit time by more than 100 times, greatly reduces the number of physical trial and error experiments, and reduces the consumption of R&D resources.
[0321] Technical Implementation: An asynchronous parallel architecture consisting of "LLM Decision Brain + Five Core Modules" is adopted. The modules are highly decoupled through task IDs and callback mechanisms to avoid waiting losses in linear processes and maximize the utilization of computing resources.
[0322] This invention replaces traditional physical trial and error with full-chain computational simulation. From molecular generation and virtual screening to delivery system optimization and performance verification, all are accomplished with the help of SE (3) equivariant models, MD simulation, digital twin simulation and other technologies, reducing the dependence on wet experiments.
[0323] This invention achieves efficient use of computational resources through a two-stage funnel-shaped screening process (rapid filtering + fine evaluation), quickly eliminating defective molecules and performing high-precision evaluation only on high-potential molecules, thereby improving screening efficiency.
[0324] Second, simultaneously optimize drug-likeness and delivery suitability to reduce the risk of R&D failure;
[0325] Core advantages: The compatibility of drug delivery systems has been improved from less than 30% in existing technologies to over 90%, avoiding industry challenges such as "molecular activity meets standards but delivery fails" and "drug-related defects are exposed later," thereby improving the success rate of clinical translation.
[0326] Technical Implementation: This invention overturns the traditional concept of "molecule first, delivery later" and constructs a pre-delivery collaborative logic of "molecule generation-delivery adaptation". Through a three-stage delivery system design process (XGBoost predicts the drug delivery route → expert rule screening of the carrier → reinforcement learning + molecular dynamics surrogate model to optimize parameters), it achieves integrated optimization of the two.
[0327] This invention introduces multi-dimensional constraints at the molecular generation stage and integrates ADMET property prediction data through a multi-modal data fusion strategy. During the generation process, it prioritizes the exploration of chemical spaces with better drug-like properties and eliminates potential problems such as toxicity and solubility in advance.
[0328] This invention provides atomic-level performance confirmation through a multi-dimensional comprehensive verification module, and verifies drug-target binding stability, drug-carrier encapsulation stability, and colloidal cycling stability through 100ns MD simulation, providing dual assurance for drugability and delivery effectiveness.
[0329] Third, it possesses the ability to evolve independently and continuously improve the accuracy of research and development;
[0330] Core advantages: Breaking through the limitations of existing tools that are "single-task execution", it achieves self-optimization through a closed-loop feedback mechanism. After iteration, the accuracy of drug-likeness prediction has increased from 85% to 88%, adapting to the research and development needs of different targets.
[0331] Technical Implementation: This invention constructs a closed loop of "execution-analysis-learning-optimization" with the LLM decision brain as the core. It examines the data of the entire process through meta-analysis and identifies systemic bottlenecks (such as the generally high value of molecular LogP).
[0332] This invention generates strategic iterative optimization instructions, which are fed back to the corresponding modules by the closed-loop iterative engine component, adjusting molecular generation constraints, screening model weights, carrier optimization strategies, etc., to achieve autonomous system evolution.
[0333] This invention utilizes transfer learning and preference alignment strategies in the molecular generation module to quickly adapt to new targets based on historical R&D data, thereby improving the accuracy of R&D in small sample scenarios.
[0334] Fourth, enhance the scientific rigor and interpretability of decision-making, and reduce reliance on human intervention;
[0335] Core advantages: It breaks away from the reliance on human experience in existing technologies, enables data-driven decision-making throughout the entire R&D process, and generates comprehensive evaluation reports with clear comparisons of advantages and disadvantages and reasons for recommendations, thereby improving the credibility of decisions.
[0336] Technical Implementation: The LLM decision brain provided by this invention has the dual roles of "task scheduling" and "AI domain expert". It can integrate and analyze multi-dimensional heterogeneous data such as molecular structure, biological activity, physical properties, and delivery parameters, and transform complex calculation results into professional insights that humans can understand.
[0337] This invention uses a Comprehensive Potential Scoring Model (CPS) to quantify the comprehensive potential of a molecule by weighted summing of 10 key indicators (including efficacy, safety, and drugability), thus avoiding the subjective bias of human decision-making.
[0338] The data transmission between modules in this invention is based on standardized APIs and task ID association, enabling automatic integration of structured data without the need for manual import and export, thus reducing human error and communication costs.
[0339] Fifth, break through the bottleneck in the research and development of drug-resistant targets and expand the boundaries of drug development;
[0340] Core advantages: For targets such as KRAS G12C that are difficult to overcome with traditional technologies, it can explore new chemical spaces, provide innovative R&D solutions, and make up for the shortcomings of existing technologies.
[0341] Technical Implementation: The drug molecule generation module of this invention uses the SE (3) isovariant hybrid generation model to achieve the integrated generation of molecular structure and three-dimensional conformation, which can break through the limitations of existing compound libraries and explore chemical spaces that traditional trial and error methods cannot reach.
[0342] This invention integrates the "molecule generation-delivery system synergistic optimization" mechanism, and solves the problems of efficiency and specificity in delivering molecules to difficult-to-generate drug targets by means of targeted pathway planning and environmentally responsive delivery design.
[0343] The transfer learning strategy of this invention can address the challenge of sparse data on drug-resistant targets. By pre-training on a general dataset and fine-tuning on a target family dataset, the performance and specificity of the model on specific targets can be improved.
[0344] The above description is merely a preferred embodiment of the present invention. It should be understood that the present invention is not limited to the forms disclosed herein and should not be construed as excluding other embodiments. It can be used in various other combinations, modifications, and environments, and can be altered within the scope of the concept described herein through the above teachings or related technologies or knowledge. Modifications and variations made by those skilled in the art that do not depart from the spirit and scope of the present invention should be within the protection scope of the appended claims.
Claims
1. A collaborative method for autonomous evolutionary drug discovery and delivery based on a large language model, characterized in that, The steps include: S1. For a given three-dimensional structure of a biological target, a candidate molecule library that is complementary to the target pocket and also takes into account drug-like properties is generated by using the SE(3) isovariant hybrid generation model; S2. According to the two-stage funnel-type high-throughput virtual screening process, the molecules with the highest comprehensive potential are screened from the candidate molecule library. The two-stage funnel-shaped high-throughput virtual screening process specifically includes: The first stage filters molecules based on physicochemical properties and rapid docking, eliminating those with significant defects. Screening rules include geometric and spatial constraints, physicochemical properties, and preliminary binding efficiency. Geometric and spatial constraints include no collisions within the ligand or ligand-pocket, and conformation minimization (RMSD < 1.0 Å). Physicochemical properties include the number of rotatable bonds < 10, 1 < LogP < 3, and 1 < SA Score < 10. Preliminary binding efficiency includes Vina Score < -0.6, VinaEfficiency < -0.3, and CNN Score > 0.
4. The second stage conducts a multi-dimensional comprehensive evaluation based on a pre-trained model, utilizing ADMETlab. Version 3.0 and the GNINA tool are used to construct a comprehensive potential scoring model through weighted summation. The weights are set as follows: GNINA affinity assessment 0.3, GNINA convolutional network prediction 0.2, hERG cardiotoxicity prediction 0.1, drug similarity 0.08, human intestinal absorption 0.08, water solubility 0.05, synthetic accessibility 0.05, lipid-water partition coefficient 0.05, molecular weight 0.05, number of hydrogen bond acceptors or donors 0.04; S3, a customized delivery plan is generated through a three-stage process; the first stage is based on XGBoost for route of administration prediction. XGBoost is trained on a dataset containing more than a dozen known drugs and their routes of administration. For each molecule in the dataset, multiple physicochemical property descriptors are extracted, including molecular weight (MW), lipid-water partition coefficient (LogP), and topological polar surface area (TPSA), and molecular fingerprints are used to capture its structural information; the second stage is based on expert rule-based carrier strategy screening, the core of which is if- Then, a rule tree is used to screen carrier types based on molecular hydrophilicity / hydrophobicity and functional requirements; in the third stage, carrier parameter optimization is performed based on reinforcement learning and a molecular dynamics surrogate model; the state space of the reinforcement learning framework is a carrier parameter vector, including core material components, carrier particle size, types and densities of surface-modified ligands, and adjustable parameters for drug loading; the action space is for parameter fine-tuning operations; and the reward function is: Where f() represents the performance term to be maximized, g() represents the penalty term to be minimized, and w1 and w2 represent weight sparsity; A molecular dynamics reward surrogate model is constructed, and the construction process is as follows: An offline database is built, firstly, a simulation system library containing dozens of known drug-carrier monomer combinations is constructed offline; Molecular dynamics simulations are performed, for each system in the simulation system library, molecular dynamics simulations are performed under simulated physiological conditions for a set time, generating physically realistic atomic motion trajectory data; Physical observables are extracted, from the simulated trajectory data, a series of physical quantities characterizing intermolecular interactions and dynamic behavior are calculated and extracted, including the interaction energy between drug molecules and carrier materials, radial distribution function, and diffusion coefficient; The molecular dynamics reward surrogate model is trained, using the physical quantities extracted from the molecular dynamics simulation as feature inputs and the known macroscopic performance index encapsulation rate as labels, a small MLP neural network is trained, and the trained MLP... The neural network is a reward agent model; S4, in a digital twin model simulating the real in vivo tumor microenvironment, the targeted delivery process of the drug-carrier complex is simulated and its efficiency is evaluated; S5, through molecular dynamics simulation, the atomic-level performance of the ternary system of drug, carrier and target is verified in multiple dimensions.
2. The autonomous evolutionary drug discovery and delivery collaborative method based on a large language model according to claim 1, characterized in that, In step S1, the generation process includes a modeling paradigm that combines continuous coordinate generation and discrete type generation. Continuous coordinate generation: Flow matching is used to model the three-dimensional coordinates of ligand atoms. The modeling process is equivalent to a controlled denoising process, that is, starting from a random atomic coordinate distribution of Gaussian noise, the model gradually optimizes the spatial position of each atom in multiple iterations based on the geometric and chemical constraints of the protein pocket environment. Discrete type generation: The Markov bridge model is used to model the element type of atoms and the chemical bond type between atoms.
3. The autonomous evolutionary drug discovery and delivery collaborative method based on a large language model according to claim 1, characterized in that, In step S1, the generation process includes an adaptive molecular size adjustment strategy, a multimodal data fusion and transfer learning strategy, and a preference alignment-driven directional optimization strategy. Adaptive molecular size adjustment strategy: Adaptive molecular size adjustment is achieved by adding virtual nodes in the atom type. During the generation process, the model autonomously learns to mark the types of redundant spatial nodes in the pocket as virtual nodes. After sampling, all virtual nodes and their associated chemical bonds are removed. Multimodal data fusion and transfer learning strategy: PDB target structure data, ChEMBL or CrossDocked2020 compound activity data, and pre-trained ADMET property prediction data are integrated. The crystal structure is preprocessed using reduce to construct a chemically complete binding pocket. The model is pre-trained and then fine-tuned using datasets related to the corresponding target family. Preference alignment-driven directional optimization strategy: Using a pre-trained base model, a large number of molecules are generated, and they are labeled as winning and losing samples according to one or more target attributes, forming a synthetic preference dataset; Starting with the base model, fine-tuning is performed under a new objective function that aims to maximize the probability of the model generating winning samples while minimizing the probability of generating failing samples.
4. The autonomous evolutionary drug discovery and delivery collaborative method based on a large language model according to claim 1, characterized in that, Step S4 specifically includes: constructing a three-dimensional digital twin model of the tumor microenvironment containing physical and chemical features; the digital twin model of the tumor microenvironment includes an extracellular matrix physical barrier, a chemokine concentration field, and a dynamic tumor cell cluster. The extracellular matrix physical barrier is simulated by randomly distributed obstacles in three-dimensional space to simulate a dense extracellular matrix network composed of collagen. The chemokine concentration field simulates vascular endothelial growth factor secreted by tumor cells. A three-dimensional scalar field with the tumor region as the center and the concentration gradually decreasing towards the periphery is constructed in the simulation environment. The hybrid navigation algorithm is divided into long-range chemokine navigation and short-range Q-learning navigation. Long-range chemokine navigation uses a greedy algorithm to move towards the direction of maximum concentration gradient. The state space of short-range Q-learning navigation is defined as the relative vector to the nearest target point, the local density of the surrounding ECM, and the local VEGF gradient direction. The action space is a discrete three-dimensional movement command. The final output is the average targeting time, path length, and successful targeting rate.
5. The autonomous evolutionary drug discovery and delivery collaborative method based on a large language model according to claim 1, characterized in that, Step S5 specifically includes: First, the validation dimension: stability and affinity of the drug-target interaction; binding mode and interaction analysis: First, AutoDock Vina is used to accurately evaluate the binding mode. Then, PLIP is used to identify and visualize the set interaction types, ensuring that the binding mode conforms to the principles of medicinal chemistry. The set interaction types include hydrogen bonding, hydrophobic interaction, salt bridging, and π-π stacking; binding free energy calculation: molecular mechanics or the generalized Born surface area method is used to calculate the binding free energy between the drug and the target; binding stability assessment: an atomic-level model of the drug-target complex is constructed, and time-limited molecular dynamics simulations are performed under an explicit water model and physiological ion concentrations. The stability of the binding posture is assessed by analyzing the root mean square deviation of the ligand relative to the target binding pocket in the simulated trajectory; Second, the validation dimension: encapsulation stability and environmental responsiveness of the drug-carrier interaction; encapsulation stability simulation: a complete drug-carrier complex model is constructed, and time-limited molecular dynamics simulations are performed in a simulated blood environment. The encapsulation stability is assessed by monitoring the interaction between the drug molecule and the carrier. The relative position and distance of the carrier core are used to determine the risk of drug leakage, thereby verifying the encapsulation stability of the carrier; Environmental responsive release simulation: Two parallel molecular dynamics simulations are conducted, with the control group in a physiological environment and the experimental group in a simulated tumor microenvironment. By comparing the changes in carrier structure and the escape behavior of drug molecules from the carrier core in the two simulations, the environmental responsiveness of the carrier is verified intuitively and quantitatively; The third verification dimension is the colloidal and cyclic stability of the entire delivery system; Surface property simulation and evaluation: For surface-modified carriers, molecular dynamics simulation is used to verify the conformation, coverage density, and hydration layer of polyethylene glycol chains on the carrier surface; Colloidal stability evaluation: By calculating the changes in the solvent-accessible surface area and gyration radius of the entire nanoparticle over time during the simulation process, its aggregation tendency in solution is indirectly evaluated.
6. A collaborative system for autonomous evolutionary drug discovery and delivery based on a large language model, used to implement the collaborative method for autonomous evolutionary drug discovery and delivery based on a large language model as described in any one of claims 1-5, characterized in that, The system comprises a central decision-making layer, a core execution layer, and a support and guarantee layer. The central decision-making layer is the LLM decision-making brain, responsible for scheduling and managing the entire drug development workflow, dynamically allocating tasks to various modules of the core execution layer, and receiving feedback from various modules of the core execution layer to generate decision instructions and iterative optimization instructions, while providing interpretability analysis throughout the entire process. The core execution layer includes a drug molecule generation module, a virtual screening and property prediction module, an intelligent delivery system design module, a target pathway planning module, and a multi-dimensional comprehensive verification module. The drug molecule generation module generates a candidate molecule library that is complementary to the target pocket and takes into account druggability based on the three-dimensional structure of a given biological target using an SE(3) isovariant hybrid generation model. The virtual screening and property prediction module, based on... A two-stage funnel-shaped high-throughput virtual screening process identifies molecules with the highest overall potential from a candidate molecule library. An intelligent delivery system design module generates customized delivery solutions through a three-stage process. A targeted pathway planning module simulates the targeted delivery process of the drug-carrier complex and evaluates its efficiency in a digital twin model simulating the real in vivo tumor microenvironment. A multi-dimensional comprehensive verification module uses molecular dynamics simulations to validate the atomic-level performance of the drug, carrier, and target ternary system. The support and assurance layer includes a task queue management component, a data interaction interface component, and a closed-loop iteration engine component. The task queue management component stores and distributes tasks, the data interaction interface component enables data transmission and integration between modules, and the closed-loop iteration engine component stores and feeds back iterative optimization instructions.
Citation Information
Patent Citations
Automatic construction method and system of active polypeptide
CN120452523A
Deep learning-based drug molecule generation and screening and targeted delivery method and system
CN121122498A