Structure-based drug design model integrating Bayesian flow network and diffusion model

By integrating the Bayesian flow network with the diffusion model in the MolSolver model, and using the SE(3) equivariant graph neural network and continuous-time loss function to optimize molecule generation, the problems of molecular docking affinity and sampling efficiency in the MolCRAFT model were solved, achieving more efficient drug design.

CN120636604AActive Publication Date: 2025-09-12CHINA UNIV OF PETROLEUM (EAST CHINA)
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202511119995.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-12
Publication Date
2025-09-12
Estimated Expiration
2045-08-12

AI Technical Summary

Technical Problem

The existing MolCRAFT model has insufficient affinity for generated molecular docking and insufficient sampling efficiency, making it difficult to effectively improve the quality and efficiency of molecule generation.

Method used

MolSolver, a drug design model based on the structure-based integrated Bayesian flow network and diffusion model, updates the atomic features of molecules and proteins through SE(3) equivariant graph neural network, combines continuous time loss function and Bayesian flow network optimization parameters, and uses hybrid sampling methods to improve sampling speed and molecule generation quality.

Benefits of technology

The binding affinity, drug-likeness index and conformational stability of the generated molecules were improved, and the sampling efficiency was increased by 25%.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120636604A_ABST
    Figure CN120636604A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of bioinformatics, and particularly relates to a structure-based drug design model integrating a Bayesian flow network and a diffusion model. The drug design model integrates the advantages of a Bayesian flow network and a diffusion model through a stochastic differential equation. Atomic coordinates are sampled through a stochastic differential equation, and atomic types are sampled through Bayesian flow, so that the sampling time is shortened, and the attributes of molecules are improved. In addition, a continuous time loss function with an analytic solution is introduced, and the stability of the stochastic differential equation is improved by realizing smoother target distribution conversion. Compared with other models on a common data set CrossDocked 2020 in the field, the drug design model achieves better effects in the aspects of binding affinity of generated molecules, drug sample indexes and conformation stability, and the sampling efficiency is improved by 25%.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of bioinformatics, and in particular relates to a drug design model based on a structure-integrated Bayesian flow network and a diffusion model. Background Art

[0002] Structure-based drug design (SBDD) plays a key role in drug discovery by leveraging structural information about target protein pockets to guide the generation of three-dimensional molecules. This approach facilitates the design of molecules with tight binding affinity and favorable drug properties.

[0003] Recent advances in deep generative modeling have provided powerful tools for tackling this complex task. Early studies employed autoregressive models to generate molecules on an atom-by-atom or fragment-by-fragment basis. However, these approaches produced unnatural sequences and failed to consider the entire molecular structure. In contrast, diffusion models (DMs) have become the dominant approach in the field of SBDD, starting with a noise prior and reconstructing the complete molecule through iterative denoising. These methods are able to simultaneously model both local and global interactions between molecular atoms, achieving superior performance.

[0004] However, molecules contain both continuous atomic coordinates and discrete atom types. Although diffusion models perform well in handling continuous variables, they still face challenges in handling mixed modes, which can lead to suboptimal or even chemically invalid molecular structures. In addition, diffusion models also face the problem of a time-consuming sampling process, often requiring 1,000 denoising steps to generate 100 molecules, which may ultimately delay the discovery of viable drug candidates. Recently, the MolCRAFT model uses a Bayesian flow network (BFN) to generate molecules in parameter space, effectively capturing the multimodal joint distribution of molecules, providing a new perspective for molecule generation and accelerating sampling speed, generating 100 molecules in just 100 steps. Although the MolCRAFT model has achieved good performance, the docking affinity of the molecules it generates still has room for improvement, and the sampling efficiency also needs to be further optimized. Summary of the Invention

[0005] The present invention proposes a drug design model based on a structure-integrated Bayesian flow network and a diffusion model, which solves the problems of insufficient molecular docking affinity and insufficient sampling efficiency generated by the existing MolCRAFT model.

[0006] The technical solution of the present invention is achieved as follows: A drug design model based on a structured integrated Bayesian flow network and diffusion model is called MolSolver. The MolSolver model consists of the following parts: Feature extraction module: reads the three-dimensional coordinates and atomic type features of protein pockets and ligand molecules, and encodes and fuses the features; Feature update and molecule prediction module: The features extracted from the feature extraction module are input into the SE(3) equivariant graph neural network to predict the atomic coordinates of the molecule. and atomic types , in In the forecast, the forecast and , and train the MolSolver model with the set continuous time loss; Parameter update module: Uses the atomic coordinates and atomic types predicted in the feature update and molecule prediction modules to update parameters, uses a hybrid sampling method to update parameters, and uses a stochastic differential equation solver to update the parameters of the atomic coordinate distribution , using Bayesian flow to update the parameters of the atom type distribution , in In the update, the updated parameters and .

[0007] Through the above technical solutions, the MolSolver model fully and efficiently represents and integrates the three-dimensional coordinate features and atomic type features of molecules and proteins, allowing for the sampling of molecules in a continuous parameter space while maintaining low variance. The SE(3) equivariant graph neural network can maintain the translational and rotational invariance of molecules and proteins during feature updates. The continuous-time loss has an analytical solution and can improve molecular conformational stability. A hybrid sampling method based on an SDE solver and Bayesian flow can improve both the sampled molecular quality and the sampling speed.

[0008] Optionally, after the feature extraction module reads the features of the protein pocket and the ligand molecule, it splices the three-dimensional coordinates of the protein pocket and the ligand molecule to construct a K-nearest neighbor graph. (KNN graph), K=32, and then construct edge features through KNN graph, including edge type and edge distance, etc.; the atomic type of ligand molecule and time The structures are spliced ​​together and input into a single-layer linear fully connected network to obtain the embedded representation of atomic type and time features. Similarly, the atomic type features of the protein pocket are also input into the single-layer linear fully connected network to obtain the embedded representation of the atomic type features of the protein pocket. Finally, the embedded representations of the atomic type features of the molecule and protein are spliced ​​together for unified processing. The spliced ​​embedded representations of the atomic coordinate features, the embedded representations of the atomic type features, and the constructed edge features are used as the initial features input into the feature update and molecule prediction modules.

[0009] Through the above technical solution, the MolSolver model successfully extracted the atomic coordinate features, atomic type features of molecules and proteins, and the edge features of the constructed KNN graph, and encoded them. These features can be used for feature updating and input of the equivariant graph neural network of the molecular prediction module.

[0010] Optionally, in the feature update and molecule prediction module, the SE(3) equivariant graph neural network is used in the +1 level to the Layer Atoms Coordinate characteristics of and type traits The update method is as follows: ; ; ; in, It is +1 layer of atoms The coordinate characteristics of It's time, It is Atoms of the layer Type characteristics, It is +1 layer of atoms Type characteristics, It is Atoms of the layer Type characteristics, It is +1 layer of atoms Type characteristics, Is the atomic The coordinate change of It is Atom in the The coordinate characteristics of the layer, Representation node In the K-nearest neighbor graph The set of neighbors in and There are two attention modules; Represents atoms and The Euclidean distance between is an additional feature used to distinguish atoms and atoms An edge between them belongs to protein-protein, ligand-ligand or protein-ligand; is the ligand molecule mask to ensure that only the ligand coordinates are updated without modifying the protein coordinates; After iterative updating, the SE(3) equivariant graph neural network obtains the atomic coordinate features and type features of the current ligand molecule.

[0011] Through this technical solution, the SE(3) equivariant graph neural network can update the atomic and type features of molecules and proteins layer by layer, predicting molecules based on the current features. This graph neural network can maintain the translation and rotation invariance of coordinate features in three-dimensional space, a property that is crucial for molecule generation.

[0012] Optionally, in the feature update and molecule prediction modules, the continuous-time loss function is implemented by calculating the KL divergence between the distributions of two noisy samples; considering that both the atomic coordinates and the noise are modeled as Gaussian distributions, the loss of the atomic coordinates is Written as: ; in, represents the hyperparameter, Parameters representing the atomic coordinate distribution of the ligand molecule, represents the real atomic coordinates, represents the predicted atomic coordinates, Indicates time, Indicates protein; Likewise, for discrete atom types, the loss of their parameters It is obtained by calculating the KL divergence between two Gaussian distributions, and the analytical solution is obtained: ; in, Indicates the number of atomic types, Indicates time, ,in, Index for categories of dimensional one-hot encoding vector, Yes predictions, yes Noise scheduling when Finally, the total training objective is the sum of the atomic coordinate loss and the atomic type loss: .

[0013] Through the above technical scheme, the MolSolver model can be trained with this loss, and the continuous-time loss with analytical solution facilitates a smoother transition from the data distribution to the target distribution, which helps to improve the molecular conformational stability.

[0014] Optionally, in the parameter update module, a second-order multi-step fast SDE solver is used to update the molecular coordinate parameters , the formula is as follows: ; in, Indicates in The parameters of the coordinate distribution at sampling steps; ; ; Indicates the Second and The coordinate prediction value of the and Finite difference slope between ; is a noise schedule, and ; is standard Gaussian noise, ,and .

[0015] Optionally, in the parameter update module, the molecular coordinates are first obtained from the standard normal distribution Sampling in the middle, so Initialized as a vector ; In each iteration, the SDE solver also refers to the Second and times, the estimated coordinates are updated , and add Gaussian noise to maintain the randomness of the process; when the parameters of the coordinate distribution converge to the final value After that, it is fed into a neural network to predict the final molecular coordinates .

[0016] Through the above technical solution, the MolSolver model can efficiently update the molecular coordinate parameters through the second-order multi-step fast SDE solver, improve the properties of the sampled molecules and improve the sampling efficiency.

[0017] Optionally, in the parameter update module, the atom type predicted by the SE(3) equivariant graph neural network is given and the noise scheduling function , MolSolver model is distributed by Bayesian flow To update : ; in, Indicates the predicted atom type; express expectations; represents a normal distribution; indicates a protein pocket; Indicates time; is the atomic type characteristic of the molecule, Index for categories of dimensional one-hot encoding vector; represents a unit vector; represents the Dirichlet function; Parameters representing the updated atomic type distribution obtained through the formula; represents the noise variable; Real atom types The distribution of is modeled as a categorical distribution ,in, represents the total number of atoms in the molecule, Indicates the number of atomic types, matrix Each row of gives the probability that an atom belongs to each possible type.

[0018] Optionally, in the parameter update module, atom types are initially sampled from a uniform distribution, so that Initialized to ; When the neural network gives an atom type prediction Then, the MolSolver model first samples the noise variable from the Gaussian distribution , and then apply the softmax function to it to get ; The MolSolver model updates atomic coordinate parameters and atomic type parameters simultaneously; In the iteration, the parameters obtained in the previous step are 、 , protein pocket and time Input neural network to get molecular prediction , which contains the coordinates and type The MolSolver model then performs two simultaneous updates: refining the coordinate parameters using a second-order multi-step stochastic differential equation solver and updating the type parameters using a Bayesian flow distribution. This simultaneous update cycle is repeated for N = 80 steps. The final molecular prediction is .

[0019] Through the above technical solution, MolSolver updates the parameters of the atomic type distribution of molecules through Bayesian flow and updates the parameters of the atomic coordinate distribution of molecules through the SDE solver. This hybrid sampling method improves the properties of the sampled molecules and the sampling efficiency.

[0020] After adopting the above technical solution, the beneficial effects of the present invention are: The present invention designs a drug design model, MolSolver, for protein structures that integrates a Bayesian flow network and a diffusion model. Specifically, the MolSolver model employs a hybrid sampling approach, using an SDE solver (stochastic differential equation solver) to update atomic coordinate parameters and a Bayesian flow distribution to update atomic type parameters. Furthermore, the MolSolver model employs a continuous-time loss function with an analytical solution to ensure a smooth transition from the data distribution to the target distribution. Comprehensive experimental evaluations have shown that the drug design model not only improves binding affinity, drug-likeness indices, and conformational stability, but also increases sampling efficiency by 25%.

[0021] Although methods based on Bayesian flow networks have achieved good performance, their sampling strategies remain to be explored, which is crucial for efficient molecule generation. To solve this problem, the present invention proposes a model called MolSolver. The MolSolver model uses a Bayesian flow network as the backbone to sample in the parameter space to ensure the quality and efficiency of molecule generation. It adopts a hybrid sampling strategy, uses an SDE solver to update atomic coordinates, and uses Bayesian flow to optimize atomic types. In order to enhance the stability of the SDE solver, the MolSolver model adopts a continuous loss function with an analytical solution. This improvement contributes to a smoother transition from the data distribution to the target distribution, thereby improving sampling stability. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0023] Figure 1 It is the MolSovler model diagram; Figure 2 This is the coordinate prediction flow chart; Figure 3 Schematic diagram of the strain energy of molecules generated by different methods, the proportion of conformations in the generated molecules that have a root mean square deviation of less than 1 angstrom from their optimal conformation, and the number of atomic collisions; Figure 4 is the sampling efficiency diagram of different methods; Figure 5 To generate a molecular schematic. DETAILED DESCRIPTION

[0024] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0025] The embodiments of the present application disclose a drug design model based on a structure-integrated Bayesian flow network and a diffusion model.

[0026] Example: According to Figures 1 to 5 As shown in the figure, the drug design model based on the structure-integrated Bayesian flow network and diffusion model is as follows: 1. Task Description In structure-based drug design, the present invention considers protein binding sites and small molecules as three-dimensional point clouds with attributes: protein pockets are represented by Indicates that It's in your pocket The three-dimensional coordinates of atoms, is the corresponding atomic feature vector; the goal is to generate small molecules ,in and Small molecules The resulting small molecule must spatially fit into the pocket while exhibiting high binding affinity, good drug-like properties, and a stable binding conformation. In other words, it must not only "fit in," but also "stick firmly" and meet all druggability criteria.

[0027] 2. Dataset The present invention uses the widely used public dataset CrossDocked2020 in this field to evaluate the performance of the MolSolver model. This dataset contains a pdf file of the protein and an sdf file of the ligand molecule, which contains the 3D coordinates of the atoms and feature information. First, the PDB structure of the ligand binding pocket defined by Pocketome is calibrated; then, each ligand is comprehensively cross-docked with each pocket receptor using the smina tool; finally, the 22.5 million conformations generated are annotated using the PDBBind 2017 affinity data. After applying RMSD benchmark filtering and a 30% sequence homology threshold, 100,000 protein-ligand pairs were selected for training, and 100 proteins were retained for testing.

[0028] 3. Overall architecture of the MolSolver model The overall architecture of the MolSolver model is combined with Figure 1The model skeleton adopts the Bayesian flow model. On this basis, the advantages of the diffusion model are combined with the stochastic differential equation to achieve better model performance. The generated molecules have improved binding affinity, drug-like properties and conformational stability, and the sampling efficiency has also been improved. The MolSolver model does not operate on data, but on the parameters of the data distribution. For a given real molecule, its coordinate distribution and type distribution are characterized by certain parameters. It is assumed that the initial given molecule obeys a certain distribution, which is called the input distribution. During the training process, the loss is reduced by calculating the KL divergence of the sender distribution and the receiver distribution to allow the model to converge. The sender distribution is obtained by adding noise to the input distribution. For the receiver distribution, the input distribution is first passed through the SE(3) equivariant graph neural network to obtain the output distribution. The same noise as the input distribution is applied to the output distribution to obtain the receiver distribution. For the parameter update, the true initial molecular information, including atomic coordinates and atomic types, is predicted by the SE(3) equivariant graph neural network, and then the parameters are updated by the SDE solver and the Bayesian flow respectively, so as to gradually reach the true parameters.

[0029] 4. Continuous time loss The continuous-time loss function is implemented by calculating the KL divergence between the distributions of two noisy samples; considering that both the atomic coordinates and the noise are modeled as Gaussian distributions, the loss of the atomic coordinates is Written as: ; in, represents the hyperparameter, Parameters representing the atomic coordinate distribution of the ligand molecule, represents the real atomic coordinates, represents the predicted atomic coordinates, Indicates time, Indicates protein.

[0030] Likewise, for discrete atom types, the parameter loss It is obtained by calculating the KL divergence between two Gaussian distributions, and the analytical solution is obtained: ; in, Indicates the number of atomic types, Indicates time, ,in, Index for categories of dimensional one-hot encoding vector, Yes predictions, yes Noise scheduling when Finally, the total training objective is the sum of the atomic coordinate loss and the atomic type loss: .

[0031] 5. Feature Extraction and SE(3) Equivariant Graph Neural Network This section will combine Figure 2 For illustration. In terms of feature representation, for ligands, atom types are encoded using a one-hot vector that simultaneously captures chemical element type (C, N, O, F, P, S, Cl) and aromatic information. For protein pockets, each atom is represented by three complementary features: (1) a one-hot encoding of the atomic element (H, C, N, O, S, Se), (2) a 20-dimensional one-hot vector indicating the type of amino acid residue to which it belongs, and (3) a binary flag indicating whether the atom belongs to the protein backbone. Molecular and protein atoms are accompanied by 3D spatial coordinates and contain a binary indicator to clearly distinguish between molecule and protein atoms. These coordinates are then concatenated to construct a K-nearest neighbor graph, where K=32, and edge types (ligand to ligand, ligand to protein, protein to ligand, protein to protein) and other edge features are constructed accordingly. RDkit is used to process ligand and protein files for feature extraction.

[0032] The molecular type features are concatenated with time and passed through a single-layer linear neural network to obtain a fused embedding representation. Protein type features are also passed through a single-layer linear neural network to obtain an embedding representation. The two sets of embeddings are then merged and input into the SE(3) equivariant graph neural network.

[0033] In order to fully capture the ligand-pocket interaction, the network backbone uses PosNet3D with SE(3) equivariance, which effectively utilizes the molecular geometry information while maintaining the invariance to the overall rotation and translation. The SE(3) equivariant graph neural network consists of nine layers, each containing two attention blocks. The atomic type features and coordinates are updated in the layers, where The characteristics of the atoms are given by Layer type characteristics and coordinates To update: ; ; ; in, It is +1 layer of atoms The coordinate characteristics of It is +1 layer of atomic type characteristics, It’s time; It is Atoms of the layer Type characteristics, It is +1 layer of atoms Type characteristics, It is Atoms of the layer Type characteristics, It is +1 layer of atoms Type characteristics of Is the atomic The amount of coordinate change; It is Atom in the Coordinate characteristics of the layer; Representation node In the K-nearest neighbor graph The set of neighbors in and There are two attention modules; Represents atoms and The Euclidean distance between is an additional feature used to distinguish atoms and atoms An edge between them belongs to protein-protein, ligand-ligand or protein-ligand; is the ligand molecule mask that ensures that only the ligand coordinates are updated and not the protein coordinates.

[0034] Ultimately, the prediction network outputs predictions of atomic coordinates and atomic types at each iteration step; in the post-processing stage, the MolSolver model uses OpenBabel to add chemical bonds between molecular atoms.

[0035] 6. Update of atomic coordinate parameters To incorporate the advantages of diffusion models (DMs) into Bayesian flow networks (BFNs), this paper models atomic coordinate parameter updates as stochastic differential equations (SDEs) and utilizes an SDE solver for sampling to improve the performance of existing BFN-based approaches. This is the first application of this strategy in structure-based drug design (SBDD).

[0036] For the molecular coordinates ( is the number of atoms in the molecule), the present invention regards it as obeying Gaussian distribution ,in, is the mean, is the variance, is the precision, is a unit vector. So the parameter set of the coordinate distribution is ; Due to the accuracy By noise factor Pre-set and meet , It is Step accuracy value, It is Step accuracy value, It is Noise factor of the step Therefore, only the mean parameter is used during training. to update.

[0037] The original update strategy is derived from the Bayesian update function. Studies have shown that this update is actually equivalent to the first-order discretization of a specific inverse stochastic differential equation in .

[0038] ; in, Represents noise, satisfying Therefore, we can The updating process of is considered as a process following the reverse time stochastic differential equation.

[0039] in, Indicates in The parameters of the coordinate distribution at sampling steps; are the predicted atomic coordinates; is Noise scheduling of time, It is in step time, and , is The noise scheduling of time, and ; is a hyperparameter; is standard Gaussian noise.

[0040] In order to solve the equation while ensuring sampling efficiency and stability, the present invention introduces a fast second-order multi-step SDE solver whose sampling formula is shown below.

[0041] ; in, Indicates in The parameters of the coordinate distribution at sampling steps; ; ; Indicates the Second and The coordinate prediction value of the and Finite difference slope between ; is a noise schedule, and ; is standard Gaussian noise, ,and ; yes of Power, yes of Power.

[0042] Specifically, the molecular coordinates are first obtained from the standard normal distribution Sampling in the middle, so Initialized as a vector ; In each iteration, the SDE solver refers to both the current step and the previous step ( Second and times) to update the estimated coordinates , and add Gaussian noise to maintain the randomness of the process; when the parameters of the coordinate distribution converge to the final value After that, it is fed into a neural network to predict the final molecular coordinates .

[0043] 7. Atom type parameter update Discrete atomic types The distribution is modeled as a categorical distribution ,in, represents the total number of atoms in the molecule, Indicates the number of atomic types. Each row gives the probability of an atom belonging to each possible type. Given the atom type predicted by the neural network and the noise scheduling function , MolSolver model is distributed by Bayesian flow To update : ; in, Indicates the predicted atom type; express expectations; represents a normal distribution; indicates a protein pocket; Indicates time; is the atomic type characteristic of the molecule, Index for categories of dimensional one-hot encoding vector; represents a unit vector; represents the Dirichlet function; Parameters representing the updated atomic type distribution obtained through the formula; represents the noise variable.

[0044] Specifically, atom types are initially sampled from a uniform distribution, so the initial atom type distribution parameter Initialized to When the neural network predicts the atom type Then, the MolSolver model first samples the noise variable from the Gaussian distribution , and then apply the softmax function to it to get .

[0045] The MolSolver model updates atomic coordinate parameters and atomic type parameters simultaneously. In the first iteration, the present invention will The parameters of the coordinate distribution obtained in step , parameters of type distribution , protein pocket and time Input the neural network to get the predicted molecules , which contains the atomic coordinates and atomic types The MolSolver model then performs two updates simultaneously: refining the coordinate parameters using a second-order multi-step stochastic differential equation solver and updating the type parameters using a Bayesian flow distribution. This synchronous update cycle is repeated for N = 80 steps. The final molecule (including atomic coordinates and atomic types) is , represents the final atomic coordinates, Indicates the final atom type.

[0046] 8. Performance comparison of MolSolver model with other prediction models In order to verify the performance of the MolSolver model, according to the convention in the field, the present invention was tested on the CrossDocked2020 dataset. According to the same division method as other models, the present invention selected a subset of 100,000 protein-ligand pairs for training, and retained 100 proteins for testing. The present invention compared and analyzed the MolSolver model with six recent representative models and test set molecules. AR, Pocket2Mol and FLAG are autoregressive models. The first two sample molecules in an atom-by-atom manner, while FLAG generates molecules in a fragment-by-fragment manner; TargetDiff and DecompDiff are currently advanced models based on diffusion methods. MolCRAFT is a currently advanced model based on Bayesian flow networks.

[0047] The present invention thoroughly evaluates these models based on binding affinity, molecular properties, conformational stability, and sampling efficiency. Binding affinity includes the docking score, docking score minimization, and redocking score. The docking score directly estimates binding affinity based on the generated three-dimensional molecule; the docking score minimization is performed after local energy minimization and then evaluated; the redocking score involves the redocking process and reflects the optimal binding affinity. Molecular properties include three metrics: synthetic accessibility, drug-likeness, and diversity. Synthetic accessibility indicates the difficulty of synthesizing the ligand; drug-likeness is a quantitative estimate of drug similarity; and diversity is assessed by calculating the average pairwise differences between generated molecules within each binding pocket. The present invention also includes a success rate metric, which is the proportion of molecules with high affinity and reasonable molecular properties (redocking score < -8.18, drug-likeness > 0.25, synthetic accessibility > 0.59). Conformational stability includes strain energy, the number of atomic collisions, and the proportion of generated molecules with the optimal conformation. Sampling efficiency includes sampling speed and generation success rate. The generation success rate refers to the proportion of molecules that meet both chemical efficacy and structural integrity requirements.

[0048] As shown in Table 1, the MolSolver model outperformed other strong baseline models in nearly all affinity metrics. For example, in docking score minimization, the MolSolver model achieved an average score of –7.42 and a median score of –7.52, representing improvements of 2.1% and 3.6%, respectively, over the strongest baseline, the MolCRAFT model. In the de novo docking score metric, the MolSolver model achieved an average score of –8.21 and a median score of –8.39, exceeding the MolCRAFT model by 3.7% and 4.7%, respectively. These improvements demonstrate the MolSolver model's superior ability to explore high-affinity candidate molecules and tailor protein structure to achieve tight binding. In terms of molecular properties such as synthetic accessibility and diversity, the MolSolver model performed competitively with other methods, achieving a synthetic accessibility score of 0.66, second only to the MolCRAFT model among non-autoregressive methods. It is important to note that the higher synthetic accessibility values ​​of the Pocket2Mol and FLAG models are partly due to the smaller size of the molecules they generate. For the drug-likeness metric, the MolSolver model achieved the highest score of 0.58, outperforming all autoregressive and non-autoregressive methods. This demonstrates that the MolSolver model can effectively generate molecules with favorable drug-like properties. The MolSolver model achieved an overall success rate of 27.8%, the highest among all benchmark methods, demonstrating its exceptional ability to generate synthesizable molecules while maintaining high affinity. Notably, with the exception of synthetic accessibility, the MolSolver model surpassed the reference molecules in all metrics, further highlighting its effectiveness in capturing molecule-protein interactions and generating compounds with excellent drug-like properties.

[0049] Table 1 Overall results of reference molecules, molecules generated by various models, and test set molecules Model and test set Docking score Docking score Docking score minimization Docking score minimization Redocking score Redocking score Synthetic accessibility Drug-like properties Diversity Success rate Average median Average median Average median Average Average Average Test set -6.36 -6.46 -6.71 -6.49 -7.45 -7.26 0.73 0.48 — 25.0 AR -5.75 -5.64 -6.18 -5.88 -6.75 -6.62 0.64 0.51 0.70 7.1 Pocket2Mol -5.14 -4.70 -6.42 -5.82 -7.15 -6.79 0.76 0.57 0.69 24.4 FLAG 16.48 4.53 1.21 -4.04 -5.63 -6.61 0.70 0.49 0.70 14.1 TargetDiff -5.47 -6.30 -6.64 -6.83 -7.80 -7.91 0.58 0.48 0.72 10.5 DecompDiff -5.19 -5.27 -6.03 -6.00 -7.03 -7.16 0.66 0.51 0.73 14.9 MolCRAFT -6.59 -7.04 -7.27 -7.26 -7.92 -8.01 0.69 0.50 0.72 26.8 MolSolver -6.45 -7.18 -7.42 -7.52 -8.21 -8.39 0.66 0.58 0.71 27.8 like Figure 3 As shown in the figure, the MolSolver model performs consistently with the strongest baseline MolCRAFT model in terms of strain energy, with a median of 208 for both. In terms of the proportion of optimal conformations and the number of atomic collisions, the MolSolver model achieves 42.6% and 6.83, respectively, both of which are the best values ​​among all models, highlighting the MolSolver model's ability to generate structurally stable conformations.

[0050] Figure 4A scatter plot shows the relationship between the generation success rate (horizontal axis) and the corresponding sampling efficiency metric for four models based on diffusion models (DM) and Bayesian flow networks: the TargetDiff model, the DecompDiff model, the MolCRAFT model, and the MolSolver model of the present invention. The generation success rate refers to the proportion of generated samples that meet both molecular chemical validity and structural integrity. Specifically, the MolSolver model achieved a generation success rate of 98.89%, ranking first among the four models, surpassing the MolCRAFT model's 96.67%, the TargetDiff model's 90.36%, and the DecompDiff model's 82.92%. In terms of sampling speed, the MolSolver model also boasts the fastest sampling speed, approximately 1.69 (generating 100 molecules in 59 seconds), significantly exceeding the TargetDiff model's 0.127, the DecompDiff model's 0.091, and the MolCRAFT model's 1.35. Compared to the best baseline model, MolCRAFT, the MolSolver model achieved a sampling speed improvement of (1.69 -1.35) / 1.35≈25%. These results demonstrate that the MolSolver model combines high efficiency with excellent performance in molecular generation tasks.

[0051] The present invention randomly selected two proteins, 3daf and 14gs, and Figure 5 The figure shows a reference molecule, a molecule generated by the MolCRAFT model, and a molecule generated by the MolSolver model. The generated molecules were prioritized by docking score, synthetic accessibility, and drug-likeness, and the top two molecules were selected for visualization. The results showed that the molecules generated by the MolSolver model performed better.

[0052] This paper proposes a new method, which is the first method in the field of structure-based drug design to combine Bayesian flow networks and diffusion models in the form of stochastic differential equations. The MolSolver model adopts a hybrid sampling strategy: atomic coordinates are updated using a high-order SDE solver, and atom types are updated using a Bayesian flow distribution. In addition, the MolSolver model introduces a continuous loss function with a closed-form solution, which can achieve a smoother transition between the data distribution and the target distribution, thereby alleviating the common instability problems of stochastic differential equations. Experiments show that the MolSolver model not only improves binding affinity, drug-likeness indicators and conformational stability, but also improves sampling efficiency by 25%.

[0053] The above are only preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the technical solutions of the present invention should be included in the protection scope of the present invention.

Claims

1. A drug design model based on the structure-integrated Bayesian flow network and diffusion model, characterized by: The name of the model is MolSolver. The MolSolver model consists of the following parts: Feature extraction module: reads the three-dimensional coordinates and atomic type features of protein pockets and ligand molecules, and encodes and fuses the features; Feature update and molecule prediction module: The features extracted from the feature extraction module are input into the SE(3) equivariant graph neural network to predict the atomic coordinates of the molecule. and atomic types , in In the forecast, the forecast and , and train the MolSolver model with the set continuous time loss; Parameter update module: Uses the atomic coordinates and atomic types predicted in the feature update and molecule prediction modules to update parameters, and uses a stochastic differential equation solver to update the parameters of the atomic coordinate distribution , using Bayesian flow to update the parameters of the atom type distribution , in In the update, the updated parameters and .

2. The drug design model based on structure-integrated Bayesian flow network and diffusion model according to claim 1, characterized in that: In the feature extraction module, after reading the features of the protein pocket and the ligand molecule, the three-dimensional coordinates of the protein pocket and the ligand molecule are spliced ​​together to construct a KNN graph with K=32, and edge features are constructed through the KNN graph; The atomic type of the ligand molecule and the time They are spliced ​​together and input into a single-layer linear fully connected network to obtain the embedded representation of atom type and time features; the atom type features of the protein pocket are also input into the single-layer linear fully connected network to obtain the embedded representation of the atom type features of the protein pocket; finally, the embedded representations of the atom type features of the molecule and protein are spliced ​​together, and the spliced ​​embedded representations of the atomic coordinate features, the embedded representations of the atom type features, and the constructed edge features are used as the initial features input into the feature update and molecule prediction module.

3. The drug design model based on structure-integrated Bayesian flow network and diffusion model according to claim 1, characterized in that: In the feature update and molecule prediction modules, the SE(3) equivariant graph neural network is used in the first +1 level to the Layer Atoms Coordinate characteristics of and type traits The update method is as follows: ; ; ; in, It is +1 layer of atoms The coordinate characteristics of It's time, It is Atoms of the layer Type characteristics, It is +1 layer of atoms Type characteristics, It is Atoms of the layer Type characteristics, It is +1 layer of atoms Type characteristics, Is the atomic The coordinate change of It is Atom in the The coordinate characteristics of the layer, Representation node In the K-nearest neighbor graph The set of neighbors in and There are two attention modules; Represents atoms and The Euclidean distance between is an additional feature used to distinguish atoms and atoms An edge between them belongs to protein-protein, ligand-ligand or protein-ligand; is the ligand molecule mask to ensure that only the ligand coordinates are updated without modifying the protein coordinates; After iterative updating, the SE(3) equivariant graph neural network obtains the atomic coordinate features and type features of the current ligand molecule.

4. The drug design model based on structure-integrated Bayesian flow network and diffusion model according to claim 1, characterized in that: In the feature update and molecule prediction modules, the continuous-time loss function is implemented by calculating the KL divergence between the distributions of two noisy samples; considering that both the atomic coordinates and the noise are modeled as Gaussian distributions, the loss of the atomic coordinates Written as: ; in, represents the hyperparameter, Parameters representing the atomic coordinate distribution of the ligand molecule, represents the real atomic coordinates, represents the predicted atomic coordinates, Indicates time, Indicates protein; Likewise, for discrete atom types, the loss of their parameters It is obtained by calculating the KL divergence between two Gaussian distributions, and the analytical solution is obtained: ; in, Indicates the number of atomic types, Indicates time, ,in, Index for categories of dimensional one-hot encoding vector, Yes predictions, yes Noise scheduling when Finally, the total training objective is the sum of the atomic coordinate loss and the atomic type loss: 。 5. The drug design model based on structure-integrated Bayesian flow network and diffusion model according to claim 1, characterized in that: In the parameter update module, a second-order multi-step fast SDE solver is used to update the molecular coordinate parameters , the formula is as follows: ; in, Indicates in The parameters of the coordinate distribution at sampling steps; ; ; Indicates the Second and The coordinate prediction value of the and Finite difference slope between ; is a noise schedule, and ; is standard Gaussian noise, ,and .

6. The drug design model based on structure-integrated Bayesian flow network and diffusion model according to claim 5, characterized in that: In the parameter update module, the molecular coordinates are first obtained from the standard normal distribution Sampling in the middle, so Initialized as a vector ; In each iteration, the SDE solver also refers to the Second and times, the estimated coordinates are updated , and add Gaussian noise to maintain the randomness of the process; when the parameters of the coordinate distribution converge to the final value After that, it is fed into a neural network to predict the final molecular coordinates .

7. The drug design model based on structure-integrated Bayesian flow network and diffusion model according to claim 1, characterized in that: In the parameter update module, the atom type predicted by the SE(3) equivariant graph neural network is given and the noise scheduling function , MolSolver model is distributed by Bayesian flow To update : ; in, Indicates the predicted atom type; express expectations; represents a normal distribution; indicates a protein pocket; Indicates time; is the atomic type characteristic of the molecule, Index for categories of dimensional one-hot encoding vector; represents a unit vector; represents the Dirichlet function; Parameters representing the updated atomic type distribution obtained by the formula; represents the noise variable; Real atom types The distribution of is modeled as a categorical distribution ,in, represents the total number of atoms in the molecule, Indicates the number of atomic types, matrix Each row of gives the probability that an atom belongs to each possible type.

8. The drug design model based on structure-integrated Bayesian flow network and diffusion model according to claim 7, characterized in that: In the parameter update module, atom types are initially sampled from a uniform distribution, so Initialized to ; When the neural network gives an atom type prediction Then, the MolSolver model first samples the noise variable from the Gaussian distribution , and then apply the softmax function to it to get ; The MolSolver model updates atomic coordinate parameters and atomic type parameters simultaneously; In the iteration, the parameters obtained in the previous step are 、 , protein pocket and time Input neural network to get molecular prediction , which contains the coordinates and type The MolSolver model then performs two simultaneous updates: refining the coordinate parameters using a second-order multi-step stochastic differential equation solver and updating the type parameters using a Bayesian flow distribution. This simultaneous update cycle is repeated for N = 80 steps. The final molecular prediction is .

Citation Information

Patent Citations

  • Intelligent molecule generation method and device based on diffusion model, equipment and medium

    CN116665807A

  • Molecular generation method and device based on Bayesian flow network

    CN118471369A

  • Three-dimensional drug molecule generation method and system based on Bayesian flow network

    CN120220879A

  • Method and apparatus for analysis of molecular configurations and combinations

    US20050119837A1

  • Diffusion model for generative protein design

    US20240161864A1