Three-dimensional drug molecule generation method and system based on Bayesian flow network
By combining Bayesian stream network and diffusion model, using Bayesian inference and dynamic noise schedule, the problem of existing diffusion models controlling discrete data and noise sensitivity when generating drug molecules is solved, achieving more efficient molecular generation and improving the rationality and effectiveness of the generated results.
Patent Information
- Application Number
- CN202510196691.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-06-27
AI Technical Summary
When generating drug molecules, it is difficult to effectively control discrete data, resulting in problems with the generated molecules in terms of geometric and chemical rationality, and is sensitive to noise, making it easy to generate molecular structures that are not in line with reality.
A three-dimensional drug molecule generation method based on Bayesian flow network is adopted, and by combining Bayesian flow network and diffusion model, Bayesian inference is used to update prior statistical knowledge, dynamically adjust the noise schedule, and improve the control ability of the generation model and the rationality of the generated results.
The model's modeling ability of the multimodal features of molecules is effectively improved, and the generated drug molecules have been improved in terms of geometric and chemical rationality, reducing noise sensitivity, and improving the effectiveness of the generated results.
Smart Images

Figure CN120220879A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of artificial intelligence and drug design, and more specifically, to a three-dimensional drug molecule generation method and system based on a Bayesian flow network. Background Art
[0002] In recent years, with the rapid development of deep learning technology, drug molecule design based on generative models has become a research hotspot in the field of computer-aided drug design. Existing molecular generative models are mainly divided into the following categories: those based on graph neural networks (GNNs), those based on autoregressive models (RNNs), variational autoencoders (VAEs), generative adversarial networks (GANs), and diffusion models. Among them, the molecular generation method based on diffusion models is relatively common. Diffusion models are a type of generative model that has emerged in recent years, which introduce noise into data step by step and learn the denoising process. When sampling to generate new molecules, diffusion models generate the three-dimensional structure of molecules with complex distribution characteristics by simulating the process of gradually denoising. In the field of molecular generation, diffusion models have been used to generate the three-dimensional structure of molecules, and by learning the distribution changes of atoms in the molecular space, they show strong flexibility and accuracy in generating the three-dimensional structure of molecules, and partly solve the problem that traditional generative models cannot effectively capture the three-dimensional information of molecules. However, the standard diffusion model has limited ability to control the generation of discrete data in molecules (such as the types of atoms, the types of chemical bonds, etc.), especially when precise optimization of molecular properties is required.
[0003] Existing diffusion model methods, such as EDM and GeoLDM, etc., have shown certain advantages in the task of molecular three-dimensional generation. However, since existing diffusion models usually forcibly convert discrete variables into continuous one-hot vectors and then perform the noise addition and denoising process based on continuous Gaussian noise, this unreasonable processing method will lead to suboptimal solutions of the model, greatly hindering the application value of the molecular generation model in actual production; diffusion models usually ignore the physicochemical constraints of molecules (such as the rationality of bond lengths, bond angles, and dihedral angles) during the generation process, which may lead to problems in the geometric and chemical rationality of the generated molecules; existing models have limited ability to integrate chemical prior knowledge and experimental data, and basically use end-to-end models to directly generate the final 3D drug molecules, resulting in the generated molecules may lack practical medicinal value; existing diffusion models are very sensitive to the noise intensity introduced during model training, and are easily mislead by the noise in the data and then generate molecular structures that are completely inconsistent with reality. Summary of the Invention
[0004] To overcome the deficiencies in the modeling ability of molecular multimodal features and the lack of geometric and chemical rationality in generating molecules in the above-mentioned prior art, the present invention provides a three-dimensional drug molecule generation method and system based on a Bayesian flow network.
[0005] To solve the above technical problems, the technical solution of the present invention is as follows:
[0006] A three-dimensional drug molecule generation method based on a Bayesian flow network, comprising the following steps:
[0007] Input real initial molecular data into a drug molecule generation model that combines a Bayesian flow network and a diffusion model; wherein:
[0008] Generate noise molecular features according to prior statistical knowledge;
[0009] Combining the noise molecular features and the random Gaussian noise obtained from the noise schedule, introduce noise into the current molecular data to generate noisy molecular data; at the same time, update the data distribution in the prior statistical knowledge based on Bayesian inference;
[0010] Denoise the noisy molecular data to generate predicted drug molecules;
[0011] Repeat the above steps until the molecular structure converges or reaches the preset number of steps T to obtain the drug molecule generation result.
[0012] Furthermore, the present invention also proposes a three-dimensional drug molecule generation system based on a Bayesian flow network, applying the three-dimensional drug molecule generation method proposed by the present invention. Among them, the system includes:
[0013] A prior module, which is equipped with prior statistical knowledge and is used to generate noise molecular features;
[0014] A noise addition module, which is equipped with a noise schedule, and is used to combine the noise molecular features and the random Gaussian noise obtained from the noise schedule, introduce noise into the input molecular data to generate noisy molecular data; at the same time, update the data distribution in the prior statistical knowledge based on Bayesian inference;
[0015] A prediction module, which is used to denoise the noisy molecular data to generate predicted drug molecules; repeat the noise addition and prediction until the molecular structure converges or reaches the preset number of steps T, and output the drug molecule generation result.
[0016] Compared with the prior art, the beneficial effects of the technical solution of the present invention are:
[0017] The present invention constructs a drug molecule generation model by combining a Bayesian flow network and a diffusion model. The Bayesian flow network is used to sample continuous atomic coordinates, discrete atomic types, and chemical bond types, and combined with Bayesian inference to model and learn the true molecular data distribution in a continuously differentiable parameter space. The modeling process of discrete features is converted into a parameter distribution that models the data distribution, resulting in a continuous unimodal data distribution. Thus, a unified modeling framework can be used to guide the training of the generation model, effectively improving the model's ability to model the multimodal features of molecules.
[0018] The present invention introduces relevant prior statistical knowledge and updates the prior in cooperation with Bayesian inference to make it increasingly approximate the true data distribution. A noise schedule is introduced to dynamically adjust the generation ability of the model, effectively improving the rationality and effectiveness of the generated drug molecules. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 It is a flowchart of a three-dimensional drug molecule generation method based on a Bayesian flow network according to an embodiment of the present invention.
[0020] Figure 2 It is a schematic diagram of the training and sampling process of a drug molecule generation model according to an embodiment of the present invention.
[0021] Figure 3 It is a curve graph of weights and a noise schedule according to an embodiment of the present invention.
[0022] Figure 4 It is a schematic diagram of the training update of a drug molecule generation model according to an embodiment of the present invention.
[0023] Figure 5 It is a schematic diagram of the drug molecule generation result according to an embodiment of the present invention.
[0024] Figure 6 It is an architecture diagram of a three-dimensional drug molecule generation system based on a Bayesian flow network according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0025] Here, exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present invention. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present invention as detailed in the appended claims.
[0026] The terms used in this invention are for the purpose of describing specific embodiments only and are not intended to limit the invention. The singular forms "a", "the", and "said" used in this invention and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise. It should also be understood that the term "and / or" used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.
[0027] It should be understood that although the terms first, second, third, etc. may be used in this invention to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of this invention, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to a determination".
[0028] The following describes the present invention in detail in conjunction with the accompanying drawings and specific embodiments.
[0029] Embodiment 1
[0030] This embodiment proposes a three-dimensional drug molecule generation method based on a Bayesian flow network, as Figure 1 shown, which is a flowchart of the three-dimensional drug molecule generation method of this embodiment.
[0031] The three-dimensional drug molecule generation method based on a Bayesian flow network proposed in this embodiment includes the following steps:
[0032] Input the real initial molecular data into a drug molecule generation model that combines a Bayesian flow network and a diffusion model; where:
[0033] Generate noise molecular features according to prior statistical knowledge;
[0034] Combine the noise molecular features and the random Gaussian noise obtained from the noise schedule to introduce noise into the current molecular data to generate noisy molecular data; at the same time, update the data distribution in the prior statistical knowledge based on Bayesian inference;
[0035] Denoise the noisy molecular data to generate predicted drug molecules;
[0036] Repeat the above steps until the molecular structure converges or reaches a preset number of steps T to obtain the drug molecule generation result.
[0037] In this embodiment, the drug molecule generation model combines a Bayesian flow network and a diffusion model. The Bayesian flow network is used to sample continuous atomic coordinates, discrete atomic types, and chemical bond types, and combined with Bayesian inference to model and learn the real molecular data distribution in a continuously differentiable parameter space, converting "modeling the distribution of multimodal molecular data composed of continuous and discrete features" into "modeling the parameter distribution of the data distribution", which can effectively improve the model's ability to model molecular multimodal features.
[0038] Furthermore, in this embodiment, relevant chemical prior knowledge is artificially introduced, such as the distribution of each atomic type in real molecules in nature, etc., and Bayesian inference is introduced. The posterior of the data distribution is continuously calculated using the observed data samples, thereby updating the prior to make it closer and closer to the real data distribution.
[0039] This embodiment also designs a noise schedule, which can be optionally designed according to different types of molecular features, such as atomic three-dimensional coordinates, atomic types, chemical bond types, etc., to cooperate with the model to have different emphases on various molecular features during the generation process at each time step, and dynamically adjust the generation ability of the model. For example, at the initial stage of generation, more emphasis is placed on modeling the chaotic and disordered atomic three-dimensional coordinates. In the middle stage, it begins to try to confirm the corresponding element types for each atom one by one. In the later stage of generation, the chemical bonds and distances between atoms (i.e., chemical bond lengths) are considered, thereby alleviating the noise sensitivity problem of the model and effectively improving the effectiveness of the generated drug molecules.
[0040] In an alternative embodiment, the prior statistical knowledge includes atomic type distribution, chemical bond distribution, and atomic number distribution.
[0041] Among them, the prior statistical knowledge can be optionally obtained by statistically analyzing a real drug molecule dataset.
[0042] In an alternative embodiment, generating noise molecular features according to the prior statistical knowledge includes the following steps:
[0043] Input the parameters of the prior statistical knowledge into the Bayesian flow network to obtain the parameters of the estimated distribution of the original data as the noise molecular features.
[0044] In this embodiment, the Bayesian flow network is used to transform the prior distribution into a parameter distribution of a single-modal continuous prior distribution to overcome the defect of insufficient ability to model molecular multimodal features.
[0045] In an alternative embodiment, generating the predicted drug molecule includes the following steps:
[0046] According to the current time step t, extract the corresponding random Gaussian noise β(t) from the noise schedule;
[0047] Combine the random Gaussian noise β(t) at the current time step t with the noise molecular features and add noise to the current molecular data to obtain noisy molecular data;
[0048] Denoise the noisy molecular data, combine the Bayesian flow framework, and predict the three-dimensional atomic coordinates, atomic types, and chemical bond types to obtain the predicted drug molecule at the current time step t.
[0049] In this embodiment, the Bayesian flow network and the diffusion model are combined, and Bayesian inference is used to quantify and update the uncertainty of the model parameters and the generation results. At the same time, the corresponding random Gaussian noise is extracted from the noise schedule to guide the generation direction of the model.
[0050] When predicting the three-dimensional atomic coordinates, atomic types, and chemical bond types, optionally apply the random noise on the three-dimensional atomic coordinates, atomic types, and chemical bond types extracted from the corresponding noise schedule, and cooperate with Bayesian inference to model the dynamic interaction between the continuous atomic coordinate features, discrete atomic type features, and chemical bond features, effectively capture the complex three-dimensional structure of the molecule, and also quantify the uncertainty of the generation results.
[0051] Further, in an alternative embodiment, predicting the three-dimensional atomic coordinates includes:
[0052] Predict the noise β x (t) at the current time step t from the noise schedule corresponding to the three-dimensional atomic coordinates, and restore the three-dimensional atomic coordinates in combination with the Bayesian flow network according to the three-dimensional atomic coordinates x in the initial molecular data; its expression is:
[0053]
[0054] where the random Gaussian noise corresponding to the three-dimensional atomic coordinates is represents the intensity of the random Gaussian noise obtained according to the noise schedule at the sampling step t, and t ∈ [0, 1]; θ x represents the predicted three-dimensional atomic coordinates; represents the three-dimensional atomic coordinates predicted at the 0th time step, which is actually equivalent to completely random Gaussian noise; x represents the original true three-dimensional atomic coordinates; represents following a Gaussian distribution; μ i represents the sampling result at the ith time step, and I is the identity matrix.
[0055] For the three-dimensional atomic coordinates, since the characteristic distribution of the three-dimensional atomic coordinates is continuous, only the random Gaussian noise needs to be applied to the three-dimensional atomic coordinates during the noise addition process, and the sampling target is to restore the original three-dimensional atomic coordinates from the Gaussian distribution.
[0056] In this embodiment, according to the original and real three-dimensional atomic coordinates, the random noise intensity β x (t) and the result predicted at the 0th step predict the three-dimensional atomic coordinates at the i = t / T step, where the three-dimensional atomic coordinates at the i-th step follow a Gaussian distribution with a mean of γ(t)x and a variance of γ(t)(1 - γ(t))I That is, the goal of the model is to learn the sampling formula for the three-dimensional atomic coordinates related to the Gaussian distribution.
[0057] Furthermore, in an alternative embodiment, predicting the atomic type and the chemical bond type includes:
[0058] According to the classification distribution of the atomic type and the chemical bond type in the initial molecular data, reparameterize the parameters of the classification distribution, transform them into Gaussian distributions of the atomic type and the chemical bond type, and use them as the corresponding noise schedules;
[0059] Predict the Gaussian noise β h (t) and β e (t) at the current time step t from the corresponding noise schedules, and according to the atomic type and the chemical bond type in the initial molecular data, restore the atomic type and the chemical bond type by combining the Bayesian flow network, and transform the sampling result to the probability simplex space of the classification distribution through softmax; its expression is:
[0060]
[0061] where θ h represents the predicted atomic type, h represents the original and real atomic type; β h (t) is the noise schedule of the atomic type, and β h (t) = t 2 β h (1); 1 h represents the one-hot vector corresponding to the real atomic type, 1 is a vector of all 1s; y h is the result sampled from the Gaussian distribution , K h is the number of types of atomic types. In common drug molecules, the number of types of atomic types is 5 - 14, and the specific value can be optionally set according to the specific task; δ(θ h - softmax(y h )) is the Dirichlet distribution function, which is used to turn single-point data into a distribution; θ e represents the predicted chemical bond type, e represents the original and real chemical bond type; β e(t) is a noise schedule for chemical bond types, and β e (t) = t r β e (1), where r is a positive integer greater than 2; 1 e represents the one - hot vector corresponding to the true chemical bond type; y e is the result sampled from the Gaussian distribution and K e is the chemical bond type of the chemical bond type.
[0062] It should be noted that for the noise schedule β e () of the chemical bond type, the exponent r is usually set as a positive integer greater than 2 because the modeling of chemical bonds needs to be based on a more stable atomic type topological structure. Otherwise, it is meaningless to predict the chemical bonds between two atoms according to the chaotic atomic topological structure.
[0063] For discrete atomic types and chemical bond types, in this embodiment, the parameters of the classification distribution to which they belong are modeled, so as to unify the modeling process with the three - dimensional coordinates of atoms into a continuously differentiable parameter space, forming a single - mode distribution. Thus, a unified modeling framework can be used to guide the training of the generation model, effectively improving the effectiveness of the drug molecule generation results.
[0064] In an optional embodiment, the method further includes the following steps:
[0065] Obtain a real molecule dataset for training the drug molecule generation model; where:
[0066] Extract random Gaussian noise from the noise schedule according to the random sampling time, and combine it with the real molecule data to generate noisy molecule data;
[0067] Input the noisy molecule data into the drug molecule generation model to obtain a predicted molecule;
[0068] Based on the real molecule data and the predicted molecule, calculate a loss function containing physicochemical constraints, and train the drug molecule generation model based on the early stopping method.
[0069] In this embodiment, a real dataset is used to train the drug molecule generation model to improve the generation effect of the model. The loss function used in the training process contains physicochemical constraints. Optionally, it includes constraints on the rotation and / or translation invariance of the three - dimensional coordinates of the molecule, and constraints on the bond lengths of the generated chemical bonds to conform to reality, etc., thereby improving the geometric and chemical rationality of the drug molecule generation results.
[0070] Exemplarily, such as Figure 2As shown in the figure, it is a schematic diagram of the training and sampling process of the drug molecule generation model of this embodiment. It includes a training part and a sampling part. In the specific implementation process, first, the drug molecule generation model is trained, and the model is optimized with the loss function until the loss value converges or reaches the preset number of iteration steps, and a pre-trained drug molecule generation model is obtained. Further, in the sampling process, the pre-trained drug molecule generation model is applied, and noisy molecular data is generated in cooperation with prior statistical knowledge and a noise schedule. After prediction through the drug molecule generation model and cycling T times, the drug molecule generation result is output.
[0071] Exemplarily, the preset number of steps T is taken as 1000.
[0072] Further, in an alternative embodiment, the physicochemical constraints in the loss function include atomic valence loss constraints and chemical bond length loss constraints; their expressions are:
[0073]
[0074] Among them, L ∞ (g) represents the loss function including physicochemical constraints, represents the reconstruction loss of the atomic three-dimensional coordinate x, represents the reconstruction loss of the atomic type h, represents the reconstruction loss of the chemical bond type e, represents the chemical bond length loss; the superscript ∞ represents the loss of the infinite-step "decomposition-reconstruction" step, and its actual operation is to find the expected loss value of the unified modeling sampling formula for different modal features of the molecule; t~U(0,1) represents randomly uniformly sampling the time step t from 0 to 1 each time during training, and U(0,1) represents following the random uniform distribution between 0 and 1; p F (θ g |g;t) represents the sampling results of the atomic three-dimensional coordinate, atomic type, and chemical bond type, g represents the predicted molecular graph including the atomic three-dimensional coordinate, atomic type, and chemical bond type; α x (t), α h (t), α e (t) represent the noise schedules of the atomic three-dimensional coordinate, atomic type, and chemical bond type; ||x - Ψ x || 2 represents the mean square error between the atomic three-dimensional coordinate Ψ x generated by the model and the current true atomic three-dimensional coordinate x, ||h - Ψ h || 2 represents the mean square error between the atomic type Ψ h generated by the model and the current true atomic type h, ||e - Ψ e || 2Represents the chemical bond type Ψ generated by the model e The mean squared error between the current true chemical bond type e; λ e (t) is the weight for scaling the losses of chemical bond types and chemical bond lengths, and it is a monotonically increasing function with the time step t as the input; ||r - Ψ r || 2 Represents the bond length of the chemical bond type Ψ generated by the model r The mean squared error between the bond length r = bond length(e) of the current true chemical bond type e and Ψ, and Ψ r = ||Ψ xi - Ψ xj ||, ψ xi 、ψ xj Represents the three - dimensional atomic coordinates of atoms i and j.
[0075] Exemplarily, as Figure 3 shown, is the curve graph of the weight and noise schedule of this embodiment. Where atom refers to the noise schedule α x (t), α h (t), edge refers to the noise schedule α e (t) of the chemical bond type; loss weight refers to the loss weight λ e (t).
[0076] Furthermore, in an alternative embodiment, the diffusion model in the drug molecule generation model is combined with a graph neural network model; its training process is represented as:
[0077]
[0078]
[0079] Among them, Represents the distance feature between atoms i and j after the l - th layer update of the model; φ rbf (·) is the radial basis function, Represents the distance between atoms i and j after the (l - 1)-th layer update of the model, respectively represent the three - dimensional atomic coordinates of atoms i and j after the (l - 1)-th layer update of the model; Represents the interaction feature information between atoms i and j after the l - th layer update of the model; φ m (·) is the message passing module function of the model, respectively represent the atomic type features of atoms i and j after the (l - 1)-th layer update of the model; Represents the chemical bond type feature between atoms i and j after the (l - 1)-th layer update of the model; φ h(·) is the atomic type update function of the model; φ e (·) is the chemical bond type update function of the model; φ x (·) is the atomic three-dimensional coordinate update function of the model; represents the unit direction vector pointing from atom j to atom i, representing the direction from atom j to atom i, while keeping the vector modulus length as 1.
[0080] Exemplarily, as Figure 4 shown, it is the training update schematic diagram of the drug molecule generation model of this embodiment. This embodiment designs a simple graph neural network model for training, which can capture the complex dependencies between nodes, and model the dynamic interactions between continuous atomic coordinate features, discrete atomic type features and chemical bond features in a scientific and reasonable way, and then update the atomic and bond representations simultaneously during the message passing process.
[0081] Exemplarily, select the mainstream dataset Drug-QM9 in the field of molecule generation as the real molecule dataset, divide it into a training set and a validation set according to a ratio of 9:1, use the early stop method to train the molecule generation model, the maximum number of iterations for model training is 1000, the batch size is 64, the learning rate is 0.0001, and the optimizer is AdamW. Further, use atomic stability (Atom Sta), molecular stability (Mol Sta), validity (Validity), total variance of atomic types (AtomTV) and total variance of chemical bond types (BondTV) to evaluate the performance of the model. The evaluation results are shown in Table 1 below, and the schematic diagrams of some generated drug molecule generation results are as Figure 5 shown.
[0082] Table 1 Model performance evaluation results
[0083]
[0084] Combining Table 1 and Figure 4 it can be seen that the method MolBFN of this embodiment has higher geometric generation and modeling capabilities, and the experimental indicators of the total variance of atomic types and the total variance of chemical bond types prove that the present invention can make more reasonable use of prior knowledge to generate molecules that are more in line with the actual data distribution.
[0085] Embodiment 2
[0086] This embodiment applies the three-dimensional drug molecule generation method proposed in Embodiment 1 to propose a three-dimensional drug molecule generation system based on Bayesian flow network. As Figure 6 shown, it is the architecture diagram of the three-dimensional drug molecule generation system based on Bayesian flow network of this embodiment.
[0087] In the three-dimensional drug molecule generation system based on the Bayesian flow network proposed in this embodiment, it includes:
[0088] A prior module, which is equipped with prior statistical knowledge and is used to generate noise molecule features;
[0089] A noise addition module, which is equipped with a noise schedule, and is used to combine the noise molecule features and the random Gaussian noise obtained from the noise schedule to introduce noise to the input molecular data and generate noisy molecular data; meanwhile, update the data distribution in the prior statistical knowledge based on Bayesian inference;
[0090] A prediction module, which is used to denoise the noisy molecular data to generate predicted drug molecules; repeatedly execute noise addition and prediction until the molecular structure converges or reaches the preset number of steps T, and output the drug molecule generation result.
[0091] It can be understood that the system in this embodiment corresponds to the method in Embodiment 1 above, and the optional items in Embodiment 1 above are also applicable to this embodiment, so they will not be repeated here.
[0092] Embodiment 3
[0093] This embodiment proposes a computer device, which includes a memory and a processor. Computer-readable instructions are stored in the memory. When the computer-readable instructions are executed by the processor, the processor executes all or part of the steps of the three-dimensional drug molecule generation method proposed in Embodiment 1.
[0094] Embodiment 4
[0095] This embodiment proposes a storage medium, on which computer-readable instructions are stored. Among them, when the computer-readable instructions are executed by a processor, all or part of the steps of the three-dimensional drug molecule generation method proposed in Embodiment 1 are realized.
[0096] Exemplarily, the storage medium includes but is not limited to various media that can store program codes such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks or optical discs.
[0097] Exemplarily, the instructions, programs, code sets or instruction sets can be implemented using conventional programming languages.
[0098] Exemplarily, the processor includes but is not limited to smart phones, personal computers, servers, network devices, etc., and is used to execute all or part of the steps of the three-dimensional drug molecule generation method described in Embodiment 1.
[0099] Each embodiment in the present invention is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the apparatus embodiments, since they are basically similar to the method embodiments, they are described relatively simply, and for the relevant parts, reference can be made to the description of the method embodiments. The apparatus embodiments described above are merely exemplary. The modules described as separate components may or may not be physically separated. When implementing the solution of the present invention, the functions of the modules can be implemented in the same or multiple software and / or hardware. It is also possible to select some or all of the modules according to actual needs to achieve the purpose of the solution of this embodiment.
[0100] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention, rather than limitations on the implementation manners of the present invention. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to list all the implementation manners here. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included in the protection scope of the claims of the present invention.
Claims
1. A three-dimensional drug molecule generation method based on Bayesian flow network, characterized in that: The following steps are involved: The real initial molecular data is input into the drug molecule generation model combining the Bayesian flow network and the diffusion model; wherein: Generate noise molecular features based on prior statistical knowledge; Integrating the noise molecular features and the random Gaussian noise obtained from the noise plan table, introducing noise into the current molecular data to generate noisy molecular data; at the same time, updating the data distribution in the prior statistical knowledge based on Bayesian inference; De-noising the noisy molecular data to generate predicted drug molecules; Repeat the above steps until the molecular structure converges or reaches the preset step number T to obtain the drug molecule generation result.
2. The method for generating three-dimensional drug molecules according to claim 1, characterized in that: The prior statistical knowledge includes atomic type distribution, chemical bond distribution and atomic number distribution.
3. The method for generating three-dimensional drug molecules according to claim 1, characterized in that: The method of generating noise molecular features according to prior statistical knowledge comprises the following steps: The parameters of the prior statistical knowledge are input into the Bayesian flow network to obtain the parameters of the estimated distribution of the original data as the noise molecule characteristics.
4. The method for generating three-dimensional drug molecules according to claim 1, characterized in that: The generation of predicted drug molecules comprises the following steps: According to the current time step t, extract the corresponding random Gaussian noise β(t) from the noise schedule; Combining the random Gaussian noise β(t) of the current time step t with the noise molecular feature and adding noise to the current molecular data to obtain noisy molecular data; The noisy molecular data is denoised, and combined with the Bayesian flow framework, the atomic three-dimensional coordinates, atomic types and chemical bond types are predicted to obtain the predicted drug molecules at the current time step t.
5. The method for generating three-dimensional drug molecules according to claim 4, characterized in that: Predict the three-dimensional coordinates of atoms, including: Predict the noise β at the current time step t from the noise schedule corresponding to the atomic 3D coordinates x (t), and according to the atomic three-dimensional coordinates x in the initial molecular data, the atomic three-dimensional coordinates are restored in combination with the Bayesian flow network; its expression is: Among them, the random Gaussian noise corresponding to the three-dimensional coordinates of the atom is And t∈[0,1];θ x represents the predicted three-dimensional coordinates of atoms; represents the predicted atomic 3D coordinates at time step 0, which is random Gaussian noise; x represents the original true atomic 3D coordinates; It means that it obeys Gaussian distribution; μi represents the sampling result of the i-th time step, and I is the unit matrix.
6. The method for generating three-dimensional drug molecules according to claim 4, characterized in that: Predictions of atom types and chemical bond types, including: According to the classification distribution of the atom type and the chemical bond type in the initial molecular data, the parameters of the classification distribution are reparameterized and converted into Gaussian distribution of the atom type and the chemical bond type as the corresponding noise schedule; Predict the Gaussian noise β at the current time step t from the corresponding noise schedule h (t) and β e (t), and according to the atomic type and chemical bond type in the initial molecular data, the atomic type and chemical bond type are restored in combination with the Bayesian flow network, and the sampling result is converted to the probability simplex space of the classification distribution through softmax; its expression is: Among them, θ h represents the predicted atomic type, h represents the original real atomic type; β h (t) is the noise schedule of the atom type, and β h (t) = t 2 β h (1); 1 h Represents the one-hot vector corresponding to the actual atom type, 1 is a vector of all 1s; y h From Gaussian distribution The result of sampling in K h is the number of atomic types; δ(θ h -softmax(y h )) is the Dirichlet distribution function; θ e represents the predicted chemical bond type, e represents the original real chemical bond type; β e (t) is the noise schedule of chemical bond type, and β e (t) = t r β e (1), r is a positive integer greater than 2; 1 e Represents the unique hot vector corresponding to the actual chemical bond type; y e From Gaussian distribution The result of sampling in K e is the bond type of the chemical bond type.
7. The method for generating three-dimensional drug molecules according to any one of claims 1 to 6, characterized in that: The method further comprises the following steps: A real molecular data set is obtained for training the drug molecule generation model; wherein: Extract random Gaussian noise from the noise schedule according to random sampling time, and generate noisy molecular data in combination with real molecular data; Inputting the noisy molecular data into the drug molecule generation model to obtain predicted molecules; Based on the real molecular data and the predicted molecules, a loss function including physical and chemical constraints is calculated, and the drug molecule generation model is trained based on the early stopping method.
8. The method for generating three-dimensional drug molecules according to claim 7, characterized in that: The physicochemical constraints in the loss function include atomic valence loss constraints and chemical bond length loss constraints; the expression is: Among them, L ∞ (g) represents the loss function including physicochemical constraints, represents the reconstruction loss of the atomic 3D coordinate x, represents the reconstruction loss of atom type h, represents the reconstruction loss of chemical bond type e, represents the chemical bond length loss; the superscript ∞ represents the loss of the infinite "disassembly-reconstruction" step; t~U(0,1) represents the random uniform sampling of time step t from 0 to 1 for each training, and U(0,1) represents the random uniform distribution from 0 to 1; p F (θ g |g; t) represents the sampling results of atomic three-dimensional coordinates, atomic types and chemical bond types, g represents the predicted molecular graph containing atomic three-dimensional coordinates, atomic types and chemical bond types; α x (t), α h (t), α e (t) A noise table representing the atomic three-dimensional coordinates, atomic type, and chemical bond type; ||x-Ψ x || 2 Represents the three-dimensional coordinates of the atoms generated by the model Ψ x The mean square error between the current true atomic three-dimensional coordinate x, ||h-Ψ h || 2 Indicates the atomic type Ψ generated by the model h The mean square error between the current true atom type h, ||e-Ψ e || 2 Indicates the chemical bond type Ψ generated by the model e The mean square error between the current true chemical bond type e; λ e (t) is the weight for scaling the chemical bond type and chemical bond length loss; ||r-Ψ r || 2 The bond length Ψ represents the type of chemical bond generated by the model r The mean square error between the bond length r = bond length (e) of the current real chemical bond type e, and Ψ r =||Ψ xi -Ψ xj ||,Ψ xi , xj Represents the atomic three-dimensional coordinates of atom i and atom j.
9. The method for generating three-dimensional drug molecules according to claim 7, characterized in that: The diffusion model in the drug molecule generation model is combined with the graph neural network model; its training process is expressed as: in, Represents the distance feature between atoms i and j after the model layer l is updated; φ rbf (·) is the radial basis function, represents the distance between atoms i and j after the model l-1 layer is updated, Respectively represent the updated atomic three-dimensional coordinates of atom i and atom j in the l-1th layer of the model; Represents the interaction characteristic information between atoms i and j after the model layer l is updated; φ m (·) is the message passing module function of the model, They represent the atomic type characteristics of atom i and atom j after updating at the l-1th layer of the model; represents the chemical bond type characteristics of atoms i and j after updating in the l-1th layer of the model; φ h (·) is the atomic type update function of the model; φ e (·) is the chemical bond type update function of the model; φ x (·) is the atomic three-dimensional coordinate update function of the model.
10. A three-dimensional drug molecule generation system based on a Bayesian flow network, applied to the three-dimensional drug molecule generation method according to any one of claims 1 to 9, characterized in that: include: A priori module, which is equipped with a priori statistical knowledge and used to generate noise molecular features; A noise adding module, which is equipped with a noise schedule, is used to combine the noise molecular features and the random Gaussian noise obtained from the noise schedule to introduce noise into the input molecular data to generate noisy molecular data; at the same time, the data distribution in the prior statistical knowledge is updated based on Bayesian inference; The prediction module is used to denoise the noisy molecular data to generate predicted drug molecules; repeatedly perform denoising and prediction until the molecular structure converges or reaches a preset step number T, and output the drug molecule generation result.
Citation Information
Cited By
Structure-based drug design model integrating Bayesian flow network and diffusion model
CN120636604A