Drug molecule generation method based on isotropic diffusion model
By using an isovariable diffusion model based on drug molecules, using geometric fully sensing networks and rotary translation mirrors to encode and latent diffusion atomic coordinates, the existing models have insufficient prediction capabilities in multimodal feature learning and noise reduction kernels, and have achieved stronger molecular generation capabilities and attribute distribution fitting effects.
Patent Information
- Application Number
- CN202510496615.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-21
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-04-21
AI Technical Summary
When processing drug molecules, existing diffusion models are difficult to effectively learn multimodal features, and the network parameterized noise reduction kernel prediction capability is insufficient, resulting in limited diversity and accuracy of molecular generation.
Using a drug molecule generation method based on isovariable diffusion model, the atomic coordinates are encoded through an improved rotational translation mirrored isovariable autoencoder designed by geometric fully sensing network, and geometrically complete latent diffusion is performed in the diffusion model, and the losses of the autoencoder and diffusion model are optimized.
Effective learning of potential representations is achieved, the robustness and effectiveness of molecular generation is improved, and the data distribution of the generated molecular physical properties shows a high degree of fit with the attribute distribution of real data.
Smart Images

Figure CN120015172A_ABST
Abstract
Description
Technical Field
[0001] The invention relates to the field of drug molecule generation, and in particular to a drug molecule generation method based on an isotropic diffusion model. Background Art
[0002] Drug molecule design and discovery methods mostly rely on experimental screening, theoretical calculations, and the empirical knowledge of chemists, but these methods often seem powerless when faced with a vast molecular space and complex biological systems. Traditional molecular discovery methods are inherently time-intensive, and human resources cannot complete the exploration of a vast chemical space in a limited and short period of time, thus limiting the diversity of molecules that can be explored.
[0003] Deep generative models provide a feasible solution to the above problems. By learning the distribution of various compound molecules through deep generative models, stable and robust molecules can be sampled from the generative model. With the help of the powerful computing power of computers, this process takes less time and does not require human input, thus freeing chemical researchers from time-consuming and labor-intensive molecular screening to conduct other more valuable research. There are five popular modeling methods for molecular generative models, namely autoregressive models, variational autoencoders, flow-based models, generative adversarial networks, and diffusion models.
[0004] In recent years, diffusion models, as a new type of generative model, have been applied to a variety of generative tasks. Diffusion models define a process of gradually perturbing data with noise, and reverse the above process by learning to gradually denoise through a neural network. There are still some problems with existing diffusion models. First, the necessary information contained in a molecule includes atomic coordinates, type, atomic charge, and certain physical and chemical properties of the molecule. This means that a molecule can be regarded as composed of multiple discrete or continuous features. Therefore, the diffusion model operating in the atomic feature space is limited in its learning ability in the face of multimodal features, and a unified Gaussian diffusion framework for multiple modes is not optimal. Secondly, what kind of network parameterized denoising kernel can more accurately predict the noise of the sample, which is also a problem that needs to be solved by the diffusion-based molecular generation model.
[0005] Therefore, there is a need for an equivariant diffusion model-based drug molecule generation method that can learn latent representations. Summary of the invention
[0006] The main purpose of the present invention is to provide a method for generating drug molecules based on an equivariant diffusion model to solve the problem in the prior art that drug molecules cannot learn potential representations.
[0007] To achieve the above object, the present invention provides a method for generating drug molecules based on an isotropic diffusion model, which specifically comprises the following steps: S1, the drug molecular structure is represented by a graph structure.
[0008] S2, a rotation-translation-mirror equivariant autoencoder improved based on the geometry-fully-aware network design, encodes the atomic coordinates.
[0009] S3, based on the diffusion model, performs geometric complete latent diffusion.
[0010] S4, add the loss of the autoencoder and the diffusion model loss to get the total loss, and optimize the total loss by gradient descent.
[0011] Furthermore, S1 specifically includes the following steps: S1.1, given a graph , Respectively represent the graph The node set and edge set of ; Represents the number of nodes in the graph, Represents the coordinates of the node set in three-dimensional space.
[0012] S1.2, The node features are composed of scalar features and Value vector feature constitute, Each edge of and Value vector feature constituted by Represent the length of the scalar feature of each node and each edge respectively, then: ,in, Indicates scalar feature of a node, Indicates The vector features of nodes, Representation Node and nodes Scalar feature of the edge between , Representation Node and nodes The vector features of the edges between .
[0013] Furthermore, S2 specifically includes the following steps: S2.1, geometric complete message passing is defined as: ; ; ; ; , , ; in, Representation Node and nodes The messages between Represents the geometric complete message passing function, through The first layer of geometrically fully aware convolution is obtained Features of nodes , is a trainable network, is an aggregate function, express The neighbor set of After aggregation News, , and is the intermediate variable, is the geometric complete frame, is the Euclidean norm.
[0014] S2.2, update the atomic coordinates in the figure: ; in, It is a fully geometrically aware module. is the updated atomic feature, For the Tier The atomic coordinates of the nodes.
[0015] S2.3, the geometric full convolution is expressed as: ; in, For the The atomic coordinates of the layers, For the Node scalar features of the layer, It is a geometric complete convolution.
[0016] The geometric complete convolution complies with the constraints of rotation, translation, and mirror equivariance, namely: and ; in, represents the rotation change in Euclidean space, Represents a translation change.
[0017] Furthermore, step S2 further includes the following steps: S2.4, the encoding process of the improved rotation and translation mirror equivariant autoencoder is: ; in, , are the expressions of atomic coordinates and node scalar features in latent space, , are the reconstructed atomic coordinates and node scalar features, are the atomic coordinates and node scalar features before reconstruction, It's noise.
[0018] S2.5, the decoding process of the improved rotation translation mirror equivariant autoencoder is: .
[0019] S2.6, Autoencoder Loss The definitions are as follows: .
[0020] Furthermore, S3 specifically includes the following steps: S3.1, the diffusion model gradually adds noise to the sample through forward diffusion: ; ; in, is the preset hyperparameter and ; is the vector of atomic coordinates and node scalar features, i.e. latent features; is the conditional probability; is a standard normal distribution; for noise; is the identity matrix.
[0021] S3.2, diffusion model uses reverse diffusion to treat noise variables Denoising to approximate clean samples : ; in, is the mean, is the variance; The distribution fitted by the diffusion model.
[0022] S3.3, by Bayes formula: ; in, is the noise randomly sampled from the standard normal distribution, that is .
[0023] S3.4, the diffusion model reduces the noise of the latent features and the decoder restores the features to the geometric space, that is: ; ; in, is the output of the noise prediction network, , is a random value sampled from a standard normal distribution, that is , Decoder for an improved rotation-translation-mirror equivariant autoencoder.
[0024] S3.5, Loss of Diffusion Model Defined as: ; in, For time.
[0025] Furthermore, S4 specifically includes the following steps: S4.1, add the loss of the autoencoder and the diffusion model loss to get the total loss: .
[0026] S4.2, by gradient descent Make optimizations.
[0027] The present invention has the following beneficial effects: The present invention is a model that satisfies rotational translational mirror image equivariance constraints and has a stronger learning ability for chiral molecules than the conventional E(3) equivariant model, which can hardly distinguish the difference between a molecule and its mirror image isomer.
[0028] After training with the same data set as other models, the molecules randomly sampled by the present invention have stronger robustness and effectiveness. Secondly, the data distribution of the physical properties of the molecules randomly sampled by the present invention shows a high degree of fit with the attribute distribution of the real data, proving that the present invention has sufficient learning ability for the attribute distribution of real drug molecules. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] In order to more clearly illustrate the specific implementation of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the specific implementation or the prior art description. Obviously, the drawings described below are some implementations of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work. In the drawings: Figure 1A flow chart of a method for generating drug molecules based on an isotropic diffusion model of the present invention is shown. DETAILED DESCRIPTION
[0030] The technical solution of the present invention will be described clearly and completely below in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0031] like Figure 1 A drug molecule generation method based on an equivariant diffusion model is shown to solve the problem that the existing technology cannot learn the potential representation.
[0032] To achieve the above object, the present invention provides a method for generating drug molecules based on an isotropic diffusion model, which specifically comprises the following steps:
[0033] S1, the drug molecular structure is represented by a graph structure.
[0034] S2, a rotation-translation-mirror equivariant autoencoder improved based on the geometry-fully-aware network design, encodes the atomic coordinates.
[0035] S3, based on the diffusion model, performs geometric complete latent diffusion.
[0036] S4, add the loss of the autoencoder and the diffusion model loss to get the total loss, and optimize the total loss by gradient descent.
[0037] Specifically, S1 includes the following steps: S1.1, given a graph , Respectively represent the graph The node set and edge set of ; Represents the number of nodes in the graph, Represents the coordinates of the node set in three-dimensional space.
[0038] S1.2, The node features are composed of scalar features and Value vector feature constitute, Each edge of and Value vector feature constituted by Represent the length of the scalar feature of each node and each edge respectively, then: ,in, Indicates scalar feature of a node, Indicates The vector features of nodes, Representation Node and nodes Scalar feature of the edge between , Representation Node and nodes The vector features of the edges between .
[0039] Specifically, S2 includes the following steps: S2.1, geometric complete message passing is defined as: ; ; ; ; , , ; in, Representation Node and nodes The messages between Represents the geometric complete message passing function, through the vector and scalar features of the node and the geometric complete frame To calculate the message of the graph neural network. The first layer of geometrically fully aware convolution is obtained Features of nodes , is a trainable network, is an aggregate function, express The neighbor set of After aggregation News, , and is the intermediate variable, is the geometric complete frame, is the Euclidean norm.
[0040] S2.2, update the atomic coordinates in the figure: ; in, It is a fully geometrically aware module. is the updated atomic feature, For the Tier By introducing two types of atomic features and geometric complete frames As input features, provide atomic scalar features that are rotationally and translationally invariant updates Updates that are equivariant to rotations with vector features , and based on To update the atomic coordinates. After updating the features once through geometric complete message passing and updating the coordinates once through the geometric complete perception module, a geometric complete perception convolution is completed, that is, the geometric complete convolution layer is composed of geometric complete message passing and coordinate update modules.
[0041] S2.3, in the above process, the vector feature of the node , the vector feature of the edge Fully framed with geometry All are the atomic coordinates of the current layer In the field of molecular generation, the entire molecular graph is often treated as a fully connected graph, that is, any two different points in a molecular graph are connected, and the scalar characteristics of the edges of the graph are In fact, it can be regarded as the fully connected adjacency matrix of the current graph. In fact, they are completely equal. It only provides information for updating node features. Geometric complete convolution does not update edge features, so the scalar features of edges are ignored in the molecular generation model. In fact, the entire geometric complete convolution only relies on the coordinates and scalar features of atoms as input to complete the update of coordinates and features. The geometric complete convolution is expressed as: ; in, For the The atomic coordinates of the layers, For the Node scalar features of the layer, It is a geometric complete convolution.
[0042] The geometric complete convolution complies with the constraints of rotation, translation, mirroring and translational equivariance, that is, translation and rotation equivariance, namely: and ; in, represents the rotation change in Euclidean space, Represents a translation change.
[0043] Compared to the SE(3) (rotation-translation mirror equivariance) constraint, the E(3) equivariance constraint (i.e., rotation-translation equivariance constraint) is insensitive to reflection changes due to its equivariance to reflection changes. For molecules that are mirror image isomers of each other, the SE(3) equivariance model can accurately capture the difference between the two, while the E(3) equivariance model cannot accurately distinguish between the two.
[0044] Specifically, step S2 also includes the following steps: S2.4, the encoding process of the improved rotation and translation mirror equivariant autoencoder is: ; in, , are the expressions of atomic coordinates and node scalar features in latent space, , are the reconstructed atomic coordinates and node scalar features, are the atomic coordinates and node scalar features before reconstruction, It's noise.
[0045] In fact, after the feature is geometrically fully convolutionally encoded, we do not output it directly, but take the encoded feature itself as the mean and sample it once with a variance close to 0. By introducing a small amount of noise in the latent space, the latent space can be smoother and more continuous. The decoding process of the improved rotation, translation, mirror and equivariant autoencoder is: .
[0046] S2.6, Autoencoder Loss The definitions are as follows: .
[0047] Specifically, thanks to the encoder and decoder of the geometric complete autoencoder, the present invention is able to implement a denoising diffusion probability model, namely a latent space diffusion model, in a low-dimensional, continuous latent space. The diffusion model models the distribution of the latent representation through a forward diffusion process. In this process, the latent representation is gradually perturbed by Gaussian noise until it satisfies the standard normal distribution. At the same time, the denoising kernel of the diffusion model learns the ability to predict sample noise. During sampling, the diffusion model samples noise from the standard normal distribution, and gradually denoises the sample with the denoising ability of the denoising kernel to obtain the latent representation of the molecule, which is then reconstructed through the decoding process of the geometric complete autoencoder. The diffusion model can be described by two Markov chains. The forward Markov chain describes the process of converting the latent representation Gradually add noise to get The process can be expressed as conditional probability To describe, the reverse Markov chain, i.e. the denoising process, can be described by S3 specifically includes the following steps: S3.1, the diffusion model gradually adds noise to the sample through forward diffusion: ; ; in, is the preset hyperparameter and , , all parameters of the forward process are pre-set hyperparameters and cannot be trained; is the vector of atomic coordinates and node scalar features, i.e. latent features; is the conditional probability; is a standard normal distribution; for noise; is the identity matrix.
[0048] S3.2, diffusion model uses reverse diffusion to treat noise variables Denoising to approximate clean samples : ; in, is the mean, is the variance; The distribution fitted by the diffusion model. is the output of the neural network, that is, using a neural network to predict the mean of the distribution of forward diffusion. Its variance The default is the same as the variance of forward diffusion.
[0049] S3.3, by Bayes formula: ; in, is the noise randomly sampled from the standard normal distribution, that is Diffusion models do not necessarily use neural networks Direct prediction The mean of All other parameters are known, so diffusion models usually use neural networks Directly predict random noise .
[0050] S3.4, the present invention uses geometric complete convolution to parameterize the denoising kernel. When using geometric complete convolution as a 3D image denoiser, the diffusion model not only achieves SE(3) equivariance, but its geometric completeness also enables the denoiser to predict the noise more accurately. When training the model of the present invention, a strategy of jointly training the autoencoder and the latent space diffusion is adopted. The drug molecules are first encoded by the encoding process of the autoencoder. After the features are mapped to the latent space by the encoder, the diffusion model learns the distribution of the latent representation and calculates At the same time, the potential representation is restored to the geometric space to obtain the reconstructed features and calculate the reconstruction loss of the autoencoder The sum of the two Optimization is performed through gradient descent. As for the sampling process, noise is directly sampled from the standard normal distribution in the latent space. The diffusion model reduces the noise of the latent features and the decoder restores the features to the geometric space, that is, the diffusion model reduces the noise of the latent features and the decoder restores the features to the geometric space, that is: ; ; in, is the output of the noise prediction network, , is the intermediate variable, is a random value sampled from a standard normal distribution, that is , Decoder for an improved rotation-translation-mirror equivariant autoencoder.
[0051] S3.5, Loss of Diffusion Model Defined as: ; in, For time.
[0052] Specifically, S4 includes the following steps: S4.1, add the loss of the autoencoder and the diffusion model loss to get the total loss: .
[0053] S4.2, by gradient descent Make optimizations.
[0054] The present invention designs a geometric complete autoencoder that satisfies SE(3) equivariant constraints based on geometric complete perceptual convolution. With the help of the geometric complete autoencoder, the present invention realizes latent space diffusion for learning potential representations.
[0055] In order to verify the method provided by the present invention, a comparative experiment was conducted. As shown in Table 1, the molecular stability, effectiveness and uniqueness of the method GCLDM proposed by the present invention are high. The experiment shown in Table 2 uses polarizability, dipole rate and heat capacity as conditional constraints, respectively, to allow each model to generate molecules that meet the conditional constraints. The smaller the number in Table 2, the smaller the property of the generated molecule and the given property constraint, that is, it is more in line with the proposed requirements. It can be seen from Table 2 that the average absolute error of the molecular property prediction of the present invention is smaller, and the controllability of the conditional generation is stronger.
[0056] Table 1 Comparison results of GCLDM and baseline methods for molecular generation
[0057] Table 2 Comparison of the mean absolute error between expected value and predicted value when generating conditional constraints
[0058] Of course, the above description is not a limitation of the present invention, and the present invention is not limited to the above examples. Changes, modifications, additions or substitutions made by technicians in this technical field within the essential scope of the present invention should also fall within the protection scope of the present invention.
Claims
1. A method for generating drug molecules based on an isotropic diffusion model, characterized in that: The specific steps include: S1, the drug molecular structure is represented by a graph structure; S2, a rotation-translation-mirror equivariant autoencoder improved based on the geometry-fully-aware network design, encodes atomic coordinates; S3, based on the diffusion model, geometric complete latent diffusion is performed; S4, add the loss of the autoencoder and the diffusion model loss to get the total loss, and optimize the total loss by gradient descent.
2. A method for generating drug molecules based on an isotropic diffusion model according to claim 1, characterized in that: S1 specifically includes the following steps: S1.1, given a graph , Respectively represent the graph The node set and edge set of ; Represents the number of nodes in the graph, Represents the coordinates of the node set in three-dimensional space; S1.2, The node features are composed of scalar features and Value vector feature constitute, Each edge of and Value vector feature constituted by Represent the length of the scalar feature of each node and each edge respectively, then: ,in, Indicates scalar feature of a node, Indicates The vector features of nodes, Representation Node and nodes Scalar feature of the edge between , Representation Node and nodes The vector features of the edges between .
3. The method for generating drug molecules based on an isotropic diffusion model according to claim 1, characterized in that: S2 specifically includes the following steps: S2.1, geometric complete message passing is defined as: ; ; ; ; , , ; in, Representation Node and nodes The messages between Represents the geometric complete message passing function, through The first layer of geometrically fully aware convolution is obtained Features of nodes , is a trainable network, is an aggregate function, express The neighbor set of After aggregation The news, , and is the intermediate variable, is the geometric complete frame, is the Euclidean norm; S2.2, update the atomic coordinates in the figure: ; in, It is a fully geometrically aware module. is the updated atomic feature, For the Tier Atomic coordinates of nodes; S2.3, the geometric full convolution is expressed as: ; in, For the The atomic coordinates of the layers, For the Node scalar features of the layer, It is a geometric complete convolution; The geometric complete convolution complies with the constraints of rotation, translation, and mirror equivariance, namely: and ; in, represents the rotation change in Euclidean space, Represents a translation change.
4. The method for generating drug molecules based on an isotropic diffusion model according to claim 3, characterized in that: Step S2 also includes the following steps: S2.4, the encoding process of the improved rotation and translation mirror equivariant autoencoder is: ; in, , are the expressions of atomic coordinates and node scalar features in latent space, , are the reconstructed atomic coordinates and node scalar features, are the atomic coordinates and node scalar features before reconstruction, It is noise; S2.5, the decoding process of the improved rotation translation mirror equivariant autoencoder is: ; S2.6, Autoencoder Loss The definitions are as follows: 。 5. The method for generating drug molecules based on an isotropic diffusion model according to claim 1, characterized in that: S3 specifically includes the following steps: S3.1, the diffusion model gradually adds noise to the sample through forward diffusion: ; ; in, is the preset hyperparameter and , is the vector of atomic coordinates and node scalar features, i.e. latent features; is the conditional probability; is a standard normal distribution; for noise; is the identity matrix; S3.2, diffusion model uses reverse diffusion to treat noise variables Denoising to approximate clean samples : ; in, is the mean, is the variance; The distribution fitted by the diffusion model; S3.3, by Bayes formula: ; in, is the noise randomly sampled from the standard normal distribution, that is ; S3.4, the diffusion model reduces the noise of the latent features and the decoder restores the features to the geometric space, that is: ; ; in, is the output of the noise prediction network, , is a random value sampled from a standard normal distribution, that is , Decoder for an improved rotation-translation-mirror equivariant autoencoder; S3.5, Loss of Diffusion Model Defined as: ; in, For time.
6. The method for generating drug molecules based on an isotropic diffusion model according to claim 1, characterized in that: S4 specifically includes the following steps: S4.1, add the loss of the autoencoder and the diffusion model loss to get the total loss: ; S4.2, by gradient descent Make optimizations.
Citation Information
Patent Citations
Intelligent molecule generation method and device based on diffusion model, equipment and medium
CN116665807A
Generation method and device of composite molecular structure, electronic equipment and storage medium
CN116978448A
Drug molecule generation method based on diffusion model and multi-scale diagram
CN118471385A
Multimodal molecule joint generation method, device and equipment based on diffusion model and medium
CN118969129A
Drug molecule optimization method based on hidden space diffusion model
CN119314590A
Cited By
Three-dimensional molecular structure generation method and model
CN120998348A