Molecular optimization method for protein pocket three-dimensional structure perception based on diffusion model
By using EGNN-based diffusion model and protein-ligand map in molecular optimization, combined with the technical means of fully-linked ligand map, the problem of the lack of a unified framework and ignoring chemical bond information in the molecular optimization process is solved, and efficient and real molecular optimization effects are achieved.
Patent Information
- Application Number
- CN202510110335.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-06-06
AI Technical Summary
The existing generative model based on 3D target perceived diffusion has limitations in the molecular optimization process, including the lack of a unified optimization framework, the ignorance of chemical bond information, and the impractical molecular structure generated.
The EGNN-based diffusion model is used for supervision and training, and the protein-ligand map and fully connected ligand map are constructed in combination with the k-nearest neighbor algorithm. The molecular model is optimized through reverse diffusion and information transmission to achieve fragment junction, skeleton modification or skeleton transition.
A unified framework for molecular optimization of three-dimensional protein pocket perception is realized, which improves the effectiveness and authenticity of the generated molecules, and solves the problem of generating invalid molecules and non-real structural compounds in the past models.
Smart Images

Figure CN120108489A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of molecular optimization, and more specifically, to a molecular optimization method for protein pocket three-dimensional structure perception based on a diffusion model. Background Art
[0002] Molecular optimization to improve the binding affinity of lead molecules is a critical and challenging step in the early stages of drug discovery. Traditionally, this is planned based on the knowledge, experience, and intuition of medicinal chemists, which are inherently limited and not scalable or automated. Many fragment-based computational methods have been developed to assist and accelerate the traditional process, including virtual fragment library construction and screening, fragment growth and ligation, and backbone hopping; however, these methods are based on computationally intuitive database searches and physical simulations and generate optimized compounds with limited novelty and diversity.
[0003] Recent data-driven deep learning methods have enabled alternative generation processes to accelerate structure-based molecular optimization. Many target-aware deep generative models have been developed to capture the interactions between proteins and ligands to generate molecules with high affinity. Early methods integrated protein pocket features as 1D sequence strings or 2D pharmacophore maps, which are limited because the atomic-level interactions between protein pockets and molecules are not explicitly trained and modeled. This led to early attempts at 3D pocket-aware generative models, which represented receptor-ligand complexes as atomic density grids and employed 3D conventional neural networks (CNNs) to capture the geometric and spatial information of protein pockets, such as Masuda et al. (Masuda, T.; Ragoza, M.; Koes, DRGenerating 3d molecular structures conditional on areceptor binding site with deep generative models.arXiv preprint arXiv:2010.14442 2020.) using 3D convolutional neural networks (CNNs) to capture spatial information and conditional variational autoencoders to sample 3D molecules. However, these 3D CNN models compress protein structural information and are difficult to extend to large protein pockets. To solve this problem, some researchers have proposed an autoregressive generative model based on a 3D graph neural network (GNN) to learn the sequential conditional distribution of different atoms in the 3D space around the binding pocket and generate molecules atom by atom in the pocket. Luo et al. (Luo, S.; Guan, J.; Ma, J.; Peng, JA 3D generative model for structure-based drug design. Advances in Neural Information Processing Systems 2021, 34, 6229-6239.) first proposed learning the probability density of different atoms in the three-dimensional space within the binding pocket, placing atoms based on the learned distribution, and generating molecules in an autoregressive manner.On this basis, Peng et al. (Peng, X.; Luo, S.; Guan, J.; Xie, Q.; Peng, J.; Ma, J. Pocket2mol: Efficient molecular sampling based on 3d protein pockets. International Conference on Machine Learning. 2022; pp 17644-17655.) proposed an E(3)-equivariant generative network consisting of two modules: a graph neural network that captures the spatial and bonding relationships between atoms in the binding protein pocket, and an efficient algorithm for sampling new drug candidates from a controllable distribution conditioned on the pocket representation without relying on the Markov chain Monte Carlo algorithm. This method uses an E(3)-equivariant graph neural network to learn the chemical and geometric constraints imposed by protein pockets, and also samples molecules in an autoregressive manner, that is, based on an atom-by-atom generation pattern. Zhang et al. (Zhang, Z.; Min, Y.; Zheng, S.; Liu, Q. Molecule generation for target protein binding with structural motifs. The Eleventh International Conference on Learning Representations. 2023.) believe that atom-based generation schemes may lead to unrealistic substructures and inefficient molecular sampling, so they proposed a fragment-based ligand generation framework, which extracts common molecular fragments from the dataset and generates three-dimensional molecules with valid and realistic substructures in a fragment-by-fragment manner under the conditions of protein pockets. However, this sequential generation method will lead to inaccurate and unreasonable 3D molecular structures because the position and element type of atoms will be affected by all other atoms in the molecule.
[0004] Recently, researchers have proposed a 3D object-aware equivariant diffusion model that learns the joint distribution of all atoms in a 3D bag to generate all atoms of a molecule or fragment at once. For example, DiffLinker (Igashov, I.; H.; Vignac, C.; Satorras, VG; Frossard, P.; Welling, M.; Bronstein, M.; Correia, B. Equivariant 3d-conditional diffusion models for molecular linker design. arXiv preprint arXiv:2210.05274 2022.) Fragment connection within 3D protein pockets using a denoised diffusion model.
[0005] Although recently developed generative models based on 3D target-aware diffusion have shown efficiency in generating molecules that bind to specific protein pockets, they still have some limitations. First, these models only focus on their current tasks such as backbone modification or fragment connection, while in actual molecular optimization scenarios, several methods are often required to assist, and there is a lack of an easy-to-use framework to unify several molecular optimization methods. Second, they often ignore chemical bond information during training, and the current diffusion models are limited to the generation of atomic types and coordinates. The bonds between atoms need to be additionally calculated and predicted by bond inference algorithms (such as OpenBabel), and their accuracy will greatly affect the effectiveness of the generated molecules, which may lead to unrealistic 3D molecular structures. Summary of the invention
[0006] In order to overcome the defects of the limitations of the generation model based on 3D target perception diffusion described in the above-mentioned prior art, the present invention provides a molecular optimization method for protein pocket three-dimensional structure perception based on a diffusion model.
[0007] In order to solve the above technical problems, the technical solution of the present invention is as follows:
[0008] In the first aspect, a molecular optimization method for protein pocket three-dimensional structure perception based on a diffusion model comprises:
[0009] The diffusion model constructed based on EGNN is forward diffused on a supervised training set to learn the correct data distribution of atomic coordinates, atomic types and bond types in the optimized molecule to obtain an initial model; the supervised training set includes the molecular fragment to be optimized, protein pocket information and the optimized complete molecule;
[0010] Constructing a protein-ligand graph based on a k-nearest neighbor algorithm, and constructing a fully connected ligand graph based on the bond types between atoms; wherein the protein-ligand graph is used to represent protein-protein, ligand-ligand and protein-ligand interactions, and the fully connected ligand graph is used to represent interactions within ligands;
[0011] The initial model is subjected to reverse diffusion, and the protein-ligand graph and the fully connected ligand graph are used for information transfer, and the molecular optimization model is obtained by iterative updating;
[0012] The molecular optimization model is used to perform fragment connection, backbone modification or backbone transition.
[0013] In a second aspect, a computer program product comprises a computer program or computer executable instructions, wherein when the computer program or computer executable instructions are executed by a processor, all or part of the steps of the method described in the first aspect are implemented.
[0014] In a third aspect, a computer-readable storage medium is provided, characterized in that at least one instruction, at least one program, code set or instruction set is stored on the storage medium, and the at least one instruction, at least one program, code set or instruction set is loaded and executed by a processor to implement the method described in the first aspect.
[0015] Compared with the prior art, the technical solution of the present invention has the following beneficial effects:
[0016] The present invention constructs a unified molecular optimization framework based on three-dimensional protein pocket perception of the diffusion model, unifies the three molecular optimization scenarios of backbone modification, fragment connection and backbone transition, and makes the small molecule optimization process more convenient; at the same time, by constructing a protein-ligand graph and introducing EGNN, the present invention achieves the effect of three-dimensional protein pocket perception, ensuring the optimization generation capability of high binding affinity with the pocket during the optimization process; in addition, by introducing a fully connected ligand graph into the diffusion model, the model has the ability to generate bond types, which solves the problem that the previous models generated more invalid molecules and non-real structure compounds. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 Schematic diagram of the process of the molecular optimization method for protein pocket three-dimensional structure perception based on the diffusion model in Example 1 of the present application. DETAILED DESCRIPTION
[0018] The terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the terms used in this way can be interchanged in appropriate cases, which is only to describe the distinction mode adopted by the objects of the same attributes in the embodiments of the present application when describing. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, so that the process, method, system, product or equipment containing a series of units need not be limited to those units, but may include other units that are not clearly listed or inherent to these processes, methods, products or equipment. The term "determine" widely covers various actions, and may include acquisition, calculation, calculation, processing, derivation, investigation, search (for example, search in a table, database or other data structure), ascertainment, and similar actions, and may also include reception (for example, receiving information), access (for example, accessing data in a memory) and similar actions, and may also include generation, creation, establishment and similar actions, as well as parsing, selection, selection and similar actions, etc. The relevant definitions of other terms will be given in the following description.
[0019] It should be noted that when an element is considered to be "connected" to another element, it can be directly connected to the other element, or connected to the other element through an intermediate element. In addition, the "connection" in the following embodiments should be understood as "electrical connection", "communication connection", etc. if there is transmission of electrical signals or data between the connected objects.
[0020] The drawings are for illustrative purposes only and should not be construed as limiting the present patent;
[0021] In order to better illustrate the present embodiment, some parts in the drawings may be omitted, enlarged or reduced, and do not represent the size of the actual product;
[0022] It is understandable to those skilled in the art that some well-known structures and their descriptions may be omitted in the drawings.
[0023] The technical solution of the present invention is further described below in conjunction with the accompanying drawings and embodiments.
[0024] Example 1
[0025] This embodiment provides a molecular optimization method for protein pocket three-dimensional structure perception based on diffusion model, see Figure 1 ,include:
[0026] The diffusion model constructed based on EGNN (E(n)Equivariant Graph Neural Networks) is forward diffused on a supervised training set to learn the correct data distribution of atomic coordinates, atomic types and bond types in the optimized molecules to obtain an initial model; the supervised training set includes the molecular fragment to be optimized, protein pocket information and the optimized complete molecule;
[0027] Constructing a protein-ligand graph based on a k-nearest neighbor algorithm, and constructing a fully connected ligand graph based on the bond types between atoms; wherein the protein-ligand graph is used to represent protein-protein, ligand-ligand and protein-ligand interactions, and the fully connected ligand graph is used to represent interactions within ligands;
[0028] The initial model is subjected to reverse diffusion, and the protein-ligand graph and the fully connected ligand graph are used for information transfer, and the molecular optimization model is obtained by iterative updating;
[0029] The molecular optimization model is used to perform fragment connection, backbone modification or backbone transition.
[0030] It should be noted that this embodiment proposes a molecular optimization unified framework (abbreviated as Diffleo) based on a diffusion model for three-dimensional protein pocket perception, which unifies the optimization models of three molecular optimization scenarios: backbone modification, fragment connection, and backbone transition, making the small molecule optimization process more convenient. At the same time, by constructing a protein-ligand graph and introducing EGNN, the molecular optimization model is helped to achieve the effect of three-dimensional protein pocket perception, ensuring the optimization generation capability of high binding affinity with protein pockets during the optimization process.
[0031] It should be understood that the molecular optimization model constructed in this embodiment can generate linkers with high binding affinity and similar to the real structure based on the input molecular fragments and specific protein pockets, greatly improving the effectiveness of the generated molecules.
[0032] In some preferred embodiments, the supervised training set is constructed based on the CrossDocked dataset, including:
[0033] Extract the original protein-ligand complex from the CrossDocked dataset, cluster the proteins in the original protein-ligand complex based on mmseqs2 to divide the training set and the test set, and segment the ligands in the original protein-ligand complex whose number of atoms does not exceed the preset atomic threshold based on the molecular segmentation method; for the backbone modification optimization task, the ligand is divided into a backbone and a modified fragment; for the fragment connection optimization task, the ligand is segmented into two fragments and a linker;
[0034] The skeleton or fragments obtained after ligand segmentation are used as the molecular fragments to be optimized, the protein pocket information is used as the conditional information, and the corresponding original protein-ligand complex is used as the optimized complete molecule to obtain a data set in the form of triples as the supervised training set; wherein the three-dimensional structures of all molecules in the supervised data set are represented by molecular graphs, including atomic coordinates, atomic types and bond types.
[0035] Those skilled in the art should understand that the data in the supervised dataset is divided into a training set and a test set.
[0036] In some specific implementation processes, mmseqs2 was used to cluster proteins with 30% protein sequence consistency, and RDKit was used to implement molecular segmentation technology (an MMPA-based algorithm) to segment molecules with no more than 40 atoms. Finally, 50,000 data were selected as training sets and 100 data were selected as test sets for the above backbone modification optimization tasks and fragment connection optimization tasks, respectively.
[0037] In some preferred embodiments, the performing forward diffusion on the supervised training set includes:
[0038] The molecular fragment to be optimized and protein pocket information are used as inputs of the diffusion model;
[0039] A time step t is randomly selected, and the diffusion model is made to perform forward diffusion noise addition on the atomic coordinates, atomic types and bond types respectively.
[0040] In some optional embodiments, the forward diffusion and noise addition for the atomic coordinates, atomic types, and bond types respectively includes:
[0041] For atomic coordinates, noise is iteratively introduced from a Gaussian distribution to the atoms within the molecular fragment to be optimized at each time step, while the coordinates of the original input fragment and protein part are kept unchanged;
[0042] For atom types and bond types, noise is introduced based on the categorical distribution.
[0043] Furthermore, the introducing noise based on classification distribution includes:
[0044] Given the state at time step t-1, the conditional distribution at time step t is expressed as follows:
[0045]
[0046] In the formula, represents the three-dimensional geometric coordinates of the i-th atom, represents the element type of the i-th atom, represents the bond type between the i-th atom and the j-th atom, represents the retained part of the original signal, represents the noise part to be added, is the identity matrix, K v Indicates the part of the atom type reserved, K b The reserved part indicating the key type;
[0047] Using the inherent Markov properties of the diffusion process, and the fact that the distribution q is independent at each time step, we can calculate the original molecule M 0 Derive the numerator M after adding noise t time step t The distribution of , that is, the molecular representation after adding noise t time steps:
[0048]
[0049] In the formula, is the predefined noise parameter, σ t is a parameter that maintains the variance of the diffusion method.
[0050] Further, the reverse diffusion includes:
[0051] After the forward diffusion is completed, the molecular fragment to be optimized and the protein pocket information are used as fixed conditional contexts to make the diffusion model perform reverse diffusion, and the protein-ligand graph and the fully connected ligand graph are used for information transfer to restore the original sample:
[0052]
[0053] In the formula, p θ It is a neural network for parameterized transformation, namely, EGNN network, and P is the protein pocket information as a condition; Indicated by Indicated by Get M t-1 The true posterior distribution of
[0054] The optimized complete molecule is used as the output label of the diffusion model, combined with the restored original sample, a loss function is calculated, and the initial model is updated according to the loss function.
[0055] In the above embodiment, the EGNN is used to perform reverse diffusion denoising on the three-dimensional molecular graph to learn the node (atom) and edge (atom type and bond type) features.
[0056] Furthermore, the loss function is the mean square error between the atomic coordinates of the original sample and the optimized complete molecule, and the KL divergence of the atomic type and bond type, and its expression is as follows:
[0057]
[0058]
[0059]
[0060]
[0061] In the formula, represents the atomic coordinate loss at time step t-1; Indicates atomic type loss; represents the key type loss; λ 1 and λ 2 represents the hyperparameter of the corresponding loss; μ θ (M t , t) i represents a θ-parameterized neural network; Indicated by M t , M 0 Get the true posterior distribution of the i-th atom type; Indicated by M t Get the predicted distribution of the i-th atom type; Indicated by M t , M 0 Get the true posterior distribution of the bond type between the i-th atom and the j-th atom; Indicated by M t Get the predicted distribution of the bond type between the i-th atom and the j-th atom.
[0062] It should be understood that during the training process, the back-propagation algorithm can be used to continuously conduct gradient updates to the model's parameter values. After iterating a specific Epoch, the model's verification loss value will be output. When the loss value decreases and converges, the model is considered stable.
[0063] Furthermore, the process of using the protein-ligand graph and the fully connected ligand graph to transfer information is as follows:
[0064]
[0065] m ij ←φ d (||x i -x j ||,e ij )#(14)
[0066]
[0067] h i ←h i +φh (Δh K,i +Δh L,i )#(16)
[0068]
[0069]
[0070]
[0071] x i ←x i +(Δx K,i +Δx L,i )·1 mask #(20)
[0072] In the formula, Δh K,i Indicates the protein-ligand diagram used to calculate h i The intermediate variable of L,i Indicates the h in the ligand-ligand diagram i The intermediate variable of K,i Indicates the protein-ligand diagram used to calculate x i The intermediate variable of L,i Indicates the value used in the ligand-ligand diagram to calculate x i The intermediate variable of φ d (·), φ h (·),φ e (·), Both represent learnable neural network layers; x i represents the three-dimensional geometric coordinates of the i-th atom; h i represents the hidden layer representation of the i-th atom, e ji represents the hidden layer representation of the bond between the i-th atom and the j-th atom, N K (i) represents the neighbors of the i-th atom in the molecule, E ij Specifies whether the edge between the i-th atom and the j-th atom corresponds to a protein-protein, ligand-ligand, or protein-ligand edge, 1 mask Indicator for the hidden part of the ligand to be generated for optimization.
[0073] It should be noted that the construction of protein-ligand graph and fully connected ligand graph is integrated into the network to enhance the representation of atoms and bonds in the molecule in the protein context.
[0074] In addition, the above embodiment proposes a method for molecular optimization combined with bond diffusion. By introducing a fully connected ligand graph into the diffusion model and performing noise addition and denoising operations on the bond types, the model has the ability to generate bond types, which solves the problem that previous models generate more invalid molecules and non-real structure compounds.
[0075] Optionally, the use of the molecular optimization model for fragment connection, backbone modification or backbone transition includes:
[0076] For the fragment connection task or the backbone modification task, the atomic coordinate part of the target molecule to be generated is represented by a Gaussian distribution, the atomic type and bond type of the target molecule to be generated are represented by a categorical distribution, and M is sampled from the Gaussian distribution and the categorical distribution. t (Can be combined with equations (5), (6), and (7)), then from p θ (M t-1 |M t ) Sample out M t-1 , where t = T, T-1, ..., 1, M t represents the molecular representation after adding noise t time steps, and gradually reduces the output noise by using the molecular optimization model, and finally obtains the optimized target molecule;
[0077] For the skeleton transition task, the target molecule part to be transitioned is masked and marked, and the part is firstly denoised with a certain step length, and then denoised using the molecular optimization model to obtain the optimized target molecule.
[0078] In some specific implementations, for the optimization of backbone transition, t-step noise is first added to a specific fragment through a forward diffusion process, and then denoising is performed starting from the Tt-th step of the molecule generation process to obtain the compound molecule after the transition.
[0079] Example 2
[0080] This embodiment provides a computer-readable storage medium, on which is stored at least one instruction, at least one program, code set or instruction set, and the at least one instruction, at least one program, code set or instruction set is loaded and executed by a processor, so that the processor executes part or all of the steps of the method provided in Example 1 of the present application.
[0081] It is understood that the storage medium may be transient or non-transient. Exemplarily, the storage medium includes, but is not limited to, a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and other media that can store program codes.
[0082] Exemplarily, the processor may be a central processing unit (CPU), a microprocessor (MPU), a digital signal processor (DSP), or a field programmable gate array (FPGA).
[0083] Exemplarily, the read-only memory includes but is not limited to MASK ROM, PROM, EPROM, EEPROM, Flash, etc.
[0084] Exemplarily, the random access memory includes but is not limited to DRAM, SRAM, SDRAM, DDR SDRAM, etc.
[0085] In some examples, a computer program product is provided, which can be implemented in hardware, software or a combination thereof. As a non-limiting example, the computer program product can be embodied as the storage medium, or as a software product, such as an SDK (Software Development Kit).
[0086] As a non-limiting example, a computer program product is provided, the computer program product includes a computer program or a computer executable instruction, the computer program or the computer executable instruction is stored in a computer-readable storage medium. A processor of an electronic device reads the computer program or the computer executable instruction from the computer-readable storage medium, and the processor executes the computer executable instruction, so that the electronic device performs some or all steps of the method described in the embodiment of the present application.
[0087] In some examples, a computer program is provided, comprising a computer-readable code, and when the computer-readable code is run in a computer device, a processor in the computer device executes a part or all of the steps in the method.
[0088] This embodiment also proposes an electronic device, including a memory and a processor, wherein the memory stores at least one instruction, at least one program, a code set or an instruction set, and when the processor executes the at least one instruction, at least one program, a code set or an instruction set, it implements part or all of the steps of the method described in Example 1.
[0089] In some examples, a hardware entity of the electronic device is provided, including: a processor, a memory and a communication interface; wherein the processor generally controls the overall operation of the electronic device; the communication interface is used to enable the electronic device to communicate with other terminals or servers through a network; the memory is configured to store instructions and applications executable by the processor, and can also cache data to be processed or processed by the processor and various modules in the electronic device (including but not limited to image data, audio data, voice communication data and video communication data), which can be implemented by flash memory (FLASH), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM) or random access memory (RAM).
[0090] The processor may include one or more processing elements. Thus, the processor may include one or more integrated circuits (ICs) configured to perform the functions of the processor. In addition, each integrated circuit may include a circuit (e.g., a first circuit, a second circuit, and other circuits, etc.) configured to perform the functions of the processor.
[0091] Furthermore, data may be transmitted between the processor, the communication interface and the memory via a bus, which may include any number of interconnected buses and bridges, and the bus connects various circuits of one or more processors and memories together.
[0092] It can be understood that the options in the above-mentioned embodiment 1 are also applicable to this embodiment, so they will not be described again here.
[0093] The same or similar reference numerals correspond to the same or similar components;
[0094] The terms used to describe the positional relationship in the drawings are only used for illustrative purposes and should not be construed as limiting the present application;
[0095] It should be noted that, in the absence of conflict, the embodiments and features in the embodiments of the present application may be combined with each other.
[0096] In different specific implementations, the method or system described in the present application can be implemented in software, hardware or a combination thereof. In addition, the order of the steps of the method can be changed, and various elements can be added, reordered, combined, omitted, modified, etc.
[0097] Obviously, the above-mentioned embodiments of the present application are merely examples for clearly illustrating the present application, and are not intended to limit the implementation methods of the present application, and are not intended to limit the present application. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. Each discrete structure / function module or unit can be integrated together to form an independent part, or each module can exist separately, or two or more modules can be integrated to form an independent part, and the structure and function of the discrete components can be implemented as a combined structure or component. It is not necessary and impossible to enumerate all the implementation methods here. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application should be included in the scope of protection of the claims of the present application.
Claims
1. A molecular optimization method for protein pocket three-dimensional structure perception based on diffusion model, characterized in that: include: The diffusion model constructed based on EGNN is forward diffused on the supervised training set to learn the correct data distribution of atomic coordinates, atomic types and bond types in the optimized molecules to obtain an initial model; wherein the supervised training set includes the molecular fragment to be optimized, protein pocket information and the optimized complete molecule; Constructing a protein-ligand graph based on a k-nearest neighbor algorithm, and constructing a fully connected ligand graph based on the bond types between atoms; wherein the protein-ligand graph is used to represent protein-protein, ligand-ligand and protein-ligand interactions, and the fully connected ligand graph is used to represent interactions within ligands; The initial model is subjected to reverse diffusion, and the protein-ligand graph and the fully connected ligand graph are used for information transfer, and the molecular optimization model is obtained by iterative updating; The molecular optimization model is used to perform fragment connection, backbone modification or backbone transition.
2. According to claim 1, a molecular optimization method for protein pocket three-dimensional structure perception based on diffusion model is characterized in that: The supervised training set is constructed based on the CrossDocked dataset, including: Extract the original protein-ligand complex from the CrossDocked dataset, cluster the proteins in the original protein-ligand complex based on mmseqs2 to divide the training set and the test set, and segment the ligands in the original protein-ligand complex whose number of atoms does not exceed the preset atomic threshold based on the molecular segmentation method; for the backbone modification optimization task, the ligand is divided into a backbone and a modified fragment; for the fragment connection optimization task, the ligand is segmented into two fragments and a linker; The skeleton or fragments obtained after ligand segmentation are used as the molecular fragments to be optimized, the protein pocket information is used as the conditional information, and the ligand in the corresponding original protein-ligand complex is used as the optimized complete molecule to obtain a data set in the form of triples as the supervised training set; wherein the three-dimensional structures of all molecules in the supervised data set are represented by molecular graphs, including atomic coordinates, atomic types and bond types.
3. A molecular optimization method for protein pocket three-dimensional structure perception based on diffusion model according to any one of claims 1-2, characterized in that: The forward diffusion is performed on the supervised training set, including: The molecular fragment to be optimized and protein pocket information are used as inputs of the diffusion model; A time step t is randomly selected, and the diffusion model is made to perform forward diffusion noise addition on the atomic coordinates, atomic types and bond types respectively.
4. The molecular optimization method for protein pocket three-dimensional structure perception based on diffusion model according to claim 3, characterized in that: The forward diffusion and noise addition for the atomic coordinates, atomic types and bond types respectively include: For atomic coordinates, noise is iteratively introduced from a Gaussian distribution to the atoms within the molecular fragment to be optimized at each time step, while the coordinates of the original input fragment and protein part are kept unchanged; For atom types and bond types, noise is introduced based on the categorical distribution.
5. The molecular optimization method for protein pocket three-dimensional structure perception based on diffusion model according to claim 4, characterized in that: The introducing of noise based on classification distribution includes: Given the state at time step t-1, the conditional distribution at time step t is expressed as follows: In the formula, represents the three-dimensional geometric coordinates of the i-th atom, represents the element type of the i-th atom, represents the bond type between the i-th atom and the j-th atom, represents the retained part of the original signal, represents the noise part to be added, is the identity matrix, K v Indicates the part of the atom type reserved, K b The reserved part indicating the key type; Using the inherent Markov properties of the diffusion process, and the fact that the distribution q is independent at each time step, the original molecule M 0 The molecule M generated after adding noise t time steps t The distribution of is as follows: In the formula, is the predefined noise parameter, σ t is a parameter that maintains the variance of the diffusion method.
6. The molecular optimization method for protein pocket three-dimensional structure perception based on diffusion model according to claim 4, characterized in that: The reverse diffusion includes: After the forward diffusion is completed, the molecular fragment to be optimized and the protein pocket information are used as fixed conditional contexts to make the diffusion model perform reverse diffusion, and the protein-ligand graph and the fully connected ligand graph are used for information transfer to restore the original sample: In the formula, p θ represents the EGNN network used for parameterized transformation, and P is the protein pocket information as a condition; Indicated by M t ,P gets M t-1 The true posterior distribution of The optimized complete molecule is used as the output label of the diffusion model, combined with the restored original sample, a loss function is calculated, and the initial model is updated according to the loss function.
7. The molecular optimization method for protein pocket three-dimensional structure perception based on diffusion model according to claim 6, characterized in that: The loss function is the mean square error between the atomic coordinates of the original sample and the optimized complete molecule, as well as the KL divergence of the atomic type and bond type, and its expression is as follows: In the formula, represents the atomic coordinate loss at time step t-1; Indicates atomic type loss; represents the key type loss; λ1 and λ2 represent the hyperparameters of the corresponding loss; μ θ (M t ,t) i represents a θ-parameterized neural network; Indicated by M t ,M 0 Get the true posterior distribution of the i-th atom type; Indicated by M t Get the predicted distribution of the i-th atom type; Indicated by M t ,M 0 Get the true posterior distribution of the bond type between the i-th atom and the j-th atom; Indicated by M t Get the predicted distribution of the bond type between the i-th atom and the j-th atom.
8. A molecular optimization method for protein pocket three-dimensional structure perception based on diffusion model according to claim 7, characterized in that: The method of using the molecular optimization model to perform fragment connection, backbone modification or backbone transition includes: For the fragment connection task or the backbone modification task, the atomic coordinate part of the target molecule to be generated is represented by a Gaussian distribution, the atomic type and bond type of the target molecule to be generated are represented by a categorical distribution, and M is sampled from the Gaussian distribution and the categorical distribution. t , then from p θ (M t-1 |M t ) Sample out M t-1 , where t = T, T-1, ..., 1, M t represents the molecular representation after adding noise t time steps, and the output noise is gradually reduced by using the molecular optimization model to finally obtain the optimized target molecule; For the skeleton transition task, the target molecule part to be transitioned is masked and marked, and the part is firstly denoised with a certain step length, and then denoised using the molecular optimization model to obtain the optimized target molecule.
9. A computer program product comprising a computer program or computer executable instructions, characterized in that: When the computer program or computer executable instructions are executed by a processor, the method according to any one of claims 1 to 8 is implemented.
10. A computer-readable storage medium, characterized in that: The storage medium stores at least one instruction, at least one program, code set or instruction set, and the at least one instruction, at least one program, code set or instruction set is loaded and executed by the processor to implement the method as described in any one of claims 1-8.
Citation Information
Cited By
Molecular docking generation method and related equipment
CN121034390A
Crystal structure optimization method based on diffusion model
CN121281688A