An information processing method based on protein target structure information
By setting virtual particles in the cavity of the protein target and driving their evolution, high-quality molecular structure information is generated, and the problem of difficult to effectively design drug molecules in the prior art is solved, and efficient and excellent molecular generation is achieved.
Patent Information
- Application Number
- CN202211191747.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-28
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2042-09-28
AI Technical Summary
The prior art is difficult to effectively use protein target structure information for drug molecules design, especially when facing complex protein target structures, design is difficult and it is difficult to fully consider the complex chemical information within the protein target.
By obtaining protein target structure information, extracting three-dimensional data of the cavity, and setting up virtual particle objects in the cavity, driving virtual particles to move and evolve through the model, generating molecular objects with specific spatial structures and atomic types, and finally obtaining high-quality molecular structure information through fine-tuning optimization.
It realizes the rapid generation of high-quality molecular information, improves the efficiency and quality of molecular design, and can consider the impact of protein target structure information globally.
Smart Images

Figure CN115831251B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of information processing technology, and in particular to an information processing method based on protein target structure information. Background Art
[0002] The basis of drug effect is the interaction between drug molecules and protein targets. Therefore, the key to structure-based drug design is to find or design drug molecules with strong binding affinity to protein targets.
[0003] There are two ways to design drug molecules based on protein targets: high-throughput screening of molecular libraries, which is limited by the inherent size of the molecular library. A large molecular library will increase the cost of screening and reduce efficiency. Another way is to rationally design based on the protein target structure. When faced with a complex protein target structure, design is very difficult.
[0004] Rational molecular design is often limited by the expert experience of medicinal chemists. In addition, these methods are limited by the definition of pharmacophores by computer scientists, making it difficult to fully consider the complex chemical information within the protein target.
[0005] Therefore, there is currently no effective method for drug molecule design based on protein target structure information. Summary of the invention
[0006] The purpose of the present invention is to address the defects of the prior art and provide an information processing method based on protein target structure information. The three-dimensional structure information of the protein target is used as a basis to perform comprehensive and sufficient information processing on it without information loss, so that the molecular information generated is fast and the efficiency is improved; in addition, the quality of molecular generation and optimization can be improved.
[0007] To this end, in a first aspect, an embodiment of the present invention provides an information processing method based on protein target structure information, the method comprising:
[0008] Step 1, obtaining protein target structure information, and extracting three-dimensional cavity data of the protein target cavity;
[0009] Step 2: setting a virtual particle object, wherein the virtual particle object includes a spatial attribute, and the spatial attribute includes particle three-dimensional data of the virtual particle;
[0010] Step 3: according to the spatial coordinates corresponding to the three-dimensional data of the cavity and the three-dimensional data of the particle, the virtual particles are evenly and continuously arranged in the cavity, so that there are a plurality of the virtual particles in the cavity;
[0011] Step 4: according to the first molecular structure model, different data settings are performed on the spatial coordinate attributes and the atomic type attributes in the virtual particle object of the virtual particle at each spatial coordinate; different stable position coordinates and different stable atomic type attributes of the virtual particle at the corresponding spatial coordinate are selected as the attributes of the virtual particle object after the corresponding virtual particle is in a stable state;
[0012] Step 5: clustering virtual particles with similar spatial coordinates based on spatial coordinate attributes and atomic type attributes to generate multiple atomic objects with specific spatial coordinate attributes and specific atomic models;
[0013] Step 6, subjecting the plurality of atomic objects to fine-tuning and optimization processing as a whole, thereby obtaining a molecular object with spatial structure information;
[0014] Step 7, obtaining multiple molecular objects with corresponding spatial structure information through steps 1-6, and performing confidence assessment processing on the spatial structure information of the multiple molecular objects by using the model to evaluate the overall structure information of the protein target-molecule complex, to obtain corresponding assessment processing result data;
[0015] Step 8: sorting the evaluation result data corresponding to the different molecular objects to obtain the molecular structure information corresponding to the molecular objects with better sorting order.
[0016] Furthermore, the first molecular structure model in step 4 includes protein target structure information and virtual particle object information; wherein the protein target structure information comprises atomic information of the protein target, or amino acid information of the protein target.
[0017] Furthermore, the basic framework of the first molecular structure model in step 4 is a Transformer, a graph neural network or a convolutional neural network architecture, or a molecular dynamics framework of a molecular force field.
[0018] Furthermore, the step 4 specifically includes: according to the first molecular structure model, performing multiple rounds of iterative data setting of the spatial coordinate attributes and the atomic type attributes in the virtual particle object of the virtual particle at each spatial coordinate, until the spatial coordinate attributes of each virtual particle object no longer change or vibrate within a fixed small area, and the atomic type attributes no longer change; selecting different stable position coordinates and different stable atomic type attributes of the virtual particles at the corresponding spatial coordinates as the attributes of the virtual particle object after the corresponding virtual particle is in a stable state.
[0019] Furthermore, the overall structural information of the protein target-molecule complex in step 7 is specifically the crystal structure information of the protein target-molecule complex obtained by experimental analysis, or the virtual data information of the protein target-molecule complex obtained based on molecular docking and molecular dynamics.
[0020] The information processing method based on protein target structure information provided by the embodiment of the present invention can use the three-dimensional structure information of the protein target as a basis to perform comprehensive, sufficient and information-free information processing, and the generated molecular information obtained is of high quality, so that it can be effectively used for molecular design, and the efficiency of molecular generation is improved; and the quality of molecular generation and optimization is improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 It is a flow chart of the information processing method based on protein target structure information of the present invention;
[0022] Figure 2 A schematic diagram of a model based on a Transformer architecture in the information processing method based on protein target structure information of the present invention;
[0023] Figure 3 Schematic diagram of the evolution of the information processing method based on protein target structure information of the present invention. DETAILED DESCRIPTION
[0024] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments.
[0025] Because the existing rational molecular design is often limited by the expert experience of medicinal chemists, but this method is limited by the definition of pharmacophore by computer scientists, it is difficult to fully consider the complex chemical information in the protein target. The present invention provides an artificial intelligence method, a method for designing druggable molecules based on information on the protein target structure. It can be called a "virtual dynamics" method, that is, covering virtual particles in the cavity of the protein target, driving the virtual particles to move to the target position and evolve the target type through the model, and finally merging and fine-tuning to obtain the process of real molecules.
[0026] Figure 1 This is a flow chart of the information processing method based on protein target structure information of the present invention. As shown in the figure, the embodiment of the present invention includes the following steps:
[0027] Step 101, obtaining protein target structure information, and extracting three-dimensional cavity data of the protein target cavity therefrom;
[0028] According to the protein target to be processed as needed, the target structure information of the protein target is obtained, so that the cavity and position of the protein target cavity can be obtained, that is, the three-dimensional cavity data of the protein target cavity.
[0029] There are three methods for obtaining targets: knowledge-based, protein experimental structure-based, and chemical informatics analysis-based. In general, the first step is to find out which amino acids make up the target, then correspond to the protein structure, and observe the cavity structure surrounded by these amino acids.
[0030] Step 102: setting a virtual particle object, the virtual particle object including space attributes, the space attributes including the three-dimensional particle data of the virtual particle;
[0031] Virtual particle objects are abstract data information for subsequent processing of virtual particles, including spatial attributes, from which the three-dimensional data of particles can be obtained for subsequent processing. In addition to the actual spatial coordinates, virtual particles can also be probability densities in space to evolve their positions and types.
[0032] Step 103, according to the spatial coordinates corresponding to the cavity three-dimensional data and the particle three-dimensional data, the virtual particles are evenly and continuously arranged in the cavity, so that there are multiple virtual particles in the cavity;
[0033] The significance of the above steps is to place a large number of "virtual particles" in the protein target cavity, so that the virtual particles fill the entire protein target cavity. These virtual particles have no atomic type information, but have spatial coordinate information, and the spatial coordinates are evenly distributed in the protein target cavity.
[0034] Step 104, according to the first molecular structure model, different data settings are performed on the spatial coordinate attributes and atomic type attributes in the virtual particle object of the virtual particle under each spatial coordinate; different stable position coordinates and different stable atomic type attributes of the virtual particles under the corresponding spatial coordinates are selected as the attributes of the virtual particle object after the corresponding virtual particle is in a stable state; in order to ensure that the molecules composed of virtual particles can better consider the entire protein target structure, the virtual particles should fully cover the entire spatial structure of the protein target, and ensure that particles can move to all positions in the entire space that the model considers to be important.
[0035] The above steps are the training process, that is, the de novo molecule generation process, in which the first molecular structure model refers to a model that can take into account an existing structure. This existing structure model is generally obtained through a large number of protein target-drug molecule complex structure training; for example, in this de novo molecule generation process, only the protein target structure information can be considered; in the subsequent molecular optimization process, the protein target structure information and the existing molecular fragment information are considered. Then, the spatial coordinate object information and the atomic type object information of the virtual particle object corresponding to the virtual particle are continuously evolved, so that each virtual particle is finally stabilized in a specific position and evolved into a specific atomic type.
[0036] The protein target structure can be represented by the atoms that make up the protein, or by the amino acid sequence, amino acid residue coarse-grained particles, amino acid main chain and side chain coarse-grained particles, etc. At the same time, the protein target can include multiple structures such as peptide chains, DNA, RNA, confactors, metal ions, etc.
[0037] During model training, training objectives can be considered hierarchically, including but not limited to particle coverage of important positions, most particle types at target positions being consistent with real particle types, and the molecular structure obtained through analysis being consistent with the target molecular structure.
[0038] The first molecular structure model includes protein target structure information and virtual particle object information; wherein the protein target structure information comprises atomic information of the protein target, or amino acid information of the protein target.
[0039] Specifically, the input of the first molecular structure model consists of a protein target structure and virtual particles, which include the type and position information of atoms to be retained in the molecular optimization task, as well as the protein target structure. The protein target structure can be input into the model using the atomic information constituting the protein target or the amino acid information constituting the protein target as a representation method. Virtual particles and atoms to be retained can be input into the model using the basic information of the particles as a representation method.
[0040] The basic framework of the first molecular structure model is Transformer (please provide the corresponding Chinese translation), graph neural network or convolutional neural network architecture, or molecular dynamics framework of molecular force field. The model needs to realize the interaction between protein target information and virtual particle information, and drive the evolution of particle coordinates and types through protein target information. The interaction method can be based on neural network model, such as using hybrid attention mechanism, or based on traditional Newtonian mechanics.
[0041] Figure 2This is a schematic diagram of a model based on the Transformer architecture in the information processing method based on protein target structure information of the present invention. Specifically, the structural model mainly includes three parts: a pocket encoder, a virtual particle encoder, and a model main module. The specific functions of each component are as follows.
[0042] Pocket encoder, assuming there are N pockets of atoms, can be constructed as an N*N pocket atom pair matrix, and then the coordinate information contained in the atom pair matrix is continuous through the Gaussian kernel to obtain the pocket atom pair representation. The information of pocket atoms is represented through the pocket, and the self-attention layer and the pocket atom pair representation interact, and the information is further extracted through the multi-layer feedforward neural network layer. After multiple such modules, the final pocket representation information is obtained.
[0043] Virtual particle encoder, assuming there are M virtual particles, can construct an M*M virtual particle pair matrix, and then use the Gaussian kernel to make the coordinate information contained in the matrix continuous to obtain the virtual particle pair distance representation. The above N pocket atoms and M virtual particles are used to construct an N*M virtual particle-pocket atom pair distance matrix, and then use the Gaussian kernel to make the coordinate information contained in the matrix continuous to obtain the virtual particle-pocket atom pair distance representation. At the same time, the representation of the virtual particle is passed in.
[0044] The main module of the model is responsible for accepting virtual particle representations and pocket atom representations, and completing information interaction through the virtual particle-pocket atom attention layer. After multiple gates and multi-layer feedforward neural networks, the virtual particle-pocket atom pair representations and virtual atom pair representations are continuously updated. The output of the last layer of the neural network is the predicted value of the virtual particle coordinate information and atom type.
[0045] In addition, this step can be specifically to perform multiple rounds of iterative data setting of the spatial coordinate attributes and atomic type attributes in the virtual particle object of the virtual particle under each spatial coordinate according to the first molecular structure model, until the spatial coordinate attributes of each virtual particle object no longer change or vibrate within a fixed small area, and the atomic type attributes no longer change; and select different stable position coordinates and different stable atomic type attributes of the virtual particles under the corresponding spatial coordinates as the attributes of the virtual particle object after the corresponding virtual particle is in a stable state.
[0046] Specifically, the output of the model is the prediction of the type and coordinates of each virtual particle. The output can be completed in one step, or it can be divided into multiple rounds of iterations of the type and coordinates of the virtual particles until the position of each particle no longer moves or vibrates in a fixed small area, and the type of the particle no longer changes.
[0047] Step 105, clustering virtual particles with similar spatial coordinates based on spatial coordinate attributes and atomic type attributes to generate multiple atomic objects with specific spatial coordinate attributes and specific atomic models;
[0048] For example, several virtual particles of the same type that are close to each other will be clustered into one particle (atom), the coordinates of which are the geometric center of these particles, and the type of which is the type of these particles.
[0049] Specifically, several virtual particles at the same position are clustered based on spatial coordinates and types, and merged into several atoms with spatial coordinates and specific types.
[0050] Step 106, subjecting the plurality of atomic objects to fine-tuning and optimization processing as a whole, thereby obtaining a molecular object with spatial structure information;
[0051] Fine-tuning is still achieved through the "virtual dynamics" method. At this time, the "virtual particle" corresponds to the "atom" obtained in step 105. The coordinates and type are also optimized through the model.
[0052] The merged atoms and the atoms in the existing molecular fragments during the molecular optimization process are fine-tuned and optimized as a whole to form a molecule with a spatial structure.
[0053] Step 107, obtaining multiple molecular objects with corresponding spatial structure information through steps 101-106, performing confidence assessment processing on the spatial structure information of the multiple molecular objects by using the model for the overall structure information of the protein target-molecule complex, and obtaining corresponding assessment processing result data;
[0054] The overall structural information of the protein target-molecule complex specifically refers to the crystal structure information of the protein target-molecule complex obtained by experimental analysis, or the virtual data information of the protein target-molecule complex obtained based on molecular docking and molecular dynamics.
[0055] The model should be able to evaluate the results it produces, judge the quality of a batch of results through the overall recognition of the protein target-molecular structure and sort them accordingly, and finally output a batch of molecules with better results.
[0056] Step 108: sorting the evaluation result data corresponding to different molecular objects to obtain molecular structure information corresponding to molecular objects with better sorting order.
[0057] Specifically, multiple rounds of virtual dynamics are performed to obtain multiple molecular structures, and the confidence of the obtained molecular structures is evaluated by the model for the overall structure of the protein target-molecule complex, and the results are sorted, and finally a batch of molecular structures with better sorting are output.
[0058] Among them, the expansion of the application scenarios of the molecular generation model expands the application space of the molecular generation model to molecular generation and molecular optimization problems, and uses a unified model architecture to solve key problems in molecular design. The model construction method is to use the protein target structure and virtual particles as inputs in the model, and realize the effective interaction of the two parts of information in the model to drive the evolution of virtual particles to form a real molecular structure.
[0059] Figure 3 This is a schematic diagram of the evolution of the information processing method based on protein target structure information of the present invention. As shown in the figure, the gray grid structure is the protein target structure, the light-colored balls are virtual particles, the dark-colored balls are virtual particles whose coordinates and type information are updated during the evolution process, and the ball-and-stick structures are molecules.
[0060] In the figure, the upper part of the two leftmost pictures is a schematic diagram of the protein pocket, and the lower part is a schematic diagram of the virtual particles; then there are schematic diagrams of the virtual particles spread throughout the protein pocket; next is a schematic diagram of the virtual particles moving to a "stable position" and evolving into a specific type; then there is a schematic diagram of nearby virtual particles of the same type clustering and merging into one atom; and finally there is a schematic diagram of the atoms forming the final molecule through fine-tuning.
[0061] The information processing method based on protein target structure information of the embodiment of the present invention has the following advantages:
[0062] 1. By directly using the cavity three-dimensional data in the protein target structure information as input information, the three-dimensional structural information of the protein target can be used for information processing of molecular design in a comprehensive and sufficient manner without information compression or loss.
[0063] 2. The present invention can perform information processing at one time to obtain complete molecular structure information, and can globally consider the influence of protein target structure information, thus having high molecular generation quality.
[0064] 3. The method of the present invention can conduct confidence assessment on the results produced to obtain ranking results, thereby having higher quality of molecule generation.
[0065] 4. The present invention can use a unified model to carry out two types of molecular design methods: protein target-based molecular design and protein target-based molecular optimization.
[0066] The professionals should further realize that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in the above description according to the function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.
[0067] The steps of the method or algorithm described in conjunction with the embodiments disclosed herein may be implemented using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.
[0068] The specific implementation methods described above further illustrate the objectives, technical solutions and beneficial effects of the present invention in detail. It should be understood that the above description is only a specific implementation method of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. An information processing method based on protein target structure information, It is characterized in that The method comprises: Step 1, obtaining protein target structure information, and extracting three-dimensional cavity data of the protein target cavity; Step 2: setting a virtual particle object, wherein the virtual particle object includes a spatial attribute, and the spatial attribute includes particle three-dimensional data of the virtual particle; Step 3: according to the spatial coordinates corresponding to the three-dimensional data of the cavity and the three-dimensional data of the particle, the virtual particles are evenly and continuously arranged in the cavity, so that there are a plurality of the virtual particles in the cavity; Step 4: according to the first molecular structure model, different data settings are performed on the spatial coordinate attributes and the atomic type attributes in the virtual particle object of the virtual particle at each spatial coordinate; different stable position coordinates and different stable atomic type attributes of the virtual particle at the corresponding spatial coordinate are selected as the attributes of the virtual particle object after the corresponding virtual particle is in a stable state; Step 5: clustering virtual particles with similar spatial coordinates based on spatial coordinate attributes and atomic type attributes to generate multiple atomic objects with specific spatial coordinate attributes and specific atomic models; Step 6, subjecting the plurality of atomic objects to fine-tuning and optimization processing as a whole, thereby obtaining a molecular object with spatial structure information; Step 7, obtaining multiple molecular objects with corresponding spatial structure information through steps 1-6, and performing confidence assessment processing on the spatial structure information of the multiple molecular objects by using the model to evaluate the overall structure information of the protein target-molecule complex, to obtain corresponding assessment processing result data; Step 8: sorting the evaluation result data corresponding to the different molecular objects to obtain the molecular structure information corresponding to the molecular objects with better sorting order.
2. The method according to claim 1, It is characterized in that The steps are specifically as follows: for de novo molecules, protein target structure information is obtained, and the three-dimensional cavity data of the protein target cavity is extracted therefrom; for subsequent molecules, protein target structure information and the obtained molecular structure information are obtained, and the three-dimensional cavity data of the protein target cavity is extracted therefrom.
3. The method according to claim 2, It is characterized in that The first molecular structure model in step 4 includes protein target structure information and virtual particle object information; wherein the protein target structure information comprises atomic information of the protein target, or amino acid information of the protein target.
4. The method according to claim 1, It is characterized in that The basic framework of the first molecular structure model in step 4 is a Transformer, a graph neural network or a convolutional neural network architecture, or a molecular dynamics framework of a molecular force field.
5. The method according to claim 1, It is characterized in that The step 4 specifically includes: according to the first molecular structure model, performing multiple rounds of iterative data setting of the spatial coordinate attributes and the atomic type attributes in the virtual particle object of the virtual particle at each spatial coordinate, until the spatial coordinate attributes of each virtual particle object no longer change or vibrate within a fixed small area, and the atomic type attributes no longer change; and selecting different stable position coordinates and different stable atomic type attributes of the virtual particle at the corresponding spatial coordinate as the attributes of the virtual particle object after the corresponding virtual particle is in a stable state.
6. The method according to claim 1, It is characterized in that The overall structural information of the protein target-molecule complex in step 7 is specifically the crystal structure information of the protein target-molecule complex obtained by experimental analysis, or the virtual data information of the protein target-molecule complex obtained based on molecular docking and molecular dynamics.
Citation Information
Patent Citations
Drug small molecule-protein target reaction prediction method based on multi-dimensional information
CN112331273A
Method for predicting unknown small molecule ligand target spot and application of method
CN113963744A