A method and system for annotating non-covalent interactions in a molecular system

By acquiring atomic coordinates and electronic density grid information of molecular structures, determining saddle point positions and their attribution relationships, and combining multi-resolution and machine learning models, the problem of accurately identifying non-covalent interactions in macromolecular systems was solved, achieving efficient and accurate NCI labeling.

CN115938500BActive Publication Date: 2026-01-02BEIJING STONEWISE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110953574.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-08-19
Publication Date
2026-01-02
Estimated Expiration
2041-08-19

AI Technical Summary

Technical Problem

Existing technologies for identifying and predicting non-covalent interactions in macromolecular systems suffer from problems such as qualitative analysis being insufficient for quantitative analysis, high computational costs, and a lack of a global perspective. In particular, they have poor identification capabilities in complex environments and lack accurate prediction models.

Method used

By acquiring atomic coordinates and electronic density grid information of the target molecular structure, the location and assignment of saddle points are determined. Non-covalent interactions are labeled using electronic density topological features. Combined with labels under multiple resolution conditions, a machine learning model is used to improve recognition accuracy.

Benefits of technology

It enables accurate identification and quantitative analysis of non-covalent interactions in macromolecular systems, improving the accuracy and reliability of identification while reducing computational and time costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115938500B_ABST
    Figure CN115938500B_ABST
Patent Text Reader

Abstract

The application discloses a non-covalent interaction labeling method and a labeling system in a molecular system, and the method comprises the following steps: obtaining atomic coordinate information and preset resolution electron density grid point information of a target molecular structure; determining a saddle point position based on a grid point position where an electron density saddle point in the electron density grid point information is located; determining an attribution relationship of the saddle point position by using the atomic coordinate information; generating an electron density topological feature of the saddle point position by using at least an electron density, a first-order gradient and a Hessian matrix of the saddle point position; and labeling a non-covalent interaction in the target molecular structure by using the saddle point position, the attribution relationship and the electron density topological feature. The technical scheme provided by the application improves the accuracy of labeling non-covalent interactions in a molecule and a molecular system.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of structural biology analysis, and in particular to a non-covalent interaction labeling method and a labeling system in a molecular system. BACKGROUND

[0002] Non-covalent interaction (NCI) is different from covalent bond, which does not involve shared electrons, but involves more dispersed electromagnetic interaction changes between or within molecules. NCI is crucial for maintaining the three-dimensional structure of macromolecules such as proteins and nucleic acids. In addition, they are also involved in many biological processes in which macromolecules specifically but transiently bind to each other. These interactions also have a serious impact on drug design, crystallinity and material design, especially self-assembly, and the synthesis of many organic molecules, so the identification of NCI in protein structure has a strong application.

[0003] The existing technology summarizes several rules for judging whether a structure is NCI. However, using rules to identify NCI has the defect of being able to determine only the qualitative but not the quantitative, and the ability to identify NCI in a complex environment is poor. In addition to NCI structure identification, NCI strength prediction is also an effective basis for accurate judgment of NCI type. In the traditional method, the NCI prediction based on the molecular force field belongs to the empirical method, which depends on the extraction of rules, and can predict a few classic NCIs such as hydrogen bonds and salt bonds, but lacks a prediction model for non-classical NCIs, which cannot consider the influence of the surrounding environment on the strength of NCI; the method of calculating NCI strength based on quantum chemistry, although it has high accuracy, is only suitable for small molecule systems, and in the case of macromolecular systems, its calculation amount and global perspective are not up to standard, thereby leading to the failure of the method. Therefore, how to accurately identify NCI is a problem to be solved. SUMMARY

[0004] Therefore, the embodiments of the present application provide a non-covalent interaction labeling method and a labeling system in a molecular system, thereby improving the accuracy of labeling non-covalent interactions in a macromolecular system.

[0005] According to a first aspect, a non-covalent interaction labeling method in a molecular system, the method comprises:

[0006] obtaining atomic coordinate information of a target molecular structure and electron density grid point information of a preset resolution, the electron density grid point information being used to store the electron density of each grid point of the target molecular structure;

[0007] determine a saddle point position based on a grid point position where an electron density saddle point in the electron density grid point information, the saddle point position being used to represent a spatial position of a non-covalent interaction in the target molecular structure;

[0008] determine an attribution relationship of the saddle point position by using the atomic coordinate information, the attribution relationship being used to represent a matching relationship of a non-covalent interaction and a nearby atom pair;

[0009] generate an electron density topological feature of the saddle point position by using at least an electron density, a first order gradient and a Hessian matrix of the saddle point position, the electron density topological feature being used to describe an attribute of a non-covalent interaction;

[0010] label a non-covalent interaction in the target molecular structure by using the saddle point position, the attribution relationship and the electron density topological feature.

[0011] Optionally, the electron density grid point information with a preset resolution is obtained, including:

[0012] convert the atomic coordinate information from a real space to a frequency domain space to generate a structure factor;

[0013] adjust the structure factor back to the real space with a preset resolution to generate the electron density grid point information, the preset resolution being a preset summation range of a frequency domain space vector component.

[0014] Optionally, the determination of the saddle point position based on the grid point position where the electron density saddle point in the electron density grid point information includes:

[0015] divide the electron density grid point information into a plurality of sub-information blocks with a first preset range size;

[0016] obtain a ligand-receptor atom pair in the sub-information block, and cover a range with a preset length as a radius and a midpoint position of a connection line of the ligand-receptor atom pair as a center as a candidate range;

[0017] calculate a reduced electron density gradient value of each grid point in the candidate range, and screen out a grid point with a minimum reduced electron density gradient value as a candidate saddle point;

[0018] calculate a reduced electron density gradient value of a neighboring grid point of the candidate saddle point, and take a grid point with a minimum reduced electron density gradient value from the candidate saddle point and the neighboring grid point as the saddle point position.

[0019] Optionally, after the calculation of the reduced electron density gradient value of the neighboring grid point of the candidate saddle point, and the taking of the grid point with the minimum reduced electron density gradient value from the candidate saddle point and the neighboring grid point as the saddle point position, the method further includes:

[0020] discard the saddle point position where the eigenvalues of the electron density Hessian matrix do not satisfy λ1<λ2<0<λ3.

[0021] Optionally, after the saddle point position where the eigenvalues of the electron density Hessian matrix do not satisfy λ1<λ2<0<λ3, the method further comprises:

[0022] When one of the saddle point positions corresponds to multiple ligand-receptor atom pairs, the atom pair with the closest distance between the ligand and the receptor is taken as the target atom pair to be labeled.

[0023] Optionally, after the saddle point position where the eigenvalues of the electron density Hessian matrix do not satisfy λ1<λ2<0<λ3, the method further comprises:

[0024] When the distance between any two of the saddle point positions is less than a preset distance, the saddle point position with a smaller electron density value is discarded.

[0025] Optionally, the preset resolution comprises a plurality of preset resolutions with different sizes, and the method further comprises:

[0026] Under each resolution, the steps of obtaining the atomic coordinate information of the target molecular structure and the electron density grid point information of the preset resolution, and labeling the non-covalent interaction in the target molecular structure according to the saddle point position, the attribution relationship and the electron density topological feature are performed respectively.

[0027] According to a second aspect, a non-covalent interaction labeling system in a molecular system, the system comprises:

[0028] An information acquisition module is configured to obtain atomic coordinate information of a target molecular structure and electron density grid point information of a preset resolution, the electron density grid point information being used to store the electron density of each grid point of the target molecular structure;

[0029] A position acquisition module is configured to determine a saddle point position based on the grid point position where the electron density saddle point is located in the electron density grid point information, the saddle point position being used to represent the spatial position of the non-covalent interaction in the target molecular structure;

[0030] An attribution matching module is configured to determine the attribution relationship of the saddle point position by using the atomic coordinate information, the attribution relationship being used to represent the matching relationship between the non-covalent interaction and the nearby atom pairs;

[0031] A feature acquisition module is configured to generate the electron density topological feature of the saddle point position by using at least the electron density, the first-order gradient and the Hessian matrix of the saddle point position, the electron density topological feature being used to describe the type and attribute of the non-covalent interaction;

[0032] a non-covalent interaction labeling module configured to label a non-covalent interaction in the target molecular structure according to the saddle point position, the attribution relationship, and the electron density topological feature.

[0033] According to a third aspect, an electronic device comprises:

[0034] a memory and a processor in communication connection with each other, and the memory has stored computer instructions, and the processor executes the computer instructions to perform the method in the first aspect or any one of the optional implementation manners of the first aspect.

[0035] According to a fourth aspect, an embodiment of the present application provides a computer readable storage medium, characterized in that the computer readable storage medium stores computer instructions, and the computer instructions are used to make the computer execute the method in the first aspect or any one of the optional implementation manners of the first aspect.

[0036] The technical scheme of the present application has the following advantages:

[0037] The method and system for labeling non-covalent interaction in a molecular system provided by the embodiments of the present application. The method divides the spatial structure of a macromolecule-ligand into multiple grid points, and calculates the electron density of each grid point based on crystallography theory to obtain more reliable electron density grid point information. According to the properties of the electron density itself, it is known that the electron density saddle point is almost the same as the position of the non-covalent interaction. Then, the grid point representing the electron density saddle point is found by the electron density gradient and other features in each grid point. The spatial coordinates of the grid point position are taken as the spatial coordinates of the NCI, and the electron density topological features calculated by the grid point are taken as the topological features of the NCI. Moreover, by dividing the grid points of the macromolecule-ligand structure at multiple resolutions, the NCI labeling results under multiple resolution conditions are obtained, so that the NCI of the target molecular structure is analyzed from the micro and macro perspectives, and the accuracy and reliability of NCI identification are further improved. BRIEF DESCRIPTION OF DRAWINGS

[0038] In order to more clearly illustrate the specific embodiments of the present application or the technical solutions in the prior art, the drawings needed in the specific embodiments or prior art description will be briefly introduced below. Obviously, the drawings in the following description are some embodiments of the present application, and those skilled in the art can also obtain other drawings according to these drawings without creative labor.

[0039] Figure 1 A schematic diagram of the steps of the method for labeling non-covalent interaction in a molecular system according to an embodiment of the present application;

[0040] Figure 2 A CP point structure schematic diagram in a molecular system non-covalent interaction annotation method of an embodiment of the present application;

[0041] Figure 3 An execution step schematic diagram of a molecular system non-covalent interaction machine learning model of an embodiment of the present application;

[0042] Figure 4 A training sample dimensionality mapping structure schematic diagram of a molecular system non-covalent interaction machine learning model of an embodiment of the present application;

[0043] Figure 5 A structure schematic diagram of a molecular system non-covalent interaction annotation system of an embodiment of the present application;

[0044] Figure 6 A structure schematic diagram of an electronic device of an embodiment of the present application. DETAILED DESCRIPTION

[0045] In order to make the objects, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are some but not all of the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the protection scope of the present application.

[0046] The technical features involved in the different embodiments of the present application described below can be combined with each other as long as there is no conflict.

[0047] As shown in the drawings, the present application provides a molecular system non-covalent interaction annotation method, which specifically comprises the following steps: Figure 1

[0048] ​Step S1: Obtain atomic coordinate information of a target molecular structure and electron density grid point information of a preset resolution, and the electron density grid point information is used to store the electron density of each grid point of the spatial structure information. Specifically, in the field of biological structures, NCI is a kind of weak interaction that is essential to maintain the three-dimensional structure of molecules such as proteins and nucleic acids, and NCI also participates in many biological processes in which molecules are specifically but transiently combined with each other. These interactions can also seriously affect scenarios such as drug design, crystallinity, and material design, so accurately labeling the NCI in the molecular system (including the interior of the molecule and between the molecules) of the protein plays a crucial role in the research and design process. In the traditional technology, a number of rules are summarized to identify NCI, such as the distance and angle between two atoms, and this way has the defect of being qualitative rather than quantitative, and has poor NCI identification capability in complex environments. The electron density represents the probability of finding electrons at a specific position around an atom or molecule, and because of the difference in the strength of the force between NCI and covalent bond, the distribution of its electron density has a certain regularity. Through a large number of experiments, the position of the electron density saddle point is very close to the position of the weak interaction, and if the electron density saddle point is used to label NCI, the accuracy of NCI identification in the target molecular structure can be greatly improved. The labeling method of non-covalent interaction in a molecular system provided in the embodiment of the application can accurately identify the electron density saddle point in the large molecular-ligand molecular system structure.

[0049] Step S2: Determine the saddle point position based on the grid point position of the electron density saddle point in the electron density grid point information, and the saddle point position is used to represent the spatial position of the non-covalent interaction in the target molecular structure. Specifically, searching for the position of the electron density saddle point in the electron density grid point information obtained in the above step S1 can obtain the position coordinates of the NCI.

[0050] Step S3: Determine the attribution relationship of the saddle point position by using the atomic coordinate information, and the attribution relationship is used to represent the matching relationship between the non-covalent interaction and the nearby atom pair. After the spatial coordinates of the saddle point position are determined, the position coordinates of the NCI in the space are determined, but it is still necessary to comprehensively analyze and judge according to the atomic coordinate information, the attributes of the NCI, and the distance of each atom pair to determine which atom pair in the molecule the NCI specifically belongs to.

[0051] Step S4: generating the electron density topological features of the saddle point position by using at least the electron density, the first order gradient and the Hessian matrix of the saddle point position, the electron density topological features being used to describe the properties of the non-covalent interaction. Specifically, the NCI can be further divided into electrostatic interaction, π-effect, van der Waals force and hydrophobic effect, etc., and different NCIs have different commonly used description methods, for example: the π-effect is usually described by the electron density, but the van der Waals force is usually described by the Lagrangian kinetic energy density. The commonly used electron density topological features for describing the properties and / or types of NCI include but are not limited to: electron density, Lagrangian kinetic energy density, Hamiltonian kinetic energy density, potential energy density, energy density, electron density Laplacian and electrostatic potential. All the above commonly used electron density topological features can be calculated by using at least the electron density, the first order gradient and the Hessian matrix of the saddle point position. Therefore, after the saddle point position is determined, the desired electron density topological features of the saddle point position can be obtained at any time according to the actual needs of those skilled in the art, and the corresponding NCI is labeled. The electron density topological features can be used to describe the properties of the NCI in detail, wherein the properties include but are not limited to: the presence or absence of the NCI, the type, the direction, the strength and the attribution of the NCI, thereby improving the convenience of NCI labeling.

[0052] Step S5: labeling the non-covalent interaction in the target molecular structure by using the saddle point position, the attribution relationship and the electron density topological features.

[0053] Specifically, in an embodiment, the above step S1 further comprises the following steps:

[0054] Step S101: converting the atomic coordinate information from the real space to the frequency domain space to generate the structure factor.

[0055] Step S102: adjusting the structure factor to return to the real space at a preset resolution to generate the electron density grid information, the preset resolution being a preset summation range of the vector components of the frequency domain space.

[0056] Specifically, in order to obtain the electron density saddle points, firstly, the electron density grid information representing the electron density distribution in the target molecular structure (single molecular structure or molecular structure composed of multiple molecules) needs to be known, that is, the molecular electron density map under the preset resolution condition is obtained, the electron density map is divided into multiple grids in the spatial domain, and the electron density value in each grid is obtained. Thus, the position of the electron density saddle point can be found through the electron density value of each grid. The method for generating the electron density grid information includes but is not limited to experimental methods and calculation methods, and the commonly used methods include: a method for obtaining electron density based on X-ray crystallography experiment, a method for obtaining electron density based on electron microscope experiment, and a method for obtaining electron density based on quantum chemistry calculation. The traditional X-ray crystallography can analyze the fine three-dimensional structure of a macromolecule at the atomic or near-atomic level, so that the electron density grid information is calculated based on the information of the macromolecular structure obtained by experiment; the experimental electron density obtained by the electron microscope is directly saved as a spatial point array file, which can be directly read; in the case where the experimental conditions are not available, the electron density is calculated based on the atomic spatial coordinates of the target molecular structure which are very easy to obtain, and the target molecular structure is divided into multiple small parts, that is, multiple small molecular systems, and the electron density value of each grid is calculated under the small molecular system.

[0057] Specifically, in an embodiment, among the above-mentioned methods for obtaining electron density, on the one hand, due to the limitation of conditions, the experimental conditions cannot be achieved in all laboratories, and on the other hand, the electron density obtained based on quantum chemistry calculation pays too much attention to small molecular systems, and often ignores the connection between small molecular systems in macromolecular systems, so the calculation result is often wrong from a macroscopic point of view. Based on this, the labeling method provided by the present application provides a preferred scheme for obtaining electron density: a method for calculating electron density based on crystallography theory, firstly, the atomic coordinate information is converted from the real space (physical space) to the frequency domain space, and the diffraction of the crystal lattice to the electromagnetic wave obtains the pattern in the frequency domain space. The frequency domain space not only continues the symmetry of the crystal lattice in the real space and the information of the molecular structure and the physicochemical properties, but also decomposes the “detail degree” of the molecule in the real space according to different “frequencies”, that is, the information with high frequency in the frequency domain space reflects the details of the molecule in the real space, and the information with low frequency reflects the outline of the molecule in the real space. After filtering the information in the frequency domain space according to the frequency, the information is transformed back to the real space, and the molecular representation with “different detail degrees” is obtained. Specifically, the molecular coordinate information in the real space is subjected to Fourier transform to obtain the converted result in the frequency domain space (referred to as structure factor in the embodiment of the present application), and then inverse Fourier transform is performed on the frequency domain space vector with a preset frequency range to return to the real space. The value obtained by converting the molecular coordinate information in the real space to the frequency domain space and then converting back to the real space is the electron density grid information.

[0058] Specifically, in an embodiment, for the above step S101, the atomic coordinate information is converted from real space to frequency domain space to generate structure factors in a Fourier transform-based manner. First, the atomic coordinate information of the target molecular structure is obtained, which is a set of spatial coordinates of a large number of atoms in the target molecular structure. Then, the atomic coordinate information is converted from real space to frequency domain space to generate structure factors using the Fourier transform formula. In the embodiment of the present application, the structure factors are calculated from the atomic coordinate information using the Fourier transform, and the specific conversion formula is as follows:

[0059]

[0060] wherein r represents the atomic coordinate information, x, y, and z represent three components of the atomic coordinate information, f n represents the diffraction factor of n different elements, s represents the frequency domain space vector, h, k, and l represent three components of the frequency domain space vector, F(s) represents the structure factor of the s vector (including diffraction amplitude and phase information), and 2πi is an imaginary number; in addition, the above formula can be improved by introducing a description of the solvent and a description of the vibration of the atoms over time, so that the conversion result is more accurate, and the specific improved formula is as follows:

[0061]

[0062] wherein f(bulk solvent) is a solvent shell description function, bulk solvent is a solvent shell, f(b) is a temperature factor description function, b is a temperature factor, the temperature factor is used to measure the degree of atomic thermal motion in the crystal, and the solvent shell describes the contribution of the solvent to the diffraction.

[0063] Specifically, in an embodiment, for the above step S102, the structure factors are adjusted to return to real space at a preset resolution to generate electron density grid information. The specific calculation formula is as follows:

[0064]

[0065] Wherein, r represents atomic coordinate information, s represents a frequency domain space vector, h, k, and l respectively represent three components of the frequency domain space vector, F(s) is a structure factor of the s vector, p(r) represents electron density grid information, Vcell represents a unit cell volume, 2p is an imaginary number, and the preset resolution is obtained from a preset summation range of h, k, and l. Specifically, the preset resolution is obtained from a summation range of components of the preset frequency domain space vector, that is, by changing the summation range of h, k, and l, electron densities of different resolutions can be obtained. The higher the resolution, the clearer the obtained electron density map, and the more accurate the representation of the bond between atoms; the lower the resolution, the more blurred the obtained electron density map, but the overall framework of the molecule is more accurately represented. Based on the method of calculating electron density according to crystallography theory, the macroscopic relationship between atoms in a large molecule system is considered more, and by adjusting different preset resolutions, electron density maps of different levels of detail are obtained, so that electron density maps under various resolution conditions can be obtained. Then, the target molecule structure is divided into a preset number of grid forms, and the x, y, and z coordinates of the grid are brought into the electron density grid information function p(r), so that the electron density map is correspondingly divided into a discrete grid form, and the electron density grid information is obtained. The final saddle point positions obtained by different resolution electron density grid information are also different, so that by analyzing and comparing the saddle point positions under high resolution and low resolution, the purpose of noise elimination can be achieved from the macroscopic and microscopic angles, and the identification result is more accurate.

[0066] Specifically, in an embodiment, the step S2 further includes:

[0067] Step S202: dividing the electron density grid information into a plurality of sub-information blocks of a first preset range size. Specifically, in an embodiment, in order to improve the search accuracy, the electron density grid information is first divided into a plurality of sub-information blocks of a first preset range size (in this embodiment, the range of 10A is used to divide the electron density grid information) in the spatial coordinate system, and the saddle point position is searched in each sub-information block to ensure the search accuracy.

[0068] Step S203: obtaining a ligand-receptor atom pair in the sub-information block, and taking a range covered by a midpoint position of a connecting line of the ligand-receptor atom pair as a center and a preset length as a radius as a candidate range. Specifically, according to the characteristics of the electron density saddle point, the position is closer to the middle position of the two atoms of the atom pair, so all ligand-receptor atom pairs in the sub-information block are obtained, and a range covered by the midpoint position of the connecting line of the atom pair as a center and a preset length as a radius is taken as the candidate range.

[0069] ​​Step S204: calculating the reduced electron density gradient value of each grid point in the candidate range, and screening out the grid point with the minimum reduced electron density gradient value as the candidate saddle point. In the electron density distribution, combined with the properties of mathematical saddle points, the gradient of the saddle point is 0, and due to the characteristics of discrete space, there is no point with an absolute gradient of 0, so the reduced electron density gradient (RDG, Reduced electron Density Gradient) of each grid point in the candidate range is calculated, and the grid point with the minimum reduced electron density gradient value is screened out as the candidate saddle point.

[0070] Step S205: calculating the reduced electron density gradient value of the adjacent grid point of the candidate saddle point, and taking the grid point with the minimum reduced electron density gradient value among the candidate saddle point and the adjacent grid point as the saddle point position. Specifically, if the candidate saddle point falls on the edge of the candidate range, the saddle point is likely to be the minimum gradient point nearby, and since the points outside the edge are not compared, but are very close to the candidate saddle point, the points outside the edge are likely to be the electron density gradient saddle point of the target atomic pair, so the reduced electron density gradient value of the adjacent grid point of the candidate saddle point is calculated, and the grid point with the minimum reduced electron density gradient value among the candidate saddle point and the adjacent grid point is taken as the saddle point position, ensuring the accuracy of the saddle point position.

[0071] Specifically, in an embodiment, after the above step S205, the following steps are further included:

[0072] Step S206: abandoning the saddle point position whose eigenvalues of the electron density Hessian matrix do not satisfy λ1<λ2<0<λ3. Specifically, since the electron density saddle point is not unique, the electron density can either decrease in two perpendicular directions in space and increase in the third direction, or increase in two perpendicular directions in space and decrease in the third direction. In order to avoid confusion caused by the simultaneous statistics of multiple saddle points, as shown in Figure 2 , the embodiment of the present application adopts the electron density saddle point [(3,-1) CP point] whose electron density decreases in two perpendicular directions in space and increases in the third direction to mark the NCI position. That is, the saddle point whose electron density is maximum in two spatial directions and minimum in the third direction. Therefore, the eigenvalues of the Hessian matrix of the point are required to be two positive and one negative. Therefore, a preferred method is to further screen the saddle point position to obtain the required (3,-1) CP point, and abandon the saddle point position whose eigenvalues of the electron density Hessian matrix do not satisfy λ1<λ2<0<λ3.

[0073] Specifically, in an embodiment, after the above step S206, step S3 includes the following steps:

[0074] Step S207: When one saddle point position corresponds to multiple ligand-receptor atom pairs, the atom pair with the closest ligand and receptor is taken as the target atom pair to be labeled. Specifically, since one saddle point position can only label one atom pair of non-covalent interaction, in the embodiments of the present application, when one saddle point position corresponds to multiple ligand-receptor atom pairs, the atom pair with the closest ligand and receptor is taken as the target atom pair to be labeled by the saddle point position.

[0075] Specifically, in an embodiment, after step S207 described above, the following steps are further included:

[0076] Step S208: When the distance between any two saddle point positions is less than a preset distance, the saddle point position with a smaller electron density value is discarded. Specifically, when the distance between two saddle point positions is too close, it is empirically judged that one of the two saddle point positions is probably an error point, and therefore in the embodiments of the present application, when the distance between any two saddle point positions is less than a preset distance, the saddle point position with a smaller electron density value is discarded. After that, the saddle point positions obtained through step S2 can be used as position labels of spatial structure information in the model training process.

[0077] Specifically, in an embodiment, the saddle point positions and the electron density topological features of the target molecular structure obtained through steps S1-S5 described above can be used as a complete process to realize the identification of NCI in the target molecular structure. However, the calculation process described in steps S1-S5 requires a large amount of capital and time cost, and for long-time and large-amount NCI identification, the cost consumption is immeasurable. Based on this, a preferred solution, the embodiments of the present application further propose a method for identifying non-covalent interaction in a molecular system based on a machine learning model, which establishes an identification model through a large amount of data in the early stage, and solves the problem of long-time and large-cost consumption in the subsequent identification process. As shown in FIG. 8, the specific steps are as follows: Figure 3

[0078] ​Step S6: based on the target molecular structure, a training sample is established, the training sample includes spatial structure information of the target molecular structure and a label corresponding to the spatial structure information, the label includes a position label and a topological feature label, the position label is used to mark the spatial position of the non-covalent interaction in the target molecular structure, and the topological feature label is used to mark the attribute of the non-covalent interaction in the target molecular structure. Specifically, the training sample required for model training is established by the method described in steps S1-S5, the spatial structure information of the target molecular structure includes atomic coordinate information, so that the obtained spatial structure information can be used for calculation in steps S1-S5, and the spatial structure information is divided into grid points with the same preset resolution as the grid point information of the electron density, so that when a plurality of preset resolutions are used to establish the training sample, for one target molecular structure, a plurality of labels under different resolutions can be obtained, and the recognition result is enriched. Then, the position coordinates of the (3,-1) CP point are used as the position label of the spatial structure information, and all the calculated electron density topological features are used as the topological feature label of the spatial structure information, so that when the machine learning model is used for recognition, the spatial structure information of a target molecular structure is input, the recognition result can directly determine the NCI position, and various topological feature values are output for each NCI, and a person skilled in the art can directly select any result according to actual needs, which greatly facilitates experimental research. In an embodiment of the present application, 1w protein complexes in the pdbbind data set are used to construct samples, and 20w corresponding (3,-1) CP point position labels and topological feature labels of the topological features of the electron density topological features of each point are obtained.

[0079] Specifically, in an embodiment, before the above-mentioned training sample is constructed, in order to improve the calculation efficiency in the training process, the target molecular structure is standardized and posture corrected, and the structure is rotated to a standard posture with the ligand atom position at the (0, 0, 0) origin and the receptor atom position on the Z axis.

[0080] Step S7: using the spatial structure information as input, the initial machine learning model is used to predict the non-covalent interaction in the target molecular structure, and the initial machine learning model is corrected according to the error between the prediction result and the label, to obtain a recognition model, which is used for recognition of the position and attribute of the non-covalent interaction in the target molecular structure.

[0081] Specifically, since the saddle point position is actually the grid coordinates of the electron density grid information, in order to facilitate the processing of the training samples in the form of grids, the embodiments of the present application adopt a convolutional neural network (CNN) to establish the recognition model. In a preferred scheme, considering that the target molecular structure has various chemical characteristics, such as atomic types, aromaticity, pre-training features, etc. Under different feature conditions, the values in the three-dimensional matrix and the distribution of the values are different, therefore, as shown in Figure 4 the original space structure information is mapped from a three-dimensional feature vector to a four-dimensional feature vector representing different atomic features, increasing the training feature dimension, so that the training result is more accurate.

[0082] Finally, the training samples are continuously input into the CNN, and the CNN model is corrected according to the error between the prediction result output by the CNN and the pre-labeled label. After the training is completed, the trained recognition model is obtained, and according to the saddle point position and the electron density topological feature output by the model, the NCI position and attribute in the target molecular structure can be recognized. And an embedding vector representing the local feature of the NCI can be obtained in the full connection layer of the CNN. And through the training samples of different resolutions obtained through the S1-S5 steps, the CNN model can directly obtain the recognition results of different resolutions for the spatial structure information of the same target molecular structure. According to the difference between different recognition results, it is convenient for the person skilled in the art to further analyze the noise in the NCI.

[0083] By performing the above steps, the present application provides a method for labeling non-covalent interactions in a molecular system. The method divides the spatial structure of the macromolecule-ligand into multiple grids, and calculates the electron density of each grid based on crystallography theory to obtain more reliable electron density grid information. According to the properties of the electron density itself, it is known that the electron density saddle point is almost the same as the position of the non-covalent interaction, and then the grid point representing the electron density saddle point is found through the electron density gradient and other features in each grid point. The spatial coordinates of the grid point are taken as the spatial coordinates of the NCI, and the electron density topological features calculated by the grid point are taken as the topological features of the NCI. And by dividing the grids of the macromolecule-ligand structure at multiple resolutions, NCI labeling results under multiple resolution conditions are obtained, thereby realizing the analysis of the NCI of the target molecular structure from the micro and macro perspectives, and further improving the accuracy and reliability of the NCI recognition. In addition, by collecting the three-dimensional spatial structure of the target molecular structure, and taking the position of the NCI and the NCI topological features in the structure as the labels corresponding to the target molecular structure, a training sample is created, and then combined with the convolutional neural network algorithm, a machine learning model capable of accurately recognizing the NCI in the macromolecule-ligand is trained, thereby improving the recognition accuracy and efficiency of the NCI.

[0084] As Figure 5 shown, the embodiment also provides a labeling system for non-covalent interaction in a molecular system, which comprises:

[0085] The information collection module 101 is configured to acquire atomic coordinate information of a target molecular structure, and generate electronic density grid point information of a preset resolution according to the atomic coordinate information, the electronic density grid point information being used to store electronic density of each grid point of the target molecular structure. For details, refer to the related description of step S1 in the method embodiment, which will not be repeated here.

[0086] The position acquisition module 102 is configured to determine a saddle point position based on a grid point position where an electronic density saddle point in the electronic density grid point information is located, the saddle point position being used to represent a spatial position of the non-covalent interaction in the target molecular structure. For details, refer to the related description of step S2 in the method embodiment, which will not be repeated here.

[0087] The attribution matching module 103 is configured to determine an attribution relationship of the saddle point position by using the atomic coordinate information, the attribution relationship being used to represent a matching relationship between the non-covalent interaction and nearby atoms. For details, refer to the related description of step S3 in the method embodiment, which will not be repeated here.

[0088] The feature acquisition module 104 is configured to generate an electronic density topological feature of the saddle point position by using at least the electronic density, the first-order gradient and the Hessian matrix of the saddle point position, the electronic density topological feature being used to describe a type and an attribute of the non-covalent interaction. For details, refer to the related description of step S4 in the method embodiment, which will not be repeated here.

[0089] The non-covalent interaction labeling module 105 is configured to label the non-covalent interaction in the target molecular structure according to the saddle point position, the attribution relationship and the electronic density topological feature. For details, refer to the related description of step S5 in the method embodiment, which will not be repeated here.

[0090] The labeling system for non-covalent interaction in a molecular system provided by the embodiment is used to execute the labeling method for non-covalent interaction in a molecular system provided by the above embodiment, and has the same implementation manner and principle. For details, refer to the related description of the method embodiment, which will not be repeated here.

[0091] Figure 6 An electronic device of the embodiment is shown, which comprises a processor 901 and a memory 902, which can be connected through a bus or other means, Figure 6 for example, through a bus connection.

[0092] The processor 901 can be a central processing unit (CPU). The processor 901 can also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, or combinations thereof.

[0093] The memory 902, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs, non-transitory computer-executable programs and modules, such as program instructions / modules corresponding to the methods in the above method embodiments. The processor 901 performs various functional applications and data processing of the processor by running the non-transitory software programs, instructions and modules stored in the memory 902, that is, implements the methods in the above method embodiments.

[0094] The memory 902 can include a program storage area and a data storage area, where the program storage area can store an operating system, at least one application required by a function; and the data storage area can store data created by the processor 901 and the like. In addition, the memory 902 can include a high-speed random access memory, and can also include a non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state memory device. In some embodiments, the memory 902 can optionally include a memory disposed remotely with respect to the processor 901, and these remote memories can be connected to the processor 901 through a network. Examples of the above network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0095] One or more modules are stored in the memory 902, and when executed by the processor 901, the methods in the above method embodiments are performed.

[0096] The above electronic device specific details can be understood by referring to the corresponding related descriptions and effects in the above method embodiments, which will not be repeated here.

[0097] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program. The implemented program can be stored in a computer readable storage medium. When the program is executed, it can include the processes of the above-mentioned embodiment methods. The storage medium can be a magnetic disc, an optical disc, a read-only memory (ROM), a random access memory (RAM), a flash memory, a hard disk drive (HDD) or a solid-state drive (SSD), etc. The storage medium can also include a combination of the above-mentioned types of memories.

[0098] Although the embodiments of the present application are described in conjunction with the drawings, various modifications and changes can be made by those skilled in the art without departing from the spirit and scope of the present application, and such modifications and changes fall within the scope defined by the appended claims.

Claims

1. A method for annotating non-covalent interactions in a molecular system, characterized in that, The method comprises: Obtaining atomic coordinate information of a target molecular structure and electron density grid point information of a preset resolution, the electron density grid point information being used to store electron density of each grid point of the target molecular structure; Determining a saddle point position based on a grid point position where an electron density saddle point in the electron density grid point information is located, the saddle point position being used to represent a spatial position of a non-covalent interaction in the target molecular structure; Determining an attribution relationship of the saddle point position by using the atomic coordinate information, the attribution relationship being used to represent a matching relationship of the non-covalent interaction and a nearby atom pair; Generating an electron density topological feature of the saddle point position by using at least an electron density, a first-order gradient and a Hessian matrix of the saddle point position, the electron density topological feature being used to describe an attribute of the non-covalent interaction; Labeling the non-covalent interaction in the target molecular structure by using the saddle point position, the attribution relationship and the electron density topological feature.

2. The method of claim 1, wherein, Obtaining electron density grid point information of a preset resolution comprises: Converting the atomic coordinate information from a real space to a frequency domain space to generate a structure factor; Adjusting the structure factor to return to the real space at a preset resolution to generate the electron density grid point information, the preset resolution being a preset summation range of frequency domain space vector components.

3. The method of claim 1, wherein, The determination of the saddle point position based on the grid point position where the electron density saddle point in the electron density grid point information is located comprises: Dividing the electron density grid point information into a plurality of sub-information blocks of a first preset range size; Obtaining a ligand-receptor atom pair in the sub-information block, and taking a range covered by a midpoint position of a connecting line of the ligand-receptor atom pair as a candidate range, the midpoint position being taken as a center and a preset length being taken as a radius; Calculating a reduced electron density gradient value of each grid point in the candidate range, and screening out a grid point with a minimum reduced electron density gradient value in the candidate range as a candidate saddle point; Calculating a reduced electron density gradient value of a neighboring grid point of the candidate saddle point, and taking the candidate saddle point and the grid point with the minimum reduced electron density gradient value as the saddle point position.

4. The method of claim 3, wherein, After the calculation of the reduced electron density gradient value of the neighboring grid point of the candidate saddle point and the taking of the candidate saddle point and the grid point with the minimum reduced electron density gradient value as the saddle point position, the method further comprises: Abandoning the saddle point position whose eigenvalue of the electron density Hessian matrix does not satisfy λ1<λ2<0<λ3.

5. The method of claim 4, wherein, The determination of the attribution relationship of the saddle point position by using the atomic coordinate information, the attribution relationship being used to represent the matching relationship of the non-covalent interaction and the nearby atom pair comprises: When one saddle point position corresponds to a plurality of ligand-receptor atom pairs, taking an atom pair with the closest distance between a ligand and a receptor as a target atom pair to be labeled.

6. The method of claim 5, wherein, After the taking of the atom pair with the closest distance between the ligand and the receptor as the target atom pair to be labeled when one saddle point position corresponds to a plurality of ligand-receptor atom pairs, the method further comprises: When a distance between any two saddle point positions is less than a preset distance, abandoning a saddle point position with a smaller electron density value.

7. The method according to any one of claims 1 to 6, characterized in that, The preset resolution comprises a plurality of preset resolutions with different sizes, and the method further comprises: The step of obtaining atomic coordinate information of a target molecular structure and preset-resolution electron density grid information, and the step of labeling non-covalent interactions in the target molecular structure according to the saddle point position, the attribution relationship and the electron density topological feature are performed respectively under each resolution condition.

8. A non-covalent interaction labeling system in a molecular system, characterized by, The system comprises: an information acquisition module configured to obtain atomic coordinate information of a target molecular structure and preset-resolution electron density grid information, the electron density grid information being configured to store electron density of each grid point of the target molecular structure; a position acquisition module configured to determine a saddle point position based on a grid point position where an electron density saddle point is located in the electron density grid information, the saddle point position being configured to represent a spatial position of a non-covalent interaction in the target molecular structure; an attribution matching module configured to determine an attribution relationship of the saddle point position by using the atomic coordinate information, the attribution relationship being configured to represent a matching relationship between the non-covalent interaction and nearby atoms; a feature acquisition module configured to generate an electron density topological feature of the saddle point position by using at least electron density, a first-order gradient and a Hessian matrix of the saddle point position, the electron density topological feature being configured to describe a type and an attribute of the non-covalent interaction; a non-covalent interaction labeling module configured to label the non-covalent interaction in the target molecular structure according to the saddle point position, the attribution relationship and the electron density topological feature.

9. An electronic device, comprising: comprise: a memory and a processor, which are communicatively connected to each other, the memory stores computer instructions, and the processor executes the computer instructions to perform the method according to any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer readable storage medium stores computer instructions, and the computer instructions are configured to cause the computer to perform the method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Complex system reaction access calculating system and implementing method thereof

    CN104021265A

  • Energy-based multi-objective optimization fitting prediction method for atomic structures and electron density maps

    CN111968707A