A protein-protein docking method and system based on multi-region division

By dividing the surface atoms of the receptor protein into multiple subspaces for parallel searching of the ligand protein posture, and combining this with an adaptive differential algorithm for optimization, the problems of low efficiency and high error in existing protein docking software are solved, achieving efficient and accurate protein docking.

CN115938469BActive Publication Date: 2025-11-28SHENZHEN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202211377734.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-04
Publication Date
2025-11-28
Estimated Expiration
2042-11-04

AI Technical Summary

Technical Problem

Existing protein docking software is inefficient and has a high error rate when searching for protein complex structures, and cannot effectively improve the accuracy of protein docking.

Method used

A multi-region partitioning approach is adopted to divide the surface atomic space of the receptor protein into multiple subspaces, and the pose search of the ligand protein is performed in parallel. The conformation of the candidate ligand protein in each subspace is optimized by an adaptive differential algorithm, and finally the optimal ligand protein pose is determined by merging.

Benefits of technology

It significantly improves the efficiency and accuracy of protein docking, shortens the search time, and improves the accuracy of protein complex structure prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115938469B_ABST
    Figure CN115938469B_ABST
Patent Text Reader

Abstract

The application provides a protein-protein docking method and system based on multi-region division, and relates to the technical field of biology. The method comprises the following steps: keeping the posture of one protein with a relatively large number of atoms in two docked proteins unchanged as a receptor protein, dividing the surface atoms of the receptor protein into a plurality of subspaces, independently and in parallel searching the posture of a ligand protein in the plurality of subspaces, determining the candidate ligand protein conformation generated by each subspace, greatly shortening the search time length of the protein-protein docking posture, improving the docking efficiency of the receptor protein and the ligand protein, and combining the candidate ligand protein conformations generated by each subspace, and then determining the posture of the ligand protein that is best docked with the receptor protein from the combined candidate ligand protein conformations, thereby improving the accuracy of the receptor protein and the ligand protein.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of biotechnology, and particularly relates to a protein-protein docking method and system based on multi-region division. BACKGROUND

[0002] Proteins are the basis of life activities and the main bearers of life activities. Almost all important components of human cells and tissues are involved in proteins, and protein-protein interactions are involved in a large number of biological processes related to human health and diseases. Therefore, understanding the interactions and binding processes between proteins at the molecular level helps to further understand the relationship between protein structure and function and promote application research in the field of biochemistry.

[0003] With the continuous development of technology, it has become possible to use computers to simulate the docking between proteins and predict the structure of protein complexes. Developing an effective molecular docking method to predict the structure of protein complexes has important guiding significance and practical value for studying the interactions between proteins and determining the structure of unknown protein complexes. SUMMARY

[0004] The embodiments of the present application provide a protein-protein docking method and system based on multi-region division, which can effectively improve the efficiency and accuracy of protein docking.

[0005] In a first aspect, the present application provides a protein-protein docking method based on multi-region division, comprising: according to the number of atoms of a target protein to be docked, one of the target protein is recorded as a receptor protein, and the other is recorded as a ligand protein, the number of atoms of the receptor protein is greater than that of the ligand protein, and the pose of the receptor protein remains unchanged; performing spatial division on the surface atoms of the receptor protein to obtain M subspaces, wherein M is an integer greater than 1; performing pose search of the ligand protein in the M subspaces in parallel to determine the candidate ligand protein conformations generated by each of the subspaces; merging the candidate ligand protein conformations generated by the M subspaces, and determining the pose of the ligand protein that is best docked with the receptor protein from the merged candidate ligand protein conformation set.

[0006] The embodiment of the present application keeps the pose of one protein with relatively large number of atoms unchanged as a receptor protein, divides the surface atoms of the receptor protein into a plurality of subspaces, independently and in parallel searches the pose of the ligand protein in each of the plurality of subspaces, determines the candidate ligand protein conformation generated by each of the plurality of subspaces, greatly shortens the search time of the protein-protein docking pose, improves the docking efficiency of the receptor protein and the ligand protein, and combines the candidate ligand protein conformations generated by each of the plurality of subspaces, and then determines the pose of the ligand protein best docked with the receptor protein from the combined candidate ligand protein conformations, thereby improving the accuracy of the receptor protein and the ligand protein.

[0007] In a second aspect, the present application provides a protein-protein docking system based on multi-region division, comprising:

[0008] A protein marking unit is configured to mark one of target proteins to be docked as a receptor protein and the other as a ligand protein according to the number of atoms of the target proteins, wherein the number of atoms of the receptor protein is greater than that of the ligand protein, and the pose of the receptor protein is kept unchanged.

[0009] A subspace division unit is configured to divide the surface atoms of the receptor protein in space to obtain M subspaces, wherein M is an integer greater than 1.

[0010] A pose searching unit is configured to search the pose of the ligand protein in the M subspaces in parallel, and determine the candidate ligand protein conformation generated by each of the M subspaces.

[0011] A pose determining unit is configured to combine the candidate ligand protein conformations generated by the M subspaces, and determine the pose of the ligand protein best docked with the receptor protein from the combined candidate ligand protein conformation set.

[0012] In a third aspect, the present application provides a protein-protein docking device based on multi-region division, comprising a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the method of the first aspect.

[0013] In a fourth aspect, the present application provides a computer readable storage medium, wherein the computer readable storage medium stores a computer program, and the computer program is executable on a processor to implement the method of the first aspect.

[0014] In a fifth aspect, the embodiments of the present application provide a computer program product, which, when running on the multi-region-based protein-protein docking device, enables the multi-region-based protein-protein docking device to perform the steps of the multi-region-based protein-protein docking method according to the first aspect.

[0015] It can be understood that the beneficial effects of the second aspect to the fifth aspect described above can be referred to the related description in the first aspect, which will not be repeated here. BRIEF DESCRIPTION OF DRAWINGS

[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiments or prior art description will be briefly introduced as follows. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0017] Figure 1 is a flowchart of a multi-region-based protein-protein docking method provided by the embodiments of the present application;

[0018] Figure 2 is a sub-space related point set graph provided by the embodiments of the present application;

[0019] Figure 3 is a center point scatter plot of each sub-space after division of surface atoms of a receptor protein provided by the embodiments of the present application;

[0020] Figure 4 is a flowchart of a sub-space division method provided by the embodiments of the present application;

[0021] Figure 5 is a flowchart of a method for generating a candidate ligand protein conformation provided by the embodiments of the present application;

[0022] Figure 6 is a flowchart of a method for determining the docking energy of a protein provided by the embodiments of the present application;

[0023] Figure 7 is a flowchart of a method for optimizing the pose of a protein in a docking process provided by the embodiments of the present application;

[0024] Figure 8 is a structural diagram of a multi-region-based protein-protein docking system provided by the embodiments of the present application;

[0025] Figure 9is a structural schematic diagram of a protein-protein docking device based on multi-region division provided by an embodiment of the present application. DETAILED DESCRIPTION

[0026] In the following description, specific details are set forth, such as particular system configurations, techniques, etc., in order to provide a thorough understanding of the embodiments of the present application. However, persons skilled in the art will understand that the present application can be practiced without these specific details. In other instances, well-known devices, systems, circuits, and methods have not been described in detail so as not to unnecessarily obscure the present application.

[0027] It should be understood that the term "and / or" used in the description of the present application and the appended claims means one or more of the associated listed items as well as all possible combinations of the items and includes these combinations. In addition, in the description of the present application and the appended claims, the terms "first", "second", "third", etc. are used only to distinguish descriptions and cannot be understood as indicating or implying relative importance.

[0028] It should also be understood that the reference "one embodiment" or "some embodiments" and the like means that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the present application. Thus, the appearances of the phrases "in one embodiment", "in some embodiments", "in other embodiments", "in additional embodiments", and the like, in various places in the specification are not necessarily all referring to the same embodiment, unless otherwise specifically noted. The terms "comprise", "comprising", "have", "having", "include", "including", and "contain", "containing", and the like, mean "including but not limited to", unless otherwise specifically noted.

[0029] Protein-protein docking is to find the best docking pose in three-dimensional space through shape complementation, energy matching, etc. between two proteins with known spatial structures. The protein-protein docking method based on multi-region division provided in the embodiments of the present application only considers rigid docking between proteins, i.e. the internal structure of two proteins does not change during docking. By fixing one protein, the other protein is docked around the surface of the fixed protein. During docking, only the pose of one protein needs to be considered. The movable protein is adjusted around the surface of the fixed protein to obtain different docking energies. In the simulation of protein-protein docking, the final goal is to find the best docking pose between the two.

[0030] The currently used protein docking software such as ZDOCK, FTDock and ClusPro adopts a docking algorithm based on fast Fourier transform and Monte Carlo simulation to search the protein conformation. Although the specific docking algorithm is different, in the full-space docking process on the protein surface, the protein posture is searched by adjusting the protein surface step by step in a serial manner, which is low in search efficiency and high in error rate.

[0031] Please refer to Figure 1 , Figure 1 is a flowchart of a protein-protein docking method based on multi-region division provided by the embodiment of the present application, which is described in detail as follows:

[0032] In step S101, according to the number of atoms of the target protein for docking, one of the above target proteins is recorded as a receptor protein, and the other is recorded as a ligand protein.

[0033] In the embodiment of the present application, before protein-protein docking, the number of atoms of the target protein, i.e. the two proteins for docking, is obtained to determine the fixed protein and the moving protein.

[0034] In the embodiment of the present application, the fixed protein is recorded as the receptor protein, and the moving protein is recorded as the ligand protein. The number of atoms of the receptor protein is greater than that of the ligand protein. The ligand protein moves in the surface space of the receptor protein to search for the best docking posture.

[0035] In step S102, the surface atoms of the above receptor protein are spatially divided to obtain M subspaces.

[0036] In the embodiment of the present application, the receptor protein is fixed, and the surface atoms of the receptor protein are spatially divided to divide the surface atoms into different subspaces to obtain a plurality of subspaces.

[0037] In order to improve the search efficiency of the protein-protein docking posture, the surface atoms of the receptor protein are spatially divided to obtain M subspaces to search the docking posture of the ligand protein in parallel in the M subspaces, shorten the search time of the protein-protein docking posture, and thus improve the efficiency of the protein-protein docking.

[0038] In the embodiment of the present application, the division number M of the subspaces is an integer greater than 1, preferably an integer greater than 100, and more preferably an integer greater than 1000. The greater the division number M of the subspaces, the higher the search efficiency of the protein-protein docking posture.

[0039] Please refer to Figure 2 ,Figure 3 and Figure 4 , Figure 2 This is a subspace-related point set diagram provided in an embodiment of this application. Figure 3 This is a scatter plot of the center points of each subspace after dividing the surface atoms of a receptor protein according to an embodiment of this application. Figure 4 This is a flowchart illustrating a subspace partitioning method provided in an embodiment of this application, combined with... Figure 2 and 3 ,right Figure 4 The steps shown are described in detail below:

[0040] Step S401: Randomly select M atoms from the first set of atoms as the initial cluster centers.

[0041] In this embodiment, all atoms on the surface of the receptor protein are extracted to obtain a first atom set, which is the set of all atoms extracted from the surface of the receptor protein. Then, based on the k-means clustering algorithm, M atoms are randomly selected from the first atom set as initial cluster centers {a1, a2, ..., a...}. M}.like Figure 2 As shown, the initial cluster centers are {a1, a2, ..., a...} M}for Figure 2 A set of points randomly distributed throughout the space (with some points overlapping the spherical point set). Figure 2 The numerical values ​​in the text represent coordinate scale values.

[0042] It should be noted that, unless otherwise specified, the atoms mentioned in the embodiments of this application are by default atoms on the surface of the receptor protein.

[0043] Step S402: Based on the distance between each atom and each of the initial cluster centers, each atom in the first atom set is assigned to the category corresponding to the initial cluster center with the smallest distance, and the cluster center of each category after re-division is calculated to obtain M second cluster centers.

[0044] In the embodiments of this application, for each atom i Calculate the atom respectively i The distance to each initial cluster center, the atom i Assign the atom to the category corresponding to the initial cluster center with the smallest distance, where i represents the atom's label. After classifying all atoms, for each category N... j Using the cluster center calculation formula, the cluster center 'a' for each category is recalculated. j The formula for calculating cluster centers is:

[0045]

[0046] Among them, a j The cluster centers are represented by j∈[1,M]; m represents each cluster N. j The number of atoms contained, where x represents the number of atoms. i Location information; N j Indicates the category, specifically using 'a' j The set of atoms that serves as the cluster center.

[0047] This application embodiment recalculates the cluster center for each category using a cluster center calculation formula, which is essentially averaging the atomic positions contained in each category. To ensure the accuracy of each second cluster center obtained in the end, thereby improving the accuracy of protein-protein docking, for each atom... i The above cluster center calculation formula needs to be used repeatedly for iterative operations, such as terminating after 100 iterations or after 200 iterations. In this application embodiment, no specific limit is made on the number of iterations, and the number of iterations with the smallest iteration error is the optimal number of iterations.

[0048] Step S403: Based on the first preset distance, generate corresponding spherical point sets with M of the above-mentioned second cluster centers as the center points, and obtain M spherical point sets.

[0049] In this embodiment of the application, the center of a certain subspace, namely the second cluster center, is taken as the center, and a first preset distance is taken as the radius from the center of the subspace. Generate a point set of number P within a certain range, resulting in a set such as... Figure 2 The center point is a five-pointed star with a radius of 1. The spherical point set T represented by points within the range k Where k is the index of each point in the spherical point set, k∈[0, P-1]. For each subspace, for example, centered at the second distance center, the radius is... Within the specified area, generate the corresponding spherical point set T using the method described above. k Each spherical point set T k The number of midpoints is P.

[0050] In this embodiment, the first preset distance is any distance ranging from one-fifth to one-third of the distance between the two atoms with the largest distance between them in the ligand protein. Preferably, the first preset distance is one-quarter of the distance between the two atoms with the largest distance between them in the ligand protein. Here, the atom refers to any one of all atoms in the ligand protein, namely, any one of the surface atoms and internal atoms.

[0051] Step S404, deleting redundant points in each spherical point set, and merging the spherical point sets after deleting the redundant points to obtain a second atom set.

[0052] In the embodiment of the present application, in order to reduce the influence of the redundant points on the position information of the subspace, after obtaining the M spherical point sets, each cluster center a j is traversed to find a subspace center meeting a preset condition, the preset condition being that the points within a second preset distance, such as a radius j , of the cluster center a are neighbors, and the subspace center meeting the preset condition is a neighbor center a k . After finding the neighbor center a j of each cluster center a k , the neighbor center a k is traversed, the distance between the cluster center a j and each point in the spherical point set to which the neighbor center a k belongs is calculated, and when the distance is less than the second preset distance, the corresponding point in the spherical point set to which the neighbor center a k belongs is deleted.

[0053] After the traversal of each cluster center a j of each subspace ends, the redundant points between the spherical point sets corresponding to the cluster centers a j that are relatively close are deleted, and the updated spherical point set, that is, the spherical point set after deleting the redundant points, is obtained.

[0054] The updated spherical point set is merged to obtain a second atom set, that is, the second atom set is the union of the spherical point sets after deleting the redundant points.

[0055] In the embodiment of the present application, the redundant points in the spherical point set are deleted, and the influence of the redundant points on the position information of the subspace is reduced, so that the accuracy of the coordinates of the divided subspace can be improved.

[0056] Step S405, randomly selecting M atoms from the second atom set as third cluster centers.

[0057] In the embodiment of the present application, M atoms are randomly selected from the second atom set as third cluster centers.

[0058] Step S406, dividing each atom in the second atom set into a category corresponding to the third cluster center with which the distance is smallest according to the distance between each atom and each third cluster center, and calculating the cluster center of each category after the division to obtain M fourth cluster centers, the fourth cluster centers being subspace center points.

[0059] In the embodiment of the present application, for each atom atomi Calculate the atom respectively i The distance to each initial cluster center, the atom i Assign the atom to the category corresponding to the initial cluster center with the smallest distance, where i represents the atom's label. After classifying all atoms, for each category N... j Using the cluster center calculation formula, the cluster center 'a' for each category is recalculated. j The formula for calculating cluster centers is:

[0060]

[0061] Among them, a j The cluster centers are represented by j∈[1,M]; m represents each cluster N. j The number of atoms contained, where x represents the number of atoms. i Location information; N j Indicates the category, specifically using 'a' j The set of atoms that serves as the cluster center.

[0062] This application embodiment recalculates the cluster center for each category using a cluster center calculation formula, which is essentially averaging the atomic positions contained in each category. To ensure the accuracy of each second cluster center obtained in the end, thereby improving the accuracy of protein-protein docking, for each atom... i The above cluster center calculation formula needs to be used repeatedly for iterative operations, such as terminating after 100 iterations or after 200 iterations. In this application embodiment, no specific limit is made on the number of iterations, and the number of iterations with the smallest iteration error is the optimal number of iterations.

[0063] In the embodiments of this application, the fourth cluster centers are uniformly distributed on the surface of the receptor protein. For example... Figure 3 As shown, the points uniformly dispersed above the surface of the clump-like object (receptor protein) are the fourth cluster centers, which are also the center points of all the subspaces ultimately generated. To clearly show the position of the fourth cluster centers on the surface of the receptor protein in the figure, the points representing all atoms of the receptor protein will be used when drawing the image. Figure 3 The clump-like objects in the middle at the same time Figure 3 The data is displayed within the structure, meaning the cluster represents the positions of all atoms in the receptor protein, including both surface and internal atoms. Figure 3 The numerical values ​​in the text represent coordinate scale values.

[0064] Step S103: Perform a parallel orientation search for ligand proteins in the M subspaces mentioned above to determine the conformation of candidate ligand proteins generated in each subspace.

[0065] In the embodiments of the present application, each subspace is an independent individual without interference, and the pose search of the ligand protein can be performed in parallel in each subspace, thereby effectively improving the efficiency of the pose search of the ligand protein. The search strategy adopted by each subspace is consistent, for example, the adaptive differential evolution algorithm is adopted as the search strategy of the subspace. The following is a specific description of the pose search of the ligand protein in a single subspace.

[0066] Since the protein-protein docking in the embodiments of the present application only considers rigid docking, the receptor protein remains fixed in its pose during docking, and the ligand protein moves on the surface of the receptor protein, aiming to find the pose of the ligand protein with the best docking energy with the receptor protein, and the ligand protein is considered as rigid, that is, the relative positions of all atoms inside the ligand protein remain unchanged during the docking process, only the position and direction of the whole body change, which is specifically manifested as movement and selection in the pose of the ligand protein. Therefore, the pose of the ligand protein in the subspace can be encoded as a six-dimensional vector {x, y, z, θ, α, β} composed of two three-dimensional vectors during the optimization process using the adaptive differential algorithm, wherein the first three-dimensional vector {x, y, z} represents the movement amount of all atomic coordinates of the ligand protein, and the second three-dimensional vector {θ, α, β} represents the overall direction movement of the ligand protein, wherein θ is the rotation angle of the ligand protein around the rotation axis, θ ∈ (0, 2π); α is the angle between the rotation axis and the z axis, α ∈ (0, π); and β is the angle between the projection of the rotation axis to the xoy plane and the x axis, β ∈ (0, 2π). The six-dimensional vector represents the movement and selection of the ligand protein in the subspace, and the purpose of the pose search of the ligand protein in the subspace is to find the pose adjustment (movement and rotation) of the ligand protein in the subspace, which can best dock with the receptor protein.

[0067] In the embodiments of the present application, for any one of the M subspaces, the steps shown in Figure 5 are performed to determine the candidate ligand protein conformations generated by each subspace. Specifically, refer to Figure 5 , Figure 5 is a flowchart of a method for generating a candidate ligand protein conformation provided by the embodiments of the present application, which is described in detail as follows:

[0068] In step S501, the population size is initialized, and the population size N individuals are randomly generated within a first preset distance range from the center of each subspace.

[0069] In the embodiments of the present application, after the center of each subspace is determined, the population initialization operation is performed, and the population size N individuals are randomly generated within a radius of the center of the subspace, for example, within a radius of , the individuals being ligand proteins with different poses.

[0070] During population initialization, the following notation is defined: the population size is N, and the initial set of individuals in the population is P = {p1, p2, ..., p...}. N}, p i ={x i y i , z i θ i α i ,β i}, p i Let represent the six-dimensional vector corresponding to individual i in the population, and let T be the number of population iterations in the subspace. Where (x... i y i , z i () is centered at the center of the subspace, with a radius of The randomly generated offset value represents the individual p. i The total amount of movement; (θ) i α i ,β i Let be the angle value randomly generated within the angle domain, representing the individual p. i The entire rotation generates random values ​​for each dimension within its defined range. After generating N individuals, the population initialization operation is complete.

[0071] During population initialization, the scaling factor F and crossover probability CR for each individual are generated simultaneously. Specifically, when each individual is generated, the scaling factor F and crossover probability CR for that individual are randomly generated from a normal distribution using the random generation formula for the scaling factor and crossover probability. The random generation formula for the scaling factor and crossover probability is as follows:

[0072] F i =randn(H F ,0.1)

[0073] CR i =randn(H CR ,0.1)

[0074] Among them, F i Individual p i The scaling factor, F∈(0,1); CR i Individual p i The crossover probability; H F and H CR The initial value is set to 0.5, H F H represents the initial mean of the normal distribution corresponding to the scaling factor F. CR This represents the initial mean of the normal distribution corresponding to the crossover probability CR. Both values ​​will be updated during subsequent population iterations.

[0075] Step S502, the docking energy of N individuals in different poses with the receptor protein in the subspace is determined by iterative optimization of the poses of the N individuals through an adaptive difference algorithm, and different docking energies correspond to different individual poses.

[0076] In the embodiment of the present application, after the population size is initialized and the scaling factor F and the crossover probability CR corresponding to each individual are determined, the pose of each initialized individual is judged according to the position information of the initialized individual and the position information of the receptor protein by using a scoring function, and the docking energy between the initialized individual and the receptor protein is calculated.

[0077] It should be noted that the result of the scoring function is used as the docking energy between the current calculated individual (the ligand protein in the current pose) and the receptor protein, and the docking between the two is judged according to the size of the score. The larger the value given by the scoring function, the higher the docking energy, and the better the docking. The scoring function in the embodiment of the present application is a conventional scoring function, as shown below:

[0078]

[0079] Among them, and are the van der Waals attraction term and the repulsion energy term, respectively. elec_sra and E elec_srr are the short-range electrostatic attraction term and the repulsion energy term, respectively. elec_lra and E elec_lrr are the long-range electrostatic attraction term and the repulsion energy term, respectively. ds is the desolvation energy; and the weights are w elec_sra = 0.31, w elec_srr = 0.34, w elec_lra = 0.44, w elec_lrr = 0.5, and w ds = 1.02.

[0080] In addition to the above scoring function, there are other scoring functions to calculate the energy value of the docking of two proteins. The process of calculating the docking energy based on the scoring function is not specifically expanded in the embodiment of the present application.

[0081] In the embodiment of the present application, after the poses of the initialized individuals are scored by using the scoring function, the initialized individuals are sorted according to the energy.

[0082] For any individual in the population, the method for determining the protein docking energy shown in Figure 6 is executed to obtain the docking energy corresponding to each individual. Please refer to Figure 6 ,Figure 6 is a flowchart of a method for determining the docking energy of a protein provided in an embodiment of the present application, which is described in detail as follows:

[0083] In step S601, a six-dimensional vector representing the pose of the individual is obtained.

[0084] In the embodiment of the present application, the individual in the population is a ligand protein, and the pose of each ligand protein is represented by a six-dimensional vector composed of a first three-dimensional vector and a second three-dimensional vector. Before determining the docking energy of the ligand protein and the receptor protein in the current subspace, the six-dimensional vector representing the pose of the ligand protein is obtained, so as to facilitate the next operation.

[0085] In step S602, the six-dimensional vector corresponding to the individual is optimized by using an adaptive differential algorithm to obtain an optimized six-dimensional vector.

[0086] In the embodiment of the present application, after the population initialization of each subspace is completed, the initial six-dimensional vector of each individual represents the ligand protein that is initially moved into the subspace. The six-dimensional vector corresponding to the individual is optimized, that is, the individual is moved and rotated in the subspace to change the pose of the individual, so as to rotate the pose of the ligand protein to the optimal pose for docking with the receptor protein in the subspace. The six-dimensional vector corresponding to each individual in the population is iteratively optimized by using an adaptive differential algorithm.

[0087] In step S603, the first position information of the individual is determined according to the optimized six-dimensional vector. The first position information is the position information of all atoms in the individual after the pose change.

[0088] In the embodiment of the present application, in the docking process, the new coordinates of each atom of the ligand protein after the pose update are calculated by using the coordinate information of the ligand protein rotation and movement, so as to further evaluate the docking energy between the ligand protein and the receptor protein. For the optimized six-dimensional vector, that is, the pose of the individual after the pose change, the coordinates of all atoms of the initial ligand protein in the subspace are added to the movement amount {x, y, z} of the vector, and then the rotation angle {θ, α, β} is converted into a quaternion. The position information of each atom of the updated individual, that is, the ligand protein, is calculated by using the calculation method of the quaternion. Finally, the docking energy between the ligand protein and the receptor protein is evaluated by using the scoring function through all the atomic coordinates of the ligand protein and the receptor protein.

[0089] It should be noted that for the optimization of the six-dimensional vector, the six-dimensional vector needs to be optimized multiple times to find the best pose of the two protein docking. For each optimized six-dimensional vector, the new positions of all atoms in the ligand protein after the change of the pose are calculated through the movement and rotation information of the atoms contained in the six-dimensional vector in the above manner.

[0090] In the embodiment of the application, the three angles in the second three-dimensional vector are converted into a quaternion, and the atomic coordinates of the updated ligand protein are updated and calculated using the quaternion, wherein the vector form of the quaternion is as follows:

[0091] q=a+bi+cj+dk

[0092] wherein a, b, c, and d are real numbers. i, j, and k are imaginary parts, and satisfy i 2 =j 2 =k 2 =ijk. The quaternion can also be written in the form of a vector, q=(w, x, y, z).

[0093] The quaternion can be regarded as a representation method of rotation. Assuming that the coordinates of an atom are h=(h x , h y , h z ), and the rotation axis represented by a unit vector is g=(g x , g y , g z ), the coordinates h of the atom are rotated by θ around the rotation axis, and this rotation method can be represented by a quaternion , and the rotated atom coordinates h′=qhq -1 , wherein

[0094] The formula for converting the three angles into a quaternion is as follows:

[0095]

[0096]

[0097]

[0098]

[0099] wherein w, x, y, and z are four quantities of the quaternion, and have no specific physical meaning.

[0100] In the embodiment of the present application, since the initial coordinates of all atoms of the ligand protein are known, it is assumed that after the optimization of the six-dimensional vector is completed, the coordinates of the atoms after the change of the pose are calculated by the optimized six-dimensional vector. At this time, the initial coordinates of all atoms of the ligand protein are added to the movement amount in the six-dimensional vector, and then the angles in the six-dimensional vector are converted into quaternions, and the coordinates of the atoms after the rotation are updated by the quaternions, so as to update the atomic position information of the ligand protein after the change of the pose.

[0101] In the example of the present application, the updated atomic coordinates are calculated according to the information of the six-dimensional vector in the following manner: assuming that the original coordinates of one atom are h = (h x , h y , h z ), and the corresponding optimized six-dimensional vector is {x, y, z, θ, α, β}, first, the original coordinates of the atom are added to the first three-dimensional vector, h' = (h x +x, h y +y, h z +z). Then, the three angles in the second three-dimensional vector {θ, α, β} are converted into the corresponding quaternion q = (w, x, y, z) by using the above formula. Then, the rotated atomic coordinates h'' = qh'q -1 are calculated by using the calculation method of the quaternion. Therefore, h'' is the updated atomic coordinates.

[0102] Similarly, all atoms in the ligand protein are updated according to the above manner to obtain the updated overall ligand protein, and the update of the atomic position information of the ligand protein after the change of the pose is completed.

[0103] In the embodiment of the present application, the initial six-dimensional vector {x, y, z, θ, α, β} is optimized to obtain the optimized six-dimensional vector {x', y', z', θ', α', β'}, that is, the six-dimensional vector corresponding to the updated pose of the ligand protein. Then, the coordinates of all atoms of the ligand protein after the change of the pose are calculated by using the optimized six-dimensional vector, and then the docking energy between the updated ligand protein and the atomic coordinates of the fixed receptor protein is calculated by using the coordinates of all atoms in the updated ligand protein and the atomic coordinates of the fixed receptor protein.

[0104] After several iterations of optimization, the six-dimensional vector {x'', y'', z'', θ'', α'', β''} of the ligand protein in the best pose after docking is obtained. The pose of each ligand protein corresponds to a docking energy, which indicates whether the docking with the receptor protein is good or not.

[0105] In step S604, the docking energy between the individual and the receptor protein in the subspace is determined according to the first position information and the second position information. The second position information is the position information of the receptor protein.

[0106] In the embodiment of the present application, the docking energy of the ligand protein and the receptor protein is calculated by the scoring function according to the first position information and the second position information.

[0107] Please refer to Figure 7 , Figure 7 is a flowchart of a method for optimizing the pose of a protein in a docking process provided in the embodiment of the present application, which is described in detail as follows:

[0108] In step S701, at least one dimension in the six-dimensional vector of the individual is subjected to a crossover mutation operation based on the crossover probability corresponding to the individual by using the adaptive differential algorithm, so as to obtain an updated individual.

[0109] In the embodiment of the present application, the population size is initialized, the scaling factor F and the crossover probability CR corresponding to each individual are determined, and the pose of each initialized individual is judged by using the scoring function according to the position information of the initialized individual and the position information of the receptor protein, the docking energy between the initialized individual and the receptor protein is calculated, and the initialized individuals are sorted according to the docking energy, and then each individual in the population, i.e., the initialized individual, is updated by using the differential formula, wherein the differential formula is specifically as follows:

[0110]

[0111] wherein v i represents the six-dimensional vector corresponding to the individual numbered i after the differential operation, represents the value of the jth dimension in the six-dimensional vector v i corresponding to the next generation individual calculated by using the differential formula for the current individual i, wherein i∈[0, N-1], j∈[0, 5], t represents the number of iterations of the current population, t∈[0, T-1], N represents the population size, i.e., the number of individuals contained in the population; p i represents the six-dimensional vector corresponding to the individual i, respectively represent the value of the jth dimension in the six-dimensional vector corresponding to the current individual i, the individual pbest randomly selected from the top 5% to 15%, such as the top 8%, the top 10%, the top 12%, the individual r1 randomly selected from the current population, and the individual r2 randomly selected from the current population in the tth iteration, wherein r1 and r2 are integers randomly generated in [0, N-1], and it is noted that pbest, r1 and r2 are different from each other; F i represents the scaling factor corresponding to the individual i.

[0112] For each individual in the population, three individuals, pbest, r1, and r2, are selected from the current population. The crossover and mutation values ​​of each individual are then calculated using the aforementioned difference formula for each dimension. Following the crossover probability control formula, the decision is made whether to use the dimension value calculated using the above difference formula or the original dimension value of the individual. The crossover probability control formula is as follows:

[0113]

[0114] Among them, u i This represents the six-dimensional vector corresponding to individual i after the crossover and mutation operation. CR represents the value of the j-th dimension in the six-dimensional vector corresponding to the next-generation individual i, determined by the crossover probability control formula; i j represents the crossover probability corresponding to the current individual i; rand j represents a random integer generated in the individual dimension (individual dimension is 6). rand ∈[0, 5], ensuring that at least one dimension of the six-dimensional vector corresponding to individual i will be updated using the above difference formula.

[0115] For each dimension of individual i, a floating-point number less than 1 is randomly generated from between 0 and 1. If the crossover probability CR corresponding to the current individual i is... i If the value is greater than the randomly generated floating-point number, the dimension of individual i is updated using the above difference formula; otherwise, the dimension of individual i remains unchanged.

[0116] In this embodiment of the application, for the six-dimensional vector p of individual i in the population i For each dimension, the value corresponding to the current dimension is first calculated using the difference formula. Then, the cross-probability control formula is used to determine the value, and the cross-probability CR of individual i is used. i To determine whether to use the value calculated using the above difference formula, once each dimension of the six-dimensional vector has been confirmed using the above difference formula and the above cross-probability control formula, the dimension update of individual i is complete, and the updated six-dimensional vector u corresponding to the individual is generated. i For each individual in the population, the dimensions are updated using the method described above.

[0117] Step S702: Based on the docking energy levels of individuals before and after the update, perform an individual selection operation, selecting the individual with the highest docking energy as the current population, and updating the population.

[0118] In this embodiment, after updating all individuals in the population through difference and crossover / mutation operations, a six-dimensional vector u corresponding to a temporary individual with the same size as the original population is obtained. i, and the docking energy of the temporary individual and the receptor protein is calculated, and the docking energy of the temporary individual and the receptor protein before the difference operation and the crossover mutation operation is compared, and the six-dimensional vector corresponding to the individual in the population is selected from the two. Specifically, the final six-dimensional vector of the individual is determined through an individual selection formula, and the individual selection formula is:

[0119]

[0120] wherein f() represents a scoring function, and the docking energy between the two proteins is calculated through the scoring function.

[0121] In some embodiments of the present application, if the docking energy of the updated individual corresponding to the six-dimensional vector is higher than that of the original six-dimensional vector, that is, the docking of the updated individual and the receptor protein is better, the absolute value of the difference between the docking energies is calculated, denoted as fitness i . Wherein:

[0122]

[0123] At the same time, the F and CR values corresponding to the updated better individual pose are recorded, denoted as S F,i and S CR,i . For the six-dimensional vector p i corresponding to the individual with lower docking energy that is eliminated between the two, it is stored in a temporary cache pool A for the difference operation and the crossover mutation operation of the next generation, that is, in the next iteration optimization process, r2 in the above is randomly selected from the current population and the cache pool A, that is, r2 ∈ (0, N+|A|-1).

[0124] When all the individual selection and recording are completed, all the individuals in the updated population are sorted in descending order according to the docking energy.

[0125] In the embodiments of the present application, while the selection operation is performed, the weight value corresponding to the selected individual is determined according to the absolute value of the difference between the docking energies corresponding to the individuals before and after the update, and the first mean value and the second mean value are updated according to the weight value, wherein the first mean value is the mean value of the normal distribution corresponding to the scaling factor, and the second mean value is the mean value of the normal distribution corresponding to the crossover probability.

[0126] In the embodiments of the present application, the weight value corresponding to each individual is calculated through a weight calculation formula, wherein the weight calculation formula is:

[0127]

[0128] wherein w i represents the weight value of the selected individual i, and N' represents the number of individuals with better docking energy after the population is updated, represents the sum of absolute values of differences between docking energies of all updated individuals in the population and their corresponding individuals before updating. After the weight value corresponding to the selected individual is calculated, the S F,i and S CR,i The first mean and the second mean are updated by a mean parameter updating formula, where the mean parameter updating formula is:

[0129]

[0130]

[0131] The updated first mean and the second mean in the embodiment of the application will be used in the next iteration optimization.

[0132] In step S703, population size attenuation operation is performed on the updated population.

[0133] In the embodiment of the application, after the selection operation and the parameter updating are completed, the size of the population is reduced by a population size reduction formula, where the population size reduction formula is:

[0134]

[0135] where N min represents the minimum size of attenuation. N max represents the maximum size of the population, which is consistent with the initial population size, N max =N; g represents the current cumulative number of individual evaluations, and G represents the maximum number of individual evaluations, and the calculation result is rounded down.

[0136] After the size of the population after attenuation is obtained, N t -N t+1 individuals in the population are deleted (from the lowest ranked individual), N t represents the number of individuals in the current population, N t+1 represents the number of individuals in the population in the next iteration.

[0137] In the embodiment of the application, after the individuals in the population are updated each time, the population size is attenuated, and the lower ranked individuals are deleted according to the size of the population after attenuation, which can effectively improve the efficiency of population updating and ensure that the poses of the individuals in the population are the best poses for the current docking.

[0138] In order to achieve the purpose of population optimization, the above steps S701-S703 need to be repeated until the target iteration number is reached, and the final population is obtained, and the above target iteration number is the population iteration number of the subspace.

[0139] In the embodiment of the present application, the individual in the population is subjected to the difference operation and the cross variation operation, so as to realize the optimization of the individual posture in the next population iteration process, and further improve the efficiency and accuracy of the ligand protein and the receptor protein docking posture.

[0140] In step S503, the individuals in different postures are sorted from high to low according to the final determined docking energy.

[0141] In step S504, the postures corresponding to the Q individuals in the front are selected as the candidate ligand protein conformations generated by the above-mentioned subspace, wherein Q≤N.

[0142] In the embodiment of the present application, after the iteration optimization of the individual posture is completed, the individuals in different postures are sorted from high to low according to the final determined docking energy of the individual, and the postures corresponding to the Q individuals in the front are selected as the candidate ligand protein conformations generated by the subspace corresponding to the population.

[0143] It can be understood that the number of the candidate ligand protein conformations finally generated in each subspace is the same as the population size after the attenuation.

[0144] In step S104, the M candidate ligand protein conformations generated by the above-mentioned subspace are merged, and the posture of the ligand protein best docked with the above-mentioned receptor protein is determined from the candidate ligand protein conformation set obtained by the merging.

[0145] In the embodiment of the present application, after the candidate ligand protein conformation generated by each subspace is determined, all the candidate ligand protein conformations generated by the subspace are merged, and the candidate ligand protein conformations are sorted from high to low according to the docking energy of the individual corresponding to the candidate ligand protein conformation, which represents the posture of the ligand protein docked with the receptor protein from good to bad.

[0146] In order to ensure the accuracy of the protein docking, the posture of the individual corresponding to the candidate ligand protein conformation in the front is selected as the best posture docked with the receptor protein.

[0147] Please refer to Table 1, which is the docking result of the docking simulation on 57 protein complexes by using the protein docking method provided in the embodiment of the present application.

[0148] Table 1

[0149]

[0150]

[0151]

[0152]

[0153] Top100, 50, 10, 5, 1 respectively represent the minimum error value of the actual docking receptor protein structure in the top 100, top 50, top 10, top 5 and first candidate ligand protein conformation pose in the final generated several candidate ligand protein conformations. The error value is calculated by the average coordinates of all alpha carbon atoms of the predicted candidate ligand protein conformation pose and the actual pose contact surface, that is, the docking contact range between the ligand protein and the receptor protein, specifically referring to all alpha carbon atoms of the ligand protein all amino acid residues located in the third preset distance such as of the receptor protein. The average distance between the alpha carbon atoms in the ligand protein satisfying the requirement and the actual complex such as the fixed receptor protein in the embodiment of the present application is calculated, considered as hit, that is, the predicted candidate ligand protein conformation pose and the pose of the receptor protein are successfully docked. Hits represent the number of successful docking in 57 tested complexes.

[0154] It should be noted that in the experimental verification part, the docking experiment is carried out by using the public data set, and there are several existing receptor protein and ligand protein complexes in the data set. When using the multi-region division based protein-protein docking method provided in the embodiment of the present application, the receptor protein and the ligand protein are separated, and then the docking experiment is carried out on the two. After the docking is completed, the obtained candidate ligand protein and the corresponding ligand protein in the data set are compared to calculate the error value between the two. Here, only the ligand protein needs to be compared because the receptor protein is fixed in the docking process, and only the ligand protein is moved and rotated.

[0155] As can be seen from the above table, for only the top candidate ligand protein conformation pose (top1), the number of successful docking is 28 in the 57 tested complexes, and the docking success rate is close to 50%. And in the top 100 candidate ligand protein conformations, the pose with the minimum docking error is taken, the number of successful docking reaches 47, and the success rate reaches 82.5%. It is proved that the method in the present application is effective.

[0156] In the embodiments of the present application, one of the two proteins to be docked with a relatively large number of atoms is kept as a receptor protein without changing its pose, and the surface atoms of the receptor protein are divided into a plurality of subspaces, and the pose search of the ligand protein is independently and in parallel performed in the plurality of subspaces, candidate ligand protein conformations generated in each subspace are determined, the search time of the protein-protein docking pose is greatly shortened, the docking efficiency of the receptor protein and the ligand protein is improved, and the accuracy of the receptor protein and the ligand protein is improved by merging the candidate ligand protein conformations generated in each subspace and determining the pose of the ligand protein best docked with the receptor protein from the merged candidate ligand protein conformations.

[0157] It should be understood that the size of the serial number of each step in the above embodiments does not mean the order of execution, and the execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.

[0158] Based on the protein-protein docking method based on multi-region division provided in the above embodiments, the embodiments of the present application further provide a system embodiment for implementing the above method embodiments.

[0159] Please refer to Figure 8 , Figure 8 is a schematic diagram of a protein-protein docking system based on multi-region division provided by the embodiments of the present application. Each unit included is used to execute Figure 1 each step in the corresponding embodiments. For details, please refer to Figure 1 the related description in the corresponding embodiments. For the sake of illustration, only the parts related to the present embodiments are shown. Please refer to Figure 8 , the protein-protein docking system based on multi-region division 8 comprises:

[0160] a protein marking unit 81, configured to mark one of the target proteins to be docked as a receptor protein and the other as a ligand protein according to the number of atoms of the target proteins, the number of atoms of the receptor protein is greater than that of the ligand protein, and the pose of the receptor protein is kept unchanged;

[0161] a subspace division unit 82, configured to divide the surface atoms of the receptor protein in space to obtain M subspaces, wherein M is an integer greater than 1;

[0162] a pose searching unit 83, configured to perform the pose search of the ligand protein in parallel in the M subspaces, and determine the candidate ligand protein conformations generated in each subspace;

[0163] The posture determination unit 84 is configured to combine the M candidate ligand protein conformations generated by the subspaces, and determine the posture of the ligand protein that is best docked with the receptor protein from the combined candidate ligand protein conformation set.

[0164] In some embodiments of the present application, the sub-space division unit 82 includes:

[0165] The first cluster center selection sub-unit is configured to randomly select M atoms from a first atom set as initial cluster centers, the first atom set being a set of all atoms extracted from the surface of the receptor protein.

[0166] The first cluster center calculation sub-unit is configured to divide each atom in the first atom set into a class corresponding to the initial cluster center with the smallest distance from the atom according to the distance between each atom and each initial cluster center, and calculate the cluster center of each class after re-division to obtain M second cluster centers.

[0167] The spherical point set generation sub-unit is configured to generate a corresponding spherical point set with the M second cluster centers as the center points according to a first preset distance, to obtain M spherical point sets, the first preset distance being any distance within one fifth to one third of the distance between the two atoms with the largest distance interval in the ligand protein.

[0168] The redundant point deletion sub-unit is configured to delete redundant points in each spherical point set, and combine the spherical point sets after deletion of the redundant points to obtain a second atom set.

[0169] The second cluster center selection sub-unit is configured to randomly select M atoms from the second atom set as third cluster centers.

[0170] The second cluster center calculation sub-unit is configured to divide each atom in the second atom set into a class corresponding to the third cluster center with the smallest distance from the atom according to the distance between each atom and each third cluster center, and calculate the cluster center of each class after re-division to obtain M fourth cluster centers, the fourth cluster centers being sub-space center points.

[0171] In some other embodiments of the present application, the posture search unit 83 includes:

[0172] The population size initialization sub-unit is configured to initialize the population size, and randomly generate N individuals within a first preset distance range of the center of each sub-space, the individuals being ligand proteins with different postures.

[0173] a pose iterative optimization subunit configured to determine docking energies of the N individuals with the receptor protein in different poses in the subspace by iteratively optimizing poses of the N individuals through an adaptive differential algorithm, wherein the docking energies correspond to the different individual poses;

[0174] a sorting subunit configured to sort the individuals in different poses from high to low according to the final determined docking energies;

[0175] a candidate ligand protein conformation selection subunit configured to select poses of the Q individuals in the front of the sorting as candidate ligand protein conformations generated by the subspace, wherein Q≤N.

[0176] In some embodiments of the present application, the pose iterative optimization subunit comprises:

[0177] a six-dimensional vector obtaining subunit configured to obtain a six-dimensional vector representing the pose of the individual, wherein the six-dimensional vector comprises a first three-dimensional vector and a second three-dimensional vector, the first three-dimensional vector represents movement amounts of all atomic coordinates in the individual, and the second three-dimensional vector represents an overall rotation of the individual;

[0178] a six-dimensional vector optimization subunit configured to optimize the six-dimensional vector corresponding to the individual through an adaptive differential algorithm to obtain an optimized six-dimensional vector;

[0179] a position information confirmation subunit configured to determine first position information of the individual according to the optimized six-dimensional vector, wherein the first position information is position information of all atoms in the individual after the pose change;

[0180] a docking energy confirmation subunit configured to determine the docking energy of the individual with the receptor protein in the subspace according to the first position information and second position information, wherein the second position information is position information of the receptor protein.

[0181] In some embodiments of the present application, the pose iterative optimization subunit comprises:

[0182] an individual updating subunit configured to perform a crossover mutation operation on at least one dimension of the six-dimensional vector of the individual based on a crossover probability corresponding to the individual through an adaptive differential algorithm to obtain an updated individual;

[0183] a population updating subunit configured to perform an individual selection operation according to the docking energies of the individuals before and after the update, and to update the population by taking the individual with the highest docking energy as an individual in the current population;

[0184] a population decay subunit configured to perform a population size decay operation on the updated population.

[0185] In some embodiments of the present application, the pose iterative optimization subunit further comprises:

[0186] a weight value determination subunit configured to determine the weight value corresponding to the selected individual according to the absolute value of the difference between the docking energy of the individual before the update and the docking energy of the individual after the update;

[0187] a parameter update subunit configured to update the first mean value and the second mean value according to the weight value, wherein the first mean value is the mean value of the normal distribution corresponding to the scaling factor, and the second mean value is the mean value of the normal distribution corresponding to the crossover probability.

[0188] In the embodiments of the present application, one of the two proteins with relatively large number of atoms is kept as the receptor protein, and the surface atoms of the receptor protein are divided into several subspaces, and the pose search of the ligand protein is independently and in parallel performed in the several subspaces, the candidate ligand protein conformations generated in each subspace are determined, the search time of the protein-protein docking pose is greatly shortened, the docking efficiency of the receptor protein and the ligand protein is improved, and the accuracy of the receptor protein and the ligand protein is improved by merging the candidate ligand protein conformations generated in each subspace and determining the pose of the ligand protein best docked with the receptor protein from the merged candidate ligand protein conformations.

[0189] It should be noted that the information interaction, execution process and the like between the above modules are based on the same concept as the method embodiments of the present application, and the specific functions and the technical effects brought by the same can be referred to the method embodiments part, which will not be described here.

[0190] Figure 9 is a schematic diagram of the protein-protein docking device based on multi-region division provided by the embodiments of the present application. As shown in Figure 9 the protein-protein docking device 9 based on multi-region division of the embodiments of the present application includes a processor 90, a memory 91, and a computer program 92 stored in the memory 91 and executable on the processor 90, such as a protein-protein docking program. The processor 90 implements the steps in each of the above protein-protein docking method embodiments based on multi-region division when executing the computer program 92, such as Figure 1 steps 101-104 shown in the figure. Alternatively, the processor 90 implements the functions of each module / unit in each of the above system embodiments when executing the computer program 92, such as Figure 8 the functions of units 81-84 shown in the figure.

[0191] For example, the computer program 92 can be divided into one or more modules / units, one or more modules / units are stored in the memory 91 and executed by the processor 90 to complete the present application. One or more modules / units can be a series of computer program instruction segments capable of completing a specific function, which are used to describe the execution process of the computer program 92 in the protein-protein docking device based on multi-region division 9. For example, the computer program 92 can be divided into protein labeling unit 81, subspace division unit 82, pose search unit 83, pose determination unit 84, and the specific functions of each unit are described in detail in the corresponding embodiments, which will not be repeated here. Figure 1 The corresponding embodiments are described in detail in the corresponding embodiments, which will not be repeated here.

[0192] The protein-protein docking device based on multi-region division can include, but is not limited to, the processor 90, the memory 91. Those skilled in the art can understand that, Figure 9 The protein-protein docking device based on multi-region division 9 is only an example and does not constitute a limitation on the protein-protein docking device based on multi-region division 9, which can include more or fewer components than the illustration, or combine certain components, or different components, for example, the protein-protein docking device based on multi-region division can also include input / output devices, network access devices, buses, etc.

[0193] The processor 90 can be a central processing unit (CPU), and can also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor.

[0194] The memory 91 can be an internal storage unit of the multi-region partition based protein-protein docking device 9, such as a hard disk or a memory of the multi-region partition based protein-protein docking device 9. The memory 91 can also be an external storage device of the multi-region partition based protein-protein docking device 9, such as a plug-in hard disk, an SMC (SMart Media Card), an SD (Secure Digital) card, a flash card, etc. equipped on the multi-region partition based protein-protein docking device 9. Further, the memory 91 can include both the internal storage unit and the external storage device of the multi-region partition based protein-protein docking device 9. The memory 91 is used to store computer programs and other programs and data required by the multi-region partition based protein-protein docking device. The memory 91 can also be used to temporarily store data that has been output or will be output.

[0195] The embodiment of the present application also provides a computer readable storage medium, which stores a computer program. The computer program is executed by a processor to implement the multi-region partition based protein-protein docking method.

[0196] The embodiment of the present application provides a computer program product. When the computer program product is run on the multi-region partition based protein-protein docking device, the multi-region partition based protein-protein docking device is caused to implement the multi-region partition based protein-protein docking method.

[0197] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above functional units and modules is exemplified, and in actual application, the above functions can be completed by different functional units and modules according to needs, that is, the internal structure of the system is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit. In addition, the specific names of the functional units and modules are only for easy distinction, and do not limit the protection scope of the present application. The specific working process of the units and modules in the system can refer to the corresponding process in the foregoing method embodiments, which will not be described here.

[0198] In the above embodiments, the description of each embodiment has its own emphasis, and the parts not described or recorded in detail in a certain embodiment can be referred to the related description of other embodiments.

[0199] Those skilled in the art can understand that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be realized in electronic hardware or a combination of computer software and electronic hardware. Whether the functions are realized in hardware or software manner depends on the specific application and design constraints of the technical solutions. The skilled person can use different methods to realize the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.

[0200] The above-described embodiments are only used to illustrate the technical solutions of the present application, but not limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that the technical solutions recorded in the foregoing embodiments can be modified, or some technical features can be replaced by equivalent replacements; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should be included in the protection scope of the present application.

Claims

1. A method for protein-protein docking based on multi-region partitioning, characterized in that, The method comprises: According to the number of atoms of the target protein to be docked, one of the target proteins is recorded as a receptor protein, and the other is recorded as a ligand protein, wherein the number of atoms of the receptor protein is greater than that of the ligand protein, and the posture of the receptor protein remains unchanged; spatially partitioning surface atoms of the acceptor protein to obtain M subspaces, wherein, M is an integer greater than 1; In M parallel in each of the subspaces, and determine the candidate ligand protein conformations generated in each of the subspaces; Will M The candidate ligand protein conformations generated in the subspace are merged, and the optimal orientation of the ligand protein for docking with the receptor protein is determined from the merged set of candidate ligand protein conformations. The surface atoms of the acceptor protein are spatially divided to obtain M subspaces, including: Randomly select from the first set of atoms M One atom serves as the initial cluster center, and the first set of atoms is the set of all atoms extracted from the surface of the receptor protein; According to the distance between each atom and each of the initial cluster centers, each atom in the first set of atoms is divided into a category corresponding to the initial cluster center with the smallest distance, and the cluster center of each category after re-division is calculated, obtaining M second cluster centers; According to the first preset distance, M The second cluster center is used as the center point to generate a corresponding spherical point set, resulting in M A set of spherical points, wherein the first preset distance is any distance within the range of one-fifth to one-third of the distance between the two atoms with the largest distance interval in the ligand protein; Delete redundant points in each spherical point set, and merge the spherical point sets after deleting redundant points to obtain a second atomic set; Randomly select from the second set of atoms M One atom serves as the third cluster center; According to the distance between each atom and each third clustering center, each atom in the second atom set is divided into the category corresponding to the third clustering center with the smallest distance, and the clustering center of each category after re-division is calculated, obtaining M fourth clustering centers, the fourth clustering centers being subspace center points.

2. The multi-region partition based protein-protein docking method of claim 1, wherein, The in M The posture search of the ligand protein in each of the subspaces is performed in parallel, and the candidate ligand protein conformations generated in each of the subspaces are determined. For M In any one of the subspaces, a method is performed to determine the candidate ligand protein conformations generated by each of the subspaces, and a population size is initialized. Within a first preset distance range from the center of each of the subspaces, a population size of N individuals is randomly generated, and the individuals are ligand proteins in different poses. Using an adaptive difference algorithm, N The poses of the individuals are iteratively optimized to determine the position within the subspace. N The docking energy of the individual with the receptor protein in different postures, with different docking energies corresponding to different individual postures; According to the final determined docking energy, the individuals in different postures are sorted from high to low; Selecting a sorting order in advance Q a pose of each of the individuals as a candidate ligand protein conformation produced by the subspace, Q N .​ 3. The multi-region partition based protein-protein docking method of claim 2, wherein, The iterative optimization of the poses of the individuals in the subspace is performed by an adaptive differential algorithm. N The docking energies of the individuals in the subspace are determined by an adaptive differential algorithm. N The docking energies of the individuals in the subspace are determined by an adaptive differential algorithm. for N For any one of the individuals, the following method is performed to determine the docking energy between each individual and the receptor protein in the subspace, and to obtain a six-dimensional vector characterizing the attitude of the individual. The six-dimensional vector includes a first three-dimensional vector and a second three-dimensional vector. The first three-dimensional vector represents the amount of movement of all atomic coordinates in the individual, and the second three-dimensional vector represents the overall rotation of the individual. By an adaptive differential algorithm, the six-dimensional vector corresponding to the individual is optimized to obtain an optimized six-dimensional vector; According to the optimized six-dimensional vector, the first position information of the individual is determined, and the first position information is the position information of all atoms in the individual after the posture change; According to the first position information and the second position information, the docking energy of the individual in the subspace with the receptor protein is determined, and the second position information is the position information of the receptor protein.

4. The multi-region partition based protein-protein docking method of claim 3, wherein, The iterative optimization of the poses of the individuals in the subspace is performed by an adaptive differential algorithm. N The iterative optimization of the poses of the individuals in the subspace is performed by an adaptive differential algorithm. N The iterative optimization of the poses of the individuals in the subspace is performed by an adaptive differential algorithm. By an adaptive differential algorithm, at least one dimension of the six-dimensional vector of the individual is subjected to a crossover mutation operation based on the crossover probability corresponding to the individual, to obtain an updated individual; According to the docking energy of the individual before and after the update, the individual selection operation is performed, and the individual with the highest docking energy is selected as the individual in the current population, and the population is updated; The population size of the updated population is attenuated.

5. The multi-region partition based protein-protein docking method of claim 4, wherein, When the individual selection operation is performed according to the docking energy of the individual before and after the update, the individual with the highest docking energy is selected as the individual in the current population, and the population is updated, comprising: According to the absolute value of the difference value of the docking energy corresponding to the individual before and after the update, the weight value corresponding to the selected individual is determined, and the first mean and the second mean are updated according to the weight value, wherein the first mean is the mean of the normal distribution corresponding to the scaling factor, and the second mean is the mean of the normal distribution corresponding to the crossover probability.

6. A multi-region partition based protein-protein docking system, comprising: The system comprises: A protein marking unit is configured to mark one of the target proteins as a receptor protein and the other as a ligand protein according to the number of atoms of the target protein to be docked, wherein the number of atoms of the receptor protein is greater than that of the ligand protein, and the posture of the receptor protein remains unchanged; A subspace division unit is configured to divide surface atoms of the receptor protein into a plurality of subspaces, to obtain M wherein, M is an integer greater than 1. A posture searching unit is configured to search for the posture of the ligand protein in M subspaces in parallel to determine candidate ligand protein conformations generated by each of the subspaces; The attitude determination unit is used to determine the attitude of the target. M The candidate ligand protein conformations generated in the subspace are merged, and the optimal orientation of the ligand protein for docking with the receptor protein is determined from the merged set of candidate ligand protein conformations. The subspace division unit comprises: The first cluster center selection subunit is used to randomly select from the first set of atoms. M One atom serves as the initial cluster center, and the first set of atoms is the set of all atoms extracted from the surface of the receptor protein; a first cluster center calculation subunit, configured to divide each atom in the first atom set into a category corresponding to an initial cluster center with the smallest distance between the atom and the initial cluster center according to the distance between each atom and each initial cluster center, and calculate a cluster center of each category after the division, to obtain M second cluster centers; The ball point set generating sub-unit is configured to generate a corresponding ball point set with a first preset distance, wherein the first preset distance is any distance within a range of one fifth to one third of a distance between two atoms with the largest distance interval in the ligand protein, and the second clustering center is a center point of the corresponding ball point set. M The ball point set generating sub-unit is configured to generate a corresponding ball point set with a first preset distance, wherein the first preset distance is any distance within a range of one fifth to one third of a distance between two atoms with the largest distance interval in the ligand protein, and the second clustering center is a center point of the corresponding ball point set. M The ball point set generating sub-unit is configured to generate a corresponding ball point set with a first preset distance, wherein the first preset distance is any distance within a range of one fifth to one third of A redundant point deletion subunit is configured to delete redundant points in each spherical point set, and merge the spherical point sets after deleting redundant points to obtain a second atomic set; The second cluster center selection subunit is used to randomly select from the second set of atoms. M One atom serves as the third cluster center; The second cluster center calculation subunit is configured to divide each atom in the second atom set into a category corresponding to a third cluster center with the smallest distance between the atom and the third cluster center according to the distance between each atom and each third cluster center, and calculate a cluster center of each category after the division, to obtain M a fourth cluster center, which is a subspace center point.

7. A protein-protein docking apparatus based on multi-region partitioning, comprising a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein, The processor executes the computer program to implement the protein-protein docking method based on multi-region division according to any one of claims 1 to 5.

8. A computer-readable storage medium storing a computer program, the computer-readable storage medium comprising: The computer program is executed by the processor to implement the protein-protein docking method based on multi-region division according to any one of claims 1 to 5.