Methods and systems for increasing the probability of blocking binding between a protein and a ligand

CN117672347BActive Publication Date: 2026-08-28SHENZHEN TAILI BIOTECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211036751.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-29
Publication Date
2026-08-28
Estimated Expiration
2042-08-29

AI Technical Summary

Technical Problem

[0006]针对现有技术的以上缺陷或改进需求,本发明提供了用于提高蛋白质与配体之间结合的阻断概率的方法,其目的在于,解决现有基于蛋白质的三级结构计算掩蔽肽覆盖率的方法中存在的覆盖率不能准确反映蛋白质与配体的阻断概率,导致应用其设计的改造的蛋白质,其与配体的阻断概率不能有效的提高的技术问题

Benefits of technology

[0039] (1) In this invention, hotspot atoms are divided into outer, middle and upper layers in space, and different sphere radii and corresponding thresholds for the number of new atoms in the sphere are set. This layered logic is adapted to the specific binding mode of proteins and ligands to a certain extent, thereby improving the accuracy of coverage in evaluating the protein-ligand blocking probability, and thus improving the blocking probability of the designed modified protein and ligand.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117672347B_ABST
    Figure CN117672347B_ABST
Patent Text Reader

Abstract

The application discloses a method for realizing high blocking probability between a protein and a ligand, comprising the following steps: obtaining all hotspot atoms of a primary protein and interaction atoms of a primary protein-ligand complex, then obtaining interaction planes thereof, dividing all the hotspot atoms of the primary protein into four layers of outer, middle, upper and lower according to the interaction planes, for each hotspot atom of the four layers, establishing a sphere with the coordinate of the hotspot atom as the center of the sphere, obtaining the number of newly added atoms contained in the sphere after the primary protein is reformed as the sphere newly added atom number corresponding to the hotspot atom, and judging whether the sphere newly added atom number of a hotspot atom is less than or equal to the sphere newly added atom number threshold of the hotspot atom from any hotspot atom in the four layers. The application can solve the technical problem of low coverage accuracy in the existing method for calculating the coverage of a masking peptide, thereby reducing the blocking probability of the combination between the protein and the ligand.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of biotechnology, and more specifically, relates to a method and system for increasing the probability of blocking the binding between proteins and ligands. Background Technology

[0002] During targeted drug therapy, since most target molecules are not only specifically expressed at the lesion site, but may also exist in normal cells or tissues outside the lesion site, this can lead to damage to normal tissues and toxic side effects when targeted drugs recognize antigens on normal tissues or cells. In view of this, targeted prodrugs have emerged.

[0003] Targeted prodrugs consist of three parts: the targeted drug itself, the linker, and the masking peptide. The masking peptide is used to mask the site where the targeted drug exerts its activity, thereby blocking the specific binding of the drug to the target molecule and achieving a non-toxic effect in normal tissues. The design of the masking peptide is a key step in the design of targeted prodrugs. Its core goal is to increase the probability of blocking the binding of the targeted drug to the target molecule. Specifically, this is achieved by modifying the protein of the targeted prodrug to increase the probability of blocking the binding of the protein to the ligand.

[0004] To achieve a high blocking probability between proteins and ligands, amino acid sequences (such as...) are typically added to both ends of the protein's primary sequence. Figure 2 As shown, the top is the ligand, the middle is the hotspot atom, and the bottom is the original protein. The resulting new protein is called the modified protein (e.g., ...). Figure 3As shown, the top part is the newly added fragment, the middle part is the hotspot atoms, and the bottom part is the original protein portion in the modified protein. The added amino acid sequence is called the newly added fragment, the atoms of the newly added fragment are called newly added atoms, and the protein before modification is called the original protein. Based on this, researchers found that the frequency of the newly added fragment blocking the active site of the original protein (i.e., coverage) is positively correlated with the blocking probability. Based on this, the paper "Development of a structure-based computational simulation to optimize the blocking efficacy of pro-antibodies" published in Chemical Science in 2021 disclosed a method for calculating the coverage of the masking peptide (i.e., the newly added fragment) on the antibody's complementarity-determining regions (CDRs) based on the protein's tertiary structure. This method first finds the crystal structure of the antibody's Fv region, then uses molecular dynamics homology modeling to simulate the tertiary structure of the pro-antibody, and then calculates the frequency (i.e., coverage) of the CDR residues directly above them by the masking peptide based on the model snapshot. This coverage is positively correlated with the masking factor of the pro-antibody and can be used as a scoring function for screening pro-antibodies.

[0005] However, the above method still has some significant drawbacks: First, for the modified protein, it only sets a spatial distance threshold when determining whether a hotspot atom belongs to the covered atom, without considering special cases. Specifically, when the spatial distance threshold between the hotspot atom in the middle and the newly added atom is set larger, it can still effectively block the binding of the ligand to the hotspot region, thus underestimating the blocking probability. Second, because it only considers the spatial distance between atoms and ignores the binding directionality when determining whether a hotspot atom belongs to the covered atom, the obtained coverage rate is inaccurate. Third, because it only considers whether there are newly added atoms within a certain spatial distance from the hotspot atom when determining whether a hotspot atom belongs to the covered atom, and ignores the number of atoms, it underestimates the blocking probability. Since existing calculation methods cannot accurately evaluate the blocking probability, the blocking probability of the modified protein designed using these methods cannot be effectively improved. Summary of the Invention

[0006] In view of the above-mentioned defects or improvement needs of the prior art, the present invention provides a method for increasing the blocking probability of binding between proteins and ligands. The purpose is to solve the technical problem that the coverage of existing methods for calculating the coverage of masked peptides based on the tertiary structure of proteins cannot accurately reflect the blocking probability between proteins and ligands, resulting in the inability to effectively improve the blocking probability of modified proteins designed using such methods.

[0007] To achieve the above objectives, according to one aspect of the present invention, a method for increasing the blocking probability of binding between a protein and a ligand is provided, comprising the following steps:

[0008] (1) Obtain all hotspot atoms of the original protein and the interacting atoms of the original protein-ligand complex;

[0009] (2) Obtain the interaction plane of the original protein-ligand complex based on the interacting atoms obtained in step (1);

[0010] (3) Based on the interaction plane obtained in step (2), all hot spots of the original protein are divided into three layers: outer, middle and inner. Based on the distance between the hot spots of the inner layer and the interaction plane, the hot spots of the inner layer are divided into two layers: upper and lower, thus forming four layers: outer, middle, upper and lower.

[0011] (4) For each hot spot atom in the four layers obtained in step (3), a sphere with the coordinates of the hot spot atom as the center is constructed, and the number of newly added atoms contained in the sphere after the original protein is modified is obtained as the number of newly added atoms in the sphere corresponding to the hot spot atom.

[0012] (5) Select any hot spot atom from the four layers obtained in step (3). For the hot spot atom, determine whether the number of new atoms added to the sphere of the hot spot atom obtained in step (4) is less than or equal to the threshold of the number of new atoms added to the sphere of the hot spot atom. If so, put the hot spot atom into the covered atom set and then proceed to step (6); otherwise, proceed to step (6).

[0013] (6) For the remaining hot spots in the four layers obtained in step (3), repeat step (5) until all hot spots have been processed.

[0014] (7) Calculate the coverage rate based on the total number of hot spot atoms in the covered atom set and the total number of hot spot atoms in the original protein obtained in step (1);

[0015] (8) Construct a search space consisting of multiple variable parameters;

[0016] (9) Obtain N masking factors between N modified proteins and ligands. Specify M different combinations of variable parameters in the search space obtained in step (8). For each combination of variable parameters, calculate N coverage rates through steps (3) to (7), thereby obtaining one Spearman correlation coefficient between the N coverage rates and the N masking factors. Finally, obtain M Spearman correlation coefficients corresponding to the M combinations of variable parameters, and select the combination of variable parameters corresponding to the largest Spearman correlation coefficient. Where N and M are arbitrary natural numbers.

[0017] (10) Based on the combination of variable parameters obtained in step (9), and combined with the above steps (1) to (7), obtain the final coverage rate, and adjust the original protein according to the final coverage rate.

[0018] Preferably, the interacting atoms of the original protein-ligand complex refer to the atoms in the region where the original protein binds to the corresponding ligand.

[0019] Preferably, the process of dividing the outer, middle and inner layers in step (3) firstly involves projecting all the hot spots of the original protein onto the interaction plane, then using an edge detection algorithm to determine the projection points on the outermost layer of the interaction plane, and using the hot spots corresponding to these projection points as the hot spots of the outer layer. Finally, based on the distance between the remaining projection points on the interaction plane and the projection points of the outermost layer, the hot spots of the middle layer and the hot spots of the inner layer in their corresponding spaces are determined.

[0020] Preferably, if the distance between a projection point on the interaction plane and the outermost projection point is less than or equal to a certain first preset value, then the hot spot atom corresponding to the projection point is a hot spot atom in the middle layer; if the distance between a projection point on the interaction plane and the outermost projection point is greater than a certain preset value, then the hot spot atom corresponding to the projection point is a hot spot atom in the inner layer; the initial value of the first preset value is freely set.

[0021] If the distance between an inner hotspot atom and the interaction plane is less than or equal to a certain second preset value, then the hotspot atom is a hotspot atom of the upper layer; if the distance between an inner hotspot atom and the interaction plane is greater than a certain second preset value, then the hotspot atom is a hotspot atom of the lower layer; the initial value of the second preset value is freely set.

[0022] Preferably, for the same layer, the spheres corresponding to all hotspot atoms in the layer have the same radius;

[0023] The radii of the spheres corresponding to hot spot atoms in different layers are not the same, and the radius of the spheres corresponding to hot spot atoms in the outer layer is less than the radius of the spheres corresponding to hot spot atoms in the inner layer and less than the radius of the spheres corresponding to hot spot atoms in the upper and lower layers.

[0024] The initial values ​​of the four radii are freely set.

[0025] Preferably, the variable parameters include a first preset value, a second preset value, the radius of the sphere corresponding to the outer hot spot atom, the radius of the sphere corresponding to the inner hot spot atom, the radius of the sphere corresponding to the upper hot spot atom, and the radius of the sphere corresponding to the lower hot spot atom.

[0026] Furthermore, the variable parameters also include the antigen-antibody distance value set when determining the hotspot atoms, the parameter alpha of the edge detection algorithm, and the number of newly added atoms in the spheres of each of the four hotspot atoms.

[0027] According to another aspect of the invention, a system for increasing the blocking probability of binding between a protein and a ligand is provided, comprising:

[0028] The first module is used to obtain all hotspot atoms of the original protein, as well as the interacting atoms of the original protein-ligand complex;

[0029] The second module is used to obtain the interaction plane of the original protein-ligand complex based on the interacting atoms of the original protein-ligand complex obtained from the first module;

[0030] The third module is used to divide all the hot spots of the original protein into three layers (outer, middle, and inner) based on the interaction plane obtained from the second module, and to divide the hot spots of the inner layer into two layers (upper and lower) based on the distance between the hot spots of the inner layer and the interaction plane, thus forming four layers: outer, middle, upper, and lower.

[0031] The fourth module is used to construct a sphere with the coordinates of the hot spot atom as the center for each hot spot atom in the four layers obtained in the third module, and obtain the number of newly added atoms contained in the sphere after the original protein is modified as the number of newly added atoms in the sphere corresponding to the hot spot atom.

[0032] The fifth module is used to select any hot spot atom from the four layers obtained from the third module. For the hot spot atom, it is determined whether the number of newly added atoms in the sphere of the hot spot atom obtained from the fourth module is less than or equal to the threshold number of newly added atoms in the sphere of the hot spot atom. If so, the hot spot atom is placed into the set of covered atoms and then the process proceeds to the sixth module; otherwise, the process proceeds to the sixth module.

[0033] The sixth module is used to repeat the fifth module above for the remaining hot spots in the four layers obtained in the third module, until all hot spots have been processed.

[0034] The seventh module is used to calculate the coverage rate based on the total number of hotspot atoms in the covered atom set and the total number of hotspot atoms in the original protein obtained by the first module.

[0035] The eighth module is used to construct the search space consisting of multiple variable parameters;

[0036] The ninth module is used to obtain N masking factors between N modified proteins and ligands. In the search space obtained by the eighth module, M different combinations of variable parameters are specified. For each combination of variable parameters, N coverage rates are calculated by the third to seventh modules, thereby obtaining one Spearman correlation coefficient between the N coverage rates and the N masking factors. Finally, M Spearman correlation coefficients corresponding to the M combinations of variable parameters are obtained, and the combination of variable parameters corresponding to the largest Spearman correlation coefficient is selected. Here, N and M are arbitrary natural numbers.

[0037] The tenth module is used to obtain the final coverage based on the variable parameter combination obtained from the ninth module, combined with the first to seventh modules mentioned above, and to adjust the original protein according to the final coverage.

[0038] In summary, compared with the prior art, the above-described technical solutions conceived by this invention can achieve the following beneficial effects:

[0039] (1) In this invention, hotspot atoms are divided into outer, middle and upper layers in space, and different sphere radii and corresponding thresholds for the number of new atoms in the sphere are set. This layered logic is adapted to the specific binding mode of proteins and ligands to a certain extent, thereby improving the accuracy of coverage in evaluating the protein-ligand blocking probability, and thus improving the blocking probability of the designed modified protein and ligand.

[0040] (2) The hierarchical logic of hot spot atoms in this invention is based on the interaction plane of protein-ligand complex. It fully considers the directionality of protein-ligand binding (which is also the effective direction for the new fragment to cover hot spot atoms), thereby improving the accuracy of coverage in evaluating the protein-ligand blocking probability, and thus improving the blocking probability of the designed modified protein and ligand.

[0041] (3) By introducing spheres and the number of new atoms in the spheres, this invention can reduce the adverse effects of ignoring the positional orientation of hotspot atoms and new atoms to a certain extent, thereby improving the accuracy of coverage in evaluating the protein-ligand blocking probability, and thus improving the blocking probability of the designed modified protein and ligand.

[0042] (4) The present invention uses the positive correlation between the hot spot atom coverage of the original protein corresponding to the modified protein and the protein ligand binding blocking probability to measure the accuracy of coverage calculation. The Spearman correlation coefficient between the coverage calculated by the method of the present invention and the masking factor obtained by the experiment can reach more than 70%. Attached Figure Description

[0043] Figure 1 This is a flowchart of the method of the present invention for increasing the blocking probability of binding between proteins and ligands;

[0044] Figure 2 This is a schematic diagram of the preprotein and its ligand;

[0045] Figure 3 This is a schematic diagram of the protein and ligand.

[0046] Figure 4 It is the Spearman correlation coefficient between coverage and masking factor calculated in the embodiments of the present invention. Detailed Implementation

[0047] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention. Furthermore, the technical features involved in the various embodiments of this invention described below can be combined with each other as long as they do not conflict with each other.

[0048] like Figure 1 As shown, the present invention provides a method for increasing the blocking probability of binding between a protein and a ligand, comprising the following steps:

[0049] (1) Obtain all hotspot atoms of the original protein and the interacting atoms of the original protein-ligand complex;

[0050] Specifically, when the original protein binds to the corresponding ligand, the atoms in the binding region are called the interacting atoms of the original protein-ligand complex.

[0051] (2) Obtain the interaction plane of the original protein-ligand complex based on the interacting atoms obtained in step (1);

[0052] (3) Based on the interaction plane obtained in step (2), all hot spots of the original protein are divided into three layers: outer, middle and inner. Based on the distance between the hot spots of the inner layer and the interaction plane, the hot spots of the inner layer are divided into two layers: upper and lower, thus forming four layers: outer, middle, upper and lower.

[0053] Specifically, the process of dividing the original protein into three layers in this step is as follows: First, all the hot spots of the original protein are projected onto the interaction plane. Then, the edge detection algorithm is used to determine the projection points on the outermost layer of the interaction plane. The hot spots corresponding to these projection points are taken as the hot spots of the outer layer. Finally, the hot spots of the middle layer and the hot spots of the inner layer are determined according to the distance of the remaining projection points on the interaction plane from the projection points of the outermost layer.

[0054] Furthermore, if the distance between a projection point on the interaction plane and the outermost projection point is less than or equal to a certain first preset value, then the hot spot atom corresponding to that projection point is a hot spot atom in the middle layer. If the distance between a projection point on the interaction plane and the outermost projection point is greater than a certain preset value, then the hot spot atom corresponding to that projection point is a hot spot atom in the inner layer. It should be noted that the initial value of the first preset value can be freely set and will be continuously updated in subsequent steps.

[0055] If the distance between an inner hotspot atom and the interaction plane is less than or equal to a certain second preset value, then the hotspot atom is a hotspot atom of the upper layer. If the distance between an inner hotspot atom and the interaction plane is greater than a certain second preset value, then the hotspot atom is a hotspot atom of the lower layer. It should be noted that the initial value of the second preset value can be set freely and will be continuously updated in subsequent steps.

[0056] (4) For each hot spot atom in the four layers obtained in step (3), a sphere with the coordinates of the hot spot atom as the center is constructed, and the number of newly added atoms contained in the sphere after the original protein is modified is obtained as the number of newly added atoms in the sphere corresponding to the hot spot atom.

[0057] Specifically, within the same layer, all spheres corresponding to hotspot atoms have the same radius. However, the spheres corresponding to hotspot atoms in different layers have different radii, with the radius of the sphere corresponding to the hotspot atoms in the outer layer ≤ the radius of the sphere corresponding to the hotspot atoms in the inner layer < the radius of the sphere corresponding to the hotspot atoms in the upper and lower layers. The initial values ​​of the four radii can be freely set and will be continuously updated in subsequent steps.

[0058] (5) Select any hot spot atom from the four layers obtained in step (3). For the hot spot atom, determine whether the number of new atoms added to the sphere of the hot spot atom obtained in step (4) is less than or equal to the threshold of the number of new atoms added to the sphere of the hot spot atom. If so, put the hot spot atom into the covered atom set and then proceed to step (6); otherwise, proceed to step (6).

[0059] (6) For the remaining hot spots in the four layers obtained in step (3), repeat step (5) until all hot spots have been processed.

[0060] (7) Calculate the coverage rate based on the total number of hot spot atoms in the covered atom set and the total number of hot spot atoms in the original protein obtained in step (1);

[0061] Specifically, coverage = total number of hotspot atoms in the covered atom set / total number of hotspot atoms in the original protein * 100%.

[0062] (8) Construct a search space consisting of multiple variable parameters;

[0063] The variable parameters include a first preset value, a second preset value, the radius of the sphere corresponding to the outer hot spot atom, the radius of the sphere corresponding to the inner hot spot atom, the radius of the sphere corresponding to the upper hot spot atom, the radius of the sphere corresponding to the lower hot spot atom, the antigen-antibody distance value set when determining the hot spot atom, the parameter alpha of the edge detection algorithm, and the number of newly added atoms in the spheres of each of the four hot spot atom layers.

[0064] (9) Obtain N masking factors between N modified proteins (where N is any natural number) and ligands (which are obtained through biological experiments). Specify M different combinations of variable parameters in the search space obtained in step (8) (where M can be any natural number, the larger the value, the higher the accuracy of the final result, and vice versa). For each combination of variable parameters, calculate N coverage rates through the above steps (3) to (7), thereby obtaining 1 Spearman correlation coefficient between N coverage rates and N masking factors, and finally obtain M Spearman correlation coefficients corresponding to the M combinations of variable parameters. Select the combination of variable parameters corresponding to the largest Spearman correlation coefficient. (10) According to the combination of variable parameters obtained in step (9), combine the above steps (1) to (7) to obtain the final coverage rate, and adjust the original protein according to the final coverage rate.

[0065] Specifically, after obtaining the variable parameter combination in step (9), these variable parameter combinations are used as the first preset value, the second preset value, the radius of the sphere corresponding to the outer hot spot atom, the radius of the sphere corresponding to the inner hot spot atom, the radius of the sphere corresponding to the upper hot spot atom, the radius of the sphere corresponding to the lower hot spot atom, the antigen-antibody distance value set when determining the hot spot atom, the parameter alpha of the edge detection algorithm, and the number of newly added atoms in the spheres of each of the four hot spot atoms. Based on these parameter combinations, the above steps (1) to (7) are repeated to calculate the final coverage rate. When the coverage rate does not reach the ideal value, the blocking probability can be improved by adjusting the composition of the original protein.

[0066] Example

[0067] A method for increasing the blocking probability of protein-ligand binding, taking the screening of Abagovomab preantibodies with high blocking probabilities as an example, includes the following steps:

[0068] Data on the blocking probabilities of several pre-antibodies, measured through biological experiments, were collected. The original antibodies (hereinafter referred to as Abs) and masking peptides (hereinafter referred to as caps) of all pre-antibodies (hereinafter referred to as cap_Abs) were obtained. The spatial structures of the complexes of all Abs and their corresponding antigens were obtained from a protein structure database, and the spatial structures of all cap_Abs were obtained using a protein structure prediction tool, and saved as PDB files.

[0069] Read the PDB files of all Abs and their corresponding antigen complexes, construct a three-dimensional coordinate system, obtain the coordinates of each atom of the complex in the three-dimensional coordinate system for each Ab, determine the atoms that are close to the antibody and antigen by calculating the spatial distance between the atoms, define these atoms as interacting atoms, and define the atoms on the Abs as hotspot atoms.

[0070] Interacting atoms are divided into two categories: antibody atoms and antigen atoms. The machine learning algorithm of support vector machine is used to calculate the hyperplanes of the interacting atoms based on the above two categories, and these hyperplanes are defined as interaction planes.

[0071] In a three-dimensional coordinate system, the coordinates of hotspot atoms are projected onto the interaction plane. The Alpha Shape2D point set edge extraction algorithm is used to determine the outermost projection points on the plane, thus identifying the outermost atoms in space. The distances from the remaining projection points on the plane to the outermost atom projection points are then calculated. Atoms corresponding to points on the plane closer to the outermost atom projection points are defined as the middle-layer hotspot atoms, while the atoms corresponding to the remaining points farther away are defined as the inner-layer hotspot atoms. The distances from the coordinates of the inner-layer hotspot atoms in space to the interaction plane are then calculated, dividing the inner-layer hotspot atoms into upper and lower layers. Thus, all hotspot atoms in space are divided into four layers.

[0072] Read all cap_Ab PDB files, construct a three-dimensional coordinate system, obtain the coordinates of each atom in the three-dimensional coordinate system for each cap_Ab, align cap_Ab and Ab in the same three-dimensional coordinate system by residue alignment, find the hot spot atoms and their layering information on the corresponding Ab on cap_Ab, and obtain the coordinates of all atoms in cap.

[0073] In a three-dimensional coordinate system, the four hot spot atoms of cap_Ab are drawn into corresponding spheres with the atomic coordinates as the center according to their respective set radii. A threshold for the number of new atoms in each sphere is set for each type of sphere, and the number of new atoms in the spheres that fall on each hot spot atom of cap_Ab for all atoms of cap is counted.

[0074] If the number of newly added atoms in the sphere of a certain hot spot atom is within its corresponding threshold, then the hot spot atom is included in the set of covered atoms; Coverage rate = number of covered atoms / number of hot spot atoms * 100%.

[0075] Implement the above process using code, and construct the search space for the variable parameters involved in the above calculation process, including the threshold for the number of new atoms added to the sphere of the four-layer hot spot atoms, the parameter alpha of the Alpha Shape algorithm, the radius of the sphere of the four-layer hot spot atoms, the distance threshold between the outer projection point and the middle projection point on the interaction plane when defining the middle layer hot spot atoms, the spatial distance threshold between the antigen and antibody atoms when defining the hot spot atoms, and the distance threshold from the atomic coordinate point to the interaction plane when dividing the inner layer hot spot atoms into upper and lower layers.

[0076] For all cap_Abs, if the number is N, then N coverage rates are generated through the above process. The values ​​of N masking folds are obtained through biological experiments. A parameter combination is specified in the search space of variable parameters, and the coverage rates of several pre-antibodies with known masking folds are calculated through the above steps, thus obtaining the Spearman correlation coefficient between coverage and masking fold. The parameter combination is adjusted, and the parameter combination with the highest Spearman correlation coefficient is selected. Based on this parameter combination and the above calculation process, the final algorithm F for coverage can be obtained. Its inputs are the PDB files of the spatial structure of the Ab and its corresponding antigen complex, and the PDB file of the spatial structure of the cap_Ab. The output is the coverage rate of the cap_Ab against the corresponding antigen.

[0077] A sequence was added to the N-terminus of both the heavy and light chains of Abagovomab, and this new sequence was named cap. A candidate sequence library of the new sequence was constructed. Candidate preantibodies cap_Abagovomab corresponding to all candidate new fragments were obtained. The spatial structures of the complexes of Abagovomab and its corresponding antigens were obtained from a protein structure database. The spatial structures of all candidate cap_Abagovomab were obtained using a protein structure prediction tool and saved as PDB files.

[0078] The PDB files of the complex of Abagovomab and its corresponding antigen, along with the PDB files of all candidate cap_Abagovomab, are input into algorithm F. The algorithm outputs the coverage rates of all candidates and selects the cap_Abagovomab with the highest coverage rate as the pre-antibody of Abagovomab with a high blocking probability.

[0079] from Figure 4As can be seen, the Spearman correlation coefficient between the coverage rate calculated using the method of this invention (coverrate in the figure) and the masking fold obtained from the experimental results (blocking fold in the figure) is 0.767.

[0080] Those skilled in the art will readily understand that the above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A method for increasing the probability of blocking the binding between a protein and a ligand, characterized in that, Includes the following steps: (1) Obtain all hotspot atoms of the original protein and the interacting atoms of the original protein-ligand complex; (2) Obtain the interaction plane of the original protein-ligand complex based on the interacting atoms obtained in step (1); (3) Based on the interaction plane obtained in step (2), all hot spots of the original protein are divided into three layers: outer, middle and inner. Based on the distance between the hot spots of the inner layer and the interaction plane, the hot spots of the inner layer are divided into two layers: upper and lower, thus forming four layers: outer, middle, upper and lower. (4) For each hot spot atom in the four layers obtained in step (3), a sphere with the coordinates of the hot spot atom as the center is constructed, and the number of newly added atoms contained in the sphere after the original protein is modified is obtained as the number of newly added atoms in the sphere corresponding to the hot spot atom. (5) Select any hot spot atom from the four layers obtained in step (3). For the hot spot atom, determine whether the number of new atoms added to the sphere of the hot spot atom obtained in step (4) is less than or equal to the threshold of the number of new atoms added to the sphere of the hot spot atom. If so, put the hot spot atom into the covered atom set and then proceed to step (6); otherwise, proceed to step (6). (6) For the remaining hot spots in the four layers obtained in step (3), repeat step (5) until all hot spots have been processed. (7) Calculate the coverage rate based on the total number of hot spot atoms in the covered atom set and the total number of hot spot atoms in the original protein obtained in step (1); (8) Construct a search space consisting of multiple variable parameters; (9) Obtain N masking factors between N modified proteins and ligands. Specify M different combinations of variable parameters in the search space obtained in step (8). For each combination of variable parameters, calculate N coverage rates through steps (3) to (7), thereby obtaining one Spearman correlation coefficient between the N coverage rates and the N masking factors. Finally, obtain M Spearman correlation coefficients corresponding to the M combinations of variable parameters, and select the combination of variable parameters corresponding to the largest Spearman correlation coefficient. Where N and M are arbitrary natural numbers. (10) Based on the combination of variable parameters obtained in step (9), and combined with the above steps (1) to (7), obtain the final coverage rate, and adjust the original protein according to the final coverage rate.

2. The method for increasing the blocking probability of binding between proteins and ligands according to claim 1, characterized in that, The interacting atoms of the proprotein-ligand complex refer to the atoms in the region where the proprotein binds to the corresponding ligand.

3. The method for increasing the blocking probability of binding between a protein and a ligand according to claim 1 or 2, characterized in that, In step (3), the process of dividing the three layers into outer, middle, and inner layers involves first projecting all the hot spots of the original protein onto the interaction plane, then using an edge detection algorithm to determine the projection points on the outermost layer of the interaction plane, and using the hot spots corresponding to these projection points as the hot spots of the outer layer. Finally, based on the distance between the remaining projection points on the interaction plane and the projection points of the outermost layer, the hot spots of the middle layer and the hot spots of the inner layer in their corresponding spaces are determined.

4. The method for increasing the blocking probability of binding between a protein and a ligand according to any one of claims 1 to 3, characterized in that, If the distance between a projection point on the interaction plane and the outermost projection point is less than or equal to a certain first preset value, then the hot spot atom corresponding to that projection point is a hot spot atom in the middle layer. If the distance between a projection point on the interaction plane and the outermost projection point is greater than a certain preset value, then the hot spot atom corresponding to that projection point is a hot spot atom in the inner layer. The initial value of the first preset value is freely set. If the distance between an inner hotspot atom and the interaction plane is less than or equal to a certain second preset value, then the hotspot atom is a hotspot atom of the upper layer; if the distance between an inner hotspot atom and the interaction plane is greater than a certain second preset value, then the hotspot atom is a hotspot atom of the lower layer; the initial value of the second preset value is freely set.

5. The method for increasing the blocking probability of binding between proteins and ligands according to claim 4, characterized in that, For the same layer, the spheres corresponding to all hotspot atoms in that layer have the same radius; The radii of the spheres corresponding to hot spot atoms in different layers are not the same, and the radius of the spheres corresponding to hot spot atoms in the outer layer is less than the radius of the spheres corresponding to hot spot atoms in the inner layer and less than the radius of the spheres corresponding to hot spot atoms in the upper and lower layers. The initial values ​​of the four radii are freely set.

6. The method for increasing the blocking probability of binding between a protein and a ligand according to claim 5, characterized in that, The variable parameters include a first preset value, a second preset value, the radius of the sphere corresponding to the outer hot spot atom, the radius of the sphere corresponding to the inner hot spot atom, the radius of the sphere corresponding to the upper hot spot atom, and the radius of the sphere corresponding to the lower hot spot atom.

7. A system for increasing the probability of blocking binding between proteins and ligands, characterized in that, include: The first module is used to obtain all hotspot atoms of the original protein, as well as the interacting atoms of the original protein-ligand complex; The second module is used to obtain the interaction plane of the original protein-ligand complex based on the interacting atoms of the original protein-ligand complex obtained from the first module; The third module is used to divide all the hot spots of the original protein into three layers (outer, middle, and inner) based on the interaction plane obtained from the second module, and to divide the hot spots of the inner layer into two layers (upper and lower) based on the distance between the hot spots of the inner layer and the interaction plane, thus forming four layers: outer, middle, upper, and lower. The fourth module is used to construct a sphere with the coordinates of the hot spot atom as the center for each hot spot atom in the four layers obtained in the third module, and obtain the number of newly added atoms contained in the sphere after the original protein is modified as the number of newly added atoms in the sphere corresponding to the hot spot atom. The fifth module is used to select any hot spot atom from the four layers obtained from the third module. For the hot spot atom, it is determined whether the number of newly added atoms in the sphere of the hot spot atom obtained from the fourth module is less than or equal to the threshold number of newly added atoms in the sphere of the hot spot atom. If so, the hot spot atom is placed into the set of covered atoms and then the process proceeds to the sixth module; otherwise, the process proceeds to the sixth module. The sixth module is used to repeat the fifth module above for the remaining hot spots in the four layers obtained in the third module, until all hot spots have been processed. The seventh module is used to calculate the coverage rate based on the total number of hotspot atoms in the covered atom set and the total number of hotspot atoms in the original protein obtained by the first module. The eighth module is used to construct the search space consisting of multiple variable parameters; Module 9 is used to obtain N masking factors between N modified proteins and ligands. Within the search space obtained in Module 8, M different combinations of variable parameters are specified. For each combination of variable parameters, N coverage rates are calculated using Modules 3 through 7, resulting in one Spearman correlation coefficient between the N coverage rates and the N masking factors. Finally, M Spearman correlation coefficients corresponding to the M combinations of variable parameters are obtained, and the combination of variable parameters corresponding to the largest Spearman correlation coefficient is selected. Here, N and M are arbitrary natural numbers. The tenth module is used to obtain the final coverage based on the variable parameter combination obtained from the ninth module, combined with the first to seventh modules mentioned above, and to adjust the original protein based on the final coverage.

Citation Information

Patent Citations

  • Method for predicting unknown small molecule ligand target spot and application of method

    CN113963744A

  • Protein ligand binding atom recognition method and device based on artificial intelligence

    CN114649053A