Iterative nearest point-based antigen-antibody binding conformation optimization method and device
By using an iterative nearest-point antigen-antibody binding conformation optimization method, epitope models are generated through coarse-graining and surface residue extraction. Spatial matching and local optimization are then performed to solve the problem of low antibody conformation error tolerance in antigen-antibody prediction, thereby improving the accuracy of antibody conformation and the precision of antigen-antibody binding.
Patent Information
- Application Number
- CN202511429424.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-09
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2045-10-09
AI Technical Summary
Existing protein docking software is not effective in predicting antigens and antibodies, mainly because the complementarity-determining region (CDR) of the antibody is not very complementary to the antigen and has a high degree of flexibility, which leads to an over-reliance on rigid complementarity and a low error tolerance.
An antigen-antibody binding conformation optimization method based on iterative nearest point is adopted. By coarsely processing antigen and antibody structural files, surface residue information is extracted to generate a list of antibody binding region residues and epitope models. Spatial matching and local optimization are then performed to screen out the target antigen-antibody binding conformation.
It improves the accuracy and error tolerance of antibody conformation, enhances the geometric matching precision of antigen-antibody binding and the hit rate of near-native structures, and solves the problem of low matching degree between the overall antibody posture and antigen in existing methods.
Smart Images

Figure CN120913633A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Embodiments of the present disclosure relate to the technical field of antigen-antibody docking model and the technical field of computer, and particularly relate to an antigen-antibody binding conformation optimization method and device based on iterative closest point. BACKGROUND
[0002] Computational biology methods for protein structure prediction, commonly known as docking techniques. The prior art has adopted different ways in dealing with these stages, such as Monte Carlo simulation, Fast Fourier Transform (FFT), spherical harmonics, energy-based techniques, etc. have been applied in the field of docking to pursue better results. At the same time, the selection of scoring functions is also very diverse, such as based on geometric complementarity, force field, knowledge, machine learning, or a combination of these functions to identify near-native structures.
[0003] However, the current public protein docking software in the prediction of antigen-antibody is not as good as the conventional protein docking. This is mainly because the antibody mainly binds to the antigen epitope through the CDR (Complementarity-Determining Region) region, and the complementarity of the CDR region to the antigen is weaker than that of the conventional protein-protein interaction, and the CDR region (especially the H3 (Heavy Chain Complementarity-Determining Region 3) heavy chain complementarity determining region) has strong flexibility, and too much reliance on rigid complementarity will result in low antibody conformation fault tolerance. SUMMARY
[0004] The summary part of the present disclosure is used to introduce the concepts in a brief form, which will be described in detail in the specific embodiments part. The summary part of the present disclosure is not intended to identify key or essential features of the claimed technical solutions, nor is it intended to limit the scope of the claimed technical solutions.
[0005] Some embodiments of the present disclosure propose an antigen-antibody binding conformation optimization method and device based on iterative closest point to solve the technical problems mentioned in the background part.
[0006] In a first aspect, some embodiments of the present disclosure provide an antigen-antibody binding conformation optimization method based on iterative closest point, which comprises: performing coarse-grained processing on an antigen structure file and an antibody structure file to obtain a coarse-grained antibody file and a coarse-grained antigen file; performing surface residue extraction on the antigen structure file and the antibody structure file based on the coarse-grained antibody file and the coarse-grained antigen file to obtain an antigen surface residue information set and an antibody surface residue information set; generating an antibody binding region residue list according to the antibody surface residue information set; generating an epitope model based on the antibody binding region residue list; performing spatial matching on the antigen surface residue information set and the epitope model to obtain a matched epitope model; generating an antibody docking conformation set according to the matched epitope model; performing local optimization on each antibody docking conformation in the antibody docking conformation set by using the coarse-grained antigen file to obtain an optimized antibody conformation set; and screening each optimized antibody conformation in the optimized antibody conformation set according to the coarse-grained antigen file to obtain a target antigen-antibody binding conformation.
[0007] In a second aspect, some embodiments of the present disclosure provide an antigen-antibody binding conformation optimization device based on iterative closest point, which comprises: a coarse-grained processing unit configured to perform coarse-grained processing on an antigen structure file and an antibody structure file to obtain a coarse-grained antibody file and a coarse-grained antigen file; an extraction unit configured to perform surface residue extraction on the antigen structure file and the antibody structure file based on the coarse-grained antibody file and the coarse-grained antigen file to obtain an antigen surface residue information set and an antibody surface residue information set; a first generation unit configured to generate an antibody binding region residue list according to the antibody surface residue information set; a second generation unit configured to generate an epitope model based on the antibody binding region residue list; a spatial matching unit configured to perform spatial matching on the antigen surface residue information set and the epitope model to obtain a matched epitope model; a third generation unit configured to generate an antibody docking conformation set according to the matched epitope model; a local optimization unit configured to perform local optimization on each antibody docking conformation in the antibody docking conformation set by using the coarse-grained antigen file to obtain an optimized antibody conformation set; and a screening unit configured to screen each optimized antibody conformation in the optimized antibody conformation set according to the coarse-grained antigen file to obtain a target antigen-antibody binding conformation.
[0008] In a third aspect, some embodiments of the present disclosure provide an electronic device, which comprises: one or more processors; a storage device having one or more programs stored thereon, when the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any of the implementation manners of the first aspect.
[0009] In a fourth aspect, some embodiments of the present disclosure provide a computer readable medium having stored thereon a computer program, wherein the program, when executed by a processor, implements the method described in any implementation manner of the first aspect.
[0010] The above various embodiments of the present disclosure have the following beneficial effects: through the antigen-antibody binding conformation optimization method based on iterative closest point of some embodiments of the present disclosure, the accuracy of the determined antibody conformation can be improved. Specifically, the antigen-antibody binding conformation optimization method based on iterative closest point of some embodiments of the present disclosure, first, the antigen structure file and the antibody structure file are subjected to coarse-grained processing to obtain a coarse-grained antibody file and a coarse-grained antigen file. Here, through the coarse-grained processing, the antigen and antibody structure can be simplified (the CA (i.e. alpha-carbon atom) atom of the side chain center simulation is retained), while reducing the computational complexity, the core spatial information related to the antigen-antibody binding is accurately retained, and an efficient calculation basis is provided for the subsequent matching of the CDR region and the epitope. Then, based on the above coarse-grained antibody file and the above coarse-grained antigen file, the above antigen structure file and the above antibody structure file are subjected to surface residue extraction to obtain an antigen surface residue information set and an antibody surface residue information set. Here, through the extraction of the surface residues, the irrelevant region interference can be reduced, and the subsequent docking pertinence can be improved. Then, according to the antibody surface residue information set, an antibody binding region residue list is generated. Here, through the generation of the antibody binding region residue list, the antibody binding core region can be accurately positioned. Then, based on the antibody binding region residue list, an epitope model is generated. Here, through the generation of the epitope model, the spatial characteristics of the original epitope can be simulated, the defect of weak complementarity between the CDR region and the antigen is made up, a more tolerant matching space is provided for the flexible CDR region, and the fault tolerance is improved. Next, the antigen surface residue information set and the epitope model are subjected to spatial matching to obtain a matched epitope model. Here, through the spatial matching, the subsequent search range can be narrowed, the inefficiency of random sampling in the whole space in the conventional method can be avoided, and the docking efficiency can be improved. Then, according to the matched epitope model, an antibody docking conformation set is generated. Here, through the generation of the docking conformation, the coarse-grained antibody and the epitope conformation can be aligned, the antibody can be preliminarily aligned to the antigen binding site, and the problem of low matching degree between the overall posture of the antibody and the antigen in the conventional method can be solved. In addition, the coarse-grained antigen file is used to locally optimize each antibody docking conformation in the antibody docking conformation set to obtain an optimized antibody conformation set. Here, through the local optimization, the flexible characteristics of the CDR region can be adapted, the problem of low fault tolerance caused by rigid complementarity can be avoided, and the geometric matching accuracy of the binding interface can be significantly improved. Finally, according to the coarse-grained antigen file, each optimized antibody conformation in the optimized antibody conformation set is screened to obtain a target antigen-antibody binding conformation. Here, through the final screening, the hit rate of the near-native structure can be improved, and the defect that the conventional scoring function ignores the antigen-antibody specific interaction can be solved. BRIEF DESCRIPTION OF DRAWINGS
[0011] The above and other features, aspects and advantages of the present disclosure will become more apparent with reference to the following detailed description when taken in conjunction with the accompanying drawings. Throughout the drawings, similar or same reference numerals are used to denote similar or same elements. It is to be understood that the drawings are schematic, and elements and features are not necessarily to scale.
[0012] Figure 1 is a flowchart of some embodiments of the antigen-antibody binding conformation optimization method based on iterative closest point according to the present disclosure; Figure 2 is a schematic diagram of the positional relationship of the amino acid chain, CDR region and epitope model under three-dimensional representation; Figure 3 is a structural schematic diagram of some embodiments of the antigen-antibody binding conformation optimization device based on iterative closest point according to the present disclosure; Figure 4 is a structural schematic diagram of an electronic device suitable for implementing some embodiments of the present disclosure. DETAILED DESCRIPTION
[0013] Embodiments of the present disclosure will be described in more detail with reference to the drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms, and should not be interpreted as being limited to the embodiments set forth herein. Rather, these embodiments are provided so that the present disclosure can be more thoroughly and completely understood. It should be understood that the drawings and embodiments of the present disclosure are only for illustrative purposes, and are not intended to limit the scope of protection of the present disclosure.
[0014] It should also be noted that, for ease of description, only the parts related to the present application are shown in the drawings. The embodiments in the present disclosure and the features in the embodiments can be combined with each other without conflict.
[0015] It should be noted that the concepts of "first", "second", etc. mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not intended to limit the order or interdependence of the functions performed by these devices, modules or units.
[0016] It should be noted that the adjectives "one", "multiple" mentioned in the present disclosure are illustrative and not limiting, and those skilled in the art should understand that, unless otherwise explicitly stated in the context, it should be understood as "one or more".
[0017] The names of the messages or information exchanged between the devices in the embodiments of the present disclosure are only for illustrative purposes, and are not intended to limit the scope of the messages or information. In practice, the FFT algorithm significantly improves the speed of exhaustive sampling of the search space in the current mainstream classical algorithm for antibody conformation. However, the FFT-based method has major obstacles: first, it needs a correlation function to express energy, and second, it cannot infer pose quality from the value of the energy function. Despite these limitations, this type of method has been widely recognized. In addition to exhaustive methods, docking can also be achieved by exploring the geometric features of protein surfaces, for example, tools such as PatchDock and SP-Dock align proteins based on complementary surface patches; histogram-based descriptors are also popular in finding complementary regions, such as the Zernike descriptor (ZD) used by LZerD, which can represent the shape of the protein surface in the term of the 3D function series expansion.
[0018] The ICP (Iterative Closest Point) algorithm is a classic algorithm for 3D shape matching. Its basic principle is to continuously calculate and optimize the rotation and translation transformation between the data point cloud and the model point cloud to minimize the distance between the two, thereby achieving point cloud alignment (registration). In recent years, the ICP algorithm has been widely used not only in the field of computer vision but also in biology and chemistry, such as for protein structure comparison.
[0019] However, the currently disclosed protein docking software is not as effective as conventional protein docking in antigen-antibody prediction. This is mainly because antibodies mainly bind to antigen epitopes through CDR (Complementarity-Determining Region) regions, and the complementarity of CDR regions with antigens is weaker than that of conventional protein-protein interactions, and CDR regions (especially H3 (Heavy Chain Complementarity-Determining Region 3) heavy chain complementarity determining region) have strong flexibility, and excessive reliance on rigid complementarity will result in low fault tolerance. For example, SP-Dock protein docking software (SP-Dock: Protein-Protein Docking using Shape and Physicochemical Complementarity) uses ICP algorithm to achieve shape complementarity by matching adjacent surface patches of receptors and ligands at the same time. The specific process is as follows: first, the MSMS (Maximal Speed Molecular Surface) algorithm is used to generate the solvent excluded surface (Solvent Excluded Surface, SES), and the key points are extracted based on local curvature to generate geodesic surface patches (Geodesic Surface Patch, GSP), and then the ICP algorithm is used for protein surface alignment to optimize the rigid transformation of ligand and receptor patch groups. However, like other protein docking software, due to excessive emphasis on the complementarity of receptors and ligands, and the number and quality of residues contained in the surface patch cannot specifically represent the antibody binding site and antigen epitope, ultimately resulting in poor support effect on antigen-antibody docking. Here, the existing antibody AI generation model has obvious defects: it generates antibodies by machine learning antigen-antibody binding rules, but the success rate is extremely low when the same strategy is used for antigen-antibody docking, indicating that machine learning has not truly mastered the underlying rules, resulting in the success rate and credibility of the generated antibodies cannot meet the actual demand, and a special antigen-antibody docking software is needed to guide the antibody AI model generation. According to this technical problem, the present disclosure solves it in the following way. The present disclosure will be described in detail below with reference to the accompanying drawings and embodiments.
[0020] Figure 1 The flow 100 of some embodiments of the antigen-antibody binding conformation optimization method based on the iterative closest point according to the present disclosure is shown. The antigen-antibody binding conformation optimization method based on the iterative closest point includes the following steps: Step 101, performing coarse-grained processing on the antigen structure file and the antibody structure file to obtain a coarse-grained antibody file and a coarse-grained antigen file.
[0021] In some embodiments, the execution subject (e.g., a computing device) of the iterative closest point-based antigen-antibody binding conformation optimization method can perform coarse-graining processing on the antigen structure file and the antibody structure file to obtain a coarse-grained antibody file and a coarse-grained antigen file. The antigen structure file includes antigen structure information of a free antigen. The antibody structure file includes antibody structure information of a free antibody.
[0022] It should be noted that the computing device described above can be hardware or software. When the computing device is hardware, it can be implemented as a distributed cluster composed of multiple servers or terminal devices, or as a single server or a single terminal device. When the computing device is software, it can be installed in the hardware devices listed above. It can be implemented as multiple software or software modules for providing distributed services, or as a single software or software module. No specific limitation is made herein.
[0023] The coarse-graining processing can be simplifying the data structure of the antigen structure file and the antibody structure file. For example, the antigen and antibody PDB (Protein Data Bank File, a standardized file format for storing three-dimensional spatial structure information of biological macromolecules such as proteins, nucleic acids, and sugars) structure files are respectively coarse-grained, and only the CA atom of the side chain center simulation is retained. The coordinates of the side chain atoms of the residues, i.e., the atoms other than the main chain 'N': nitrogen atom, 'CA': alpha-carbon atom, 'C': carbon atom, and 'O': oxygen atom, are extracted, the center of mass of the side chain atoms is calculated, and the center of mass is represented by a virtual CA atom. Glycine (GLY) has no side chain atoms, and the CA coordinates of the main chain are used instead.
[0024] Specifically, the coarse-graining processing of the execution subject on the antigen structure file and the antibody structure file to obtain the coarse-grained antibody file and the coarse-grained antigen file includes the following steps: Step S1, extract the antigen side chain atom coordinates and the antibody side chain atom coordinates in the antigen structure file and the antibody structure file to obtain an antigen side chain atom coordinate sequence and an antibody side chain atom coordinate sequence. The antigen structure file (i.e., a PDB file) includes various atomic coordinates, amino acid sequences, and other information. Thus, the antigen side chain atom coordinates and the antibody side chain atom coordinates can be extracted by indexing.
[0025] Step S2, determine the centroid coordinates corresponding to the sequence of side chain atomic coordinates of the antigen, obtain the antigen centroid coordinate set, and combine each antigen centroid coordinate in the antigen centroid coordinate set into a coarse-grained antigen file according to the order of the antigen structure in the antigen structure file. Wherein, first obtain the atomic mass corresponding to each antigen side chain atomic coordinate. Then, the average isotopic mass of the atom can be substituted into the centroid calculation formula to obtain the antigen centroid coordinates. Here, the centroid is calculated by mass-weighted average of side chain atoms, rather than simple geometric average, because it can more truly reflect the spatial distribution center of the residue side chain.
[0026] Step S3, determine the centroid coordinates corresponding to the sequence of side chain atomic coordinates of the antibody, obtain the antibody centroid coordinate set, and combine each antibody centroid coordinate in the antibody centroid coordinate set into a coarse-grained antibody file according to the order of the antibody structure in the antibody structure file. Wherein, the generation method of antibody centroid coordinates and the achieved technical effects can refer to the above step S2, which will not be described in detail here. Here, the centroid coordinates can be three-dimensional coordinates.
[0027] Step 102, based on the coarse-grained antibody file and the coarse-grained antigen file, surface residue extraction is performed on the antigen structure file and the antibody structure file to obtain the antigen surface residue information set and the antibody surface residue information set.
[0028] In some embodiments, the execution subject can perform surface residue extraction on the antigen structure file and the antibody structure file based on the coarse-grained antibody file and the coarse-grained antigen file to obtain the (coarse-grained) antigen surface residue information set and the (coarse-grained) antibody surface residue information set.
[0029] In some optional implementations of some embodiments, the execution subject performs surface residue extraction on the antigen structure file and the antibody structure file based on the coarse-grained antibody file and the coarse-grained antigen file to obtain the antigen surface residue information set and the antibody surface residue information set, including: Step S1, determine the solvent accessibility surface area of each amino acid residue in the antibody structure file to obtain the sequence of antibody solvent accessibility surface areas. Wherein, the antibody structure file can be input into the FreeSASA tool to calculate the solvent accessibility surface area of each amino acid residue, SASA (Solvent accessible surface area, solvent accessibility surface area).
[0030] Step S2, the antibody solvent solubility surface area and the corresponding residue chain information in the antibody solvent solubility surface area sequence that meet the preset residue screening condition are determined as the antibody surface residue information, and are applied to the coarse-grained antibody file to obtain the antibody surface residue information set. The residue chain information can include residue chain identification, residue sequence number, amino acid type, solvent accessibility surface area, and surface residue atomic coordinates. Secondly, the preset residue screening condition can be that the antibody solvent solubility surface area or the antigen solvent solubility surface area is within a preset area interval (for example, [10-30] square angstroms). The application of the residue chain information to the coarse-grained antibody file can be to delete the non-exposed residues in the coarse-grained antibody file except for the residue chain information, so that the remaining residue chain information in the coarse-grained antibody file is consistent with the screened residue chain information.
[0031] Step S3, the solvent accessibility surface area of each amino acid residue in the antigen structure file is determined to obtain an antigen solvent solubility surface area sequence.
[0032] Step S4, the antigen solvent solubility surface area and the corresponding residue chain information in the antigen solvent solubility surface area sequence that meet the preset residue screening condition are determined as the antigen surface residue information, and are applied to the coarse-grained antigen file to obtain the antigen surface residue information set.
[0033] Specifically, each residue (amino acid) in the PDB file of the antigen and the antibody has an original number in the structure record (defined by the "resSeq" field in the ATOM line of the PDB file). For example, the residues in the PDB file of a certain antibody light chain can be numbered in the order of "1, 2, 3...". The antigen residues can be numbered as "101, 102... " or independently (the residue numbers of different chains can be repeated, and need to be distinguished in combination with chain ID, such as chain L, chain H, chain A, etc.). Here, the residue sequence number directly follows the original number of the PDB file, and is associated with the chain ID (such as L chain residues 1-100, H chain residues 1-120, and antigen A chain residues 1-200), which is used to mark the exposed residues. Here, the application of the residue chain information to the coarse-grained antigen file can be to delete the non-exposed residues in the coarse-grained antigen file except for the residue chain information, so that the remaining residue chain information in the coarse-grained antigen file is consistent with the screened residue chain information.
[0034] Step 103, generating an antibody binding region residue list according to the antibody surface residue information set.
[0035] In some embodiments, the execution subject can generate an antibody binding region residue list according to the antibody surface residue information set.
[0036] In some optional implementations of some embodiments, the antibody surface residue information in the antibody surface residue information set comprises an amino acid sequence. The execution subject generates a list of antibody binding region residues according to the antibody surface residue information set, comprising: In step S1, the amino acid sequence included in the antibody surface residue information in the antibody surface residue information set is standardized numbered to obtain an amino acid numbering information sequence, and a complementarity determining region residue numbering sequence group is marked in the amino acid numbering information sequence. The amino acid numbering information comprises a standardized heavy chain numbering sequence, a first numbering mapping relationship sequence, a standardized light chain numbering sequence, and a second numbering mapping relationship sequence. The amino acid sequence in each antibody surface residue information can be standardized numbered by an antibody CDR region division method (for example, KABAT rule) to obtain the amino acid numbering information sequence. Specifically, the amino acid sequence can be numbered by a Kabat numbering tool (such as Abnum numbering software, IgBLAST numbering tool, etc.), and the CDR1, CDR2, and CDR3 regions of the heavy chain and light chain are marked according to the KABAT rule (i.e., the residue sequence number included in each CDR region is determined). Thus, the first numbering mapping relationship and the second numbering mapping relationship can be a mapping relationship between the Kabat numbering and the PDB original numbering represented by the standardized heavy chain numbering. Here, at least one complementarity determining region and the residue numbering in the complementarity determining region can be marked in the amino acid numbering information sequence by the KABAT rule to obtain the complementarity determining region residue numbering sequence.
[0037] In step S2, the solvent solubility surface area corresponding to each complementarity determining region residue numbering in the complementarity determining region residue numbering sequence group is determined to obtain a residue surface area sequence group. The solvent solubility surface area corresponding to the complementarity determining region residue numbering can be determined by a FreeSASA tool.
[0038] In practice, the SASA threshold (10-30) used when generating the extracted surface residues in the surface residue extraction step is a general standard suitable for the preliminary screening of global residues. However, the "exposure" of the residues in the CDR region, as the functional core of the antibody, has special significance: 1. The SASA value of part of the CDR region residues may be slightly lower than the threshold (such as 8-10 square angstroms), but because of the key position (such as the edge of the binding site) in the CDR region, it may still be involved in antigen binding, and the accurate SASA value needs to be calculated separately to confirm whether it is exposed; 2. Conversely, although part of the CDR region residues are included in the "surface residues", it is necessary to verify whether the SASA value truly reflects the exposure degree by separate calculation (to avoid misjudgment due to errors in global calculation).
[0039] Step S3, sort each CDR residue number in each CDR residue number sequence in the above residue surface area sequence set to obtain a sorted residue number sequence set. In this process, the residues can be sorted in descending order of residue surface area.
[0040] Step S4, determine the structure curvature corresponding to each sorted residue number in the above sorted residue number sequence set to obtain a structure curvature sequence set. The calculation of structure curvature is based on the three-dimensional structure of the antibody variable region, and the core is to quantify the local structure bending degree of the region through the atomic coordinates / amino acid coordinates around the residue. Specifically, the structure curvature of the sorted residue number can be calculated by the following steps: 1. For each residue in the CDR region, select all atoms (including main chain and side chain atoms) within a certain range (e.g. 5-10 angstroms) around it to obtain the corresponding three-dimensional coordinates (x, y, z). 2. Use the least squares method to fit the extracted atomic coordinates to an ideal surface (usually a spherical or planar surface). If it is fitted to a spherical surface, calculate the radius (curvature radius R) of the spherical surface, then the curvature K=1 / R (the smaller the radius, the greater the curvature, indicating that the local structure is more "convex" or "concave"). If it is fitted to a plane, the curvature is 0 (indicating that the local structure is flat and not easy to participate in specific binding). Finally, normalize the calculated curvature value (e.g. scale to 0-1 according to the curvature range of all residues in the CDR region) to facilitate comparison between different residues and highlight high-curvature residues. In this way, a structure curvature sequence set is obtained. Each sorted residue number can correspond to a structure curvature.
[0041] Step S5, according to the above structure curvature sequence set, screen out a target number (e.g. 1-5) of CDR residue numbers from each CDR residue number sequence in the above CDR residue number sequence set, and combine the target number of CDR residue numbers into an antibody binding region residue list. In practice, the residue surface area reflects the degree of exposure of the residue to the solvent, and a residue surface area > 20 square angstroms indicates that the side chain of the residue is mostly exposed, with the possibility of spatial binding with the antigen (if the residue is deeply buried in the protein, even if the curvature is high, it cannot contact the antigen). Secondly, the side chain structure of the residue with high curvature (such as located at the top of the CDR loop or the edge of the "pocket") is more likely to form specific interactions (such as hydrogen bonds, van der Waals forces) with the complementary structure (such as concave, convex) of the antigen epitope. Therefore, the residue number screening condition can be: preferentially screening the target number of CDR residue numbers with a residue surface area greater than 20 square angstroms and / or a structure curvature ranking at the front.
[0042] Step 104, generating an epitope model based on the antibody binding region residue list.
[0043] In some embodiments, the execution subject can generate an epitope model based on the above-mentioned antibody binding region residue list.
[0044] As an example, the epitope model can be generated by the following steps: 1. Extract relevant information from the antibody surface residue PDB file and the CDR file, calculate the coordinate range of the CDR region residues, and then generate uniformly distributed coordinate points within the range.
[0045] 2. Translate the coordinate points along the antibody direction, with the distance between the coordinate points and the CDR residues as the random stop condition, and delete the conflicting points if the distance between the coordinate points and the non-CDR residues is less than the distance between the coordinate points and the CDR residues.
[0046] As shown in Figure 2 , blue is the amino acid chain, yellow is the CDR region, white is the coordinate point generated according to the CDR coordinates, and green is the generated epitope model.
[0047] In some optional implementations of some embodiments, the execution subject generates an epitope model based on the above-mentioned antibody binding region residue list, including: Step S1, using the above-mentioned antibody binding region residue list as a template, determine the corresponding residue space distribution range. Wherein, the surface can be discretized into grid points (each grid point represents a point on the surface, recording three-dimensional coordinates and normal vector) using surface generation tools (such as show surface of PyMOL molecular visualization system, MSMS software), that is, a spatial distribution grid model. Here, each square angstrom can contain 1-2 grid points, thereby determining the residue space distribution range of each antibody binding region residue. For example, the minimum circumscribed sphere radius of the residue space, the center coordinates.
[0048] Step S2, according to the above-mentioned residue space distribution range and the preset number of candidate points (for example, 20-150), set the candidate point distribution range.
[0049] As an example, the candidate point distribution range can be a spherical region with a radius of 5-10 angstroms around the center of the residue space distribution range, obtaining the candidate point distribution range.
[0050] Step S3, generate an initial candidate point coordinate sequence within the above-mentioned candidate point distribution range. Here, the initial candidate point coordinate sequence can be generated within the above-mentioned candidate point distribution range by a random sampling algorithm (for example, uniform distribution sampling, Gaussian distribution sampling, etc.). Wherein, the distance between adjacent initial candidate point coordinates in the initial candidate point coordinate sequence is greater than or equal to the preset distance threshold (for example, 3 angstroms).
[0051] Step S4, based on the above-mentioned residue space distribution range, each initial candidate point coordinate in the above-mentioned initial candidate point coordinate sequence is screened to obtain an epitope model. Wherein, the screening can be carried out by the following steps: 1. Calculate the relative position of the initial candidate point to the CDR region surface. Wherein, for the generated initial candidate point coordinates (20-150 random distribution points), the shortest distance of each point to the surface grid point of the CDR region is calculated, and the normal vector direction relationship between the point and the surface grid point is calculated.
[0052] 2. Match the convex and concave features and screen. Wherein, if the shortest distance of the candidate point to the convex region of the CDR region is 3-5 angstroms, and located in the opposite direction of the normal vector of the region (i.e. pointing to the inside of the convex), it meets the "convex-concave" complementarity and is retained. If the shortest distance of the candidate point to the concave region of the CDR region is 3-5 angstroms, and located in the direction of the normal vector of the region (i.e. pointing to the outside of the concave), it meets the "concave-convex" complementarity and is retained. If the candidate point is too close to the surface (such as distance <2 angstroms, which may collide), too far away (distance >6 angstroms, which cannot be combined), or does not match the convex and concave features (such as convex corresponding to convex), it is removed.
[0053] 3. Control the number of candidate points. If the number of retained candidate points after screening is less than 20, the distance constraint can be relaxed (such as 2-7 angstroms) or a batch of initial candidate points can be generated for re-screening. If it exceeds 150, sort according to the "complementary matching degree to the surface of the CDR region" (such as the closer the distance is to 3-5 angstroms, the more consistent the normal vector direction is, and the higher the matching degree is), and retain the first 150 to obtain the epitope model. That is, the epitope model is composed of a plurality of epitope candidate points.
[0054] Step 105, spatially matching the antigen surface residue information set and the epitope model to obtain a matched epitope model.
[0055] In some embodiments, the above-mentioned execution subject can spatially match the above-mentioned antigen surface residue information set and the above-mentioned epitope model to obtain a matched epitope model. Wherein, the antigen surface residue information set is taken as the alignment target, and the epitope model is spatially matched to obtain a matched epitope model.
[0056] In some optional implementations of some embodiments, the above-mentioned execution subject spatially matches the above-mentioned antigen surface residue information set and the above-mentioned epitope model to obtain a matched epitope model, comprising: Step S1, determining the overall epitope candidate point centroid coordinates of each epitope candidate point in the above-mentioned epitope model. Wherein, the overall epitope candidate point centroid coordinates of each epitope candidate point in the above-mentioned epitope model can be determined by the centroid formula.
[0057] Step S2, align the whole epitope candidate point centroid coordinates and each antigen surface residue information in the antigen surface residue information set to obtain a matched epitope model. The alignment can be performed by the following steps: First, the centroid coordinates of the coordinates of each antigen surface residue in the antigen surface residue information set are determined and moved to 0 point as the residue centroid coordinates. Then, the whole epitope candidate point centroid coordinates are translated to the residue centroid coordinates (i.e. 0 point), and then translated by a distance d (0 < d < 20 angstroms) in the negative direction of the horizontal axis to make the whole epitope candidate point centroid coordinates located at the position of (-d, 0, 0). Here, the distance d can be adjusted according to the size of the antigen, and the small molecule antigen takes 5-10 angstroms, and the protein antigen takes 10-20 angstroms (to ensure that the initial distance is reasonable, to avoid too close collision or too far interaction), to obtain the matched epitope model.
[0058] (− d ,0,0) Step 106, generating an antibody docking conformation set according to the matched epitope model.
[0059] In some embodiments, the execution subject can generate an antibody docking conformation set according to the matched epitope model.
[0060] In some optional implementations of some embodiments, the execution subject generates an antibody docking conformation set according to the matched epitope model, including: Step S1, performing whole rotation and translation operations on the matched epitope model according to a preset rotation angle group to obtain a rotation conformation set. The ICP (Iterative Closest Point) iteration can be performed by the following steps.
[0061] First, each matched epitope candidate point in the matched epitope model is sequentially rotated around the X-Y-Z axis with the residue centroid coordinates as the rotation center. At the same time, the whole epitope candidate point centroid coordinates of the epitope model are located and rotated around the YZ axis to compensate for the rotation. Thus, with a fixed direction to the antigen, the angle is set to 2-15°, and thus a current epitope model is obtained as a rotation conformation for each rotation of a preset rotation angle, to obtain the rotation conformation set.
[0062] Here, the rotation compensation is to ensure that the binding surface of the matched epitope candidate point is always facing the antigen when it rotates around the residue centroid coordinate. Specifically, the rotation process is divided into two steps, first rotating around the residue centroid coordinate (revolution), and then rotating around the centroid of the epitope model (rotation). When revolving around the Y axis by θ angle, the Y axis is simultaneously rotated by θ angle; when revolving around the Z axis by φ angle, the Z axis is simultaneously rotated by φ angle. Rotation around the centroid (rotation compensation). Through this compensation, the binding surface of the epitope candidate point always maintains the relative orientation with the antigen surface during the global rotation, avoiding the deviation of the binding surface caused by rotation.
[0063] Step S2, using the antigen surface residue information set as the alignment target, performing iterative closest point space matching on the rotation conformation set to obtain an epitope model set after iteration.
[0064] The rotation conformation set is iterated by ICP (Iterative Closest Point) through the following steps: the above rotation conformation set is translated with the connecting line of the overall epitope candidate point centroid coordinate and the residue centroid coordinate of the rotation conformation set as the axis, and the translation direction is towards the antigen surface, i.e. away from the residue centroid coordinate, i.e. the antigen surface is taken as the target point set. The ICP rotation angle is limited to 2-60° on the XYZ axis, and finally the epitope model set after iteration is obtained through translation and rotation.
[0065] Step S3, from the above epitope model set after iteration, sorting and screening out the epitope model after iteration, and taking each epitope candidate point in the selected epitope model after iteration as an epitope positioning point group to obtain an epitope positioning point group set.
[0066] The epitope model is sorted and screened from the above epitope candidate point group epitope model set to obtain an epitope positioning point group set.
[0067] Optionally, when sorting and screening, the RMSD (Root Mean Square Deviation) of the multiple point pairs with the shortest distance between the epitope candidate points in the epitope model after iteration and the antigen surface residue information set is sorted, thereby the preset number of epitope models after iteration with the smallest RMSD between the antigen surface residue information set can be screened, and the epitope candidate points included therein are taken as the epitope positioning point group to obtain an epitope positioning point group set. For example, 2000-5000 epitope positioning point groups can be obtained to form an epitope positioning point group set.
[0068] Specifically, the sorting screening is: for each iteration epitope model, determining the n (for example, 5-15) closest point pairs between the iteration epitope model and the antigen surface residue information set to obtain a point pair group. Then, the RMSD (Root Mean Square Deviation) value of each point pair group is determined. Then, the iteration epitope models are sorted according to the RMSD values to obtain a sorted epitope model sequence. Finally, the M (for example, 2000-5000) selected sorted epitope models with the smallest RMSD values in the sorted epitope model sequence are selected, and the included epitope candidate points are determined as the epitope positioning point group to obtain the epitope positioning point group set.
[0069] Step S4, aligning the above epitope model with each epitope positioning point group in the epitope positioning point group set to obtain an aligned rotation matrix and an aligned translation vector, and applying each aligned rotation matrix and each aligned translation vector to the above coarse-grained antibody file to generate an antibody docking conformation to obtain an antibody docking conformation set. Wherein, the above epitope model can be aligned with each epitope positioning point group in the epitope positioning point group set to obtain an aligned rotation matrix and an aligned translation vector according to the alignment processing step in the above step 104. Secondly, the antibody coordinates in the coarse-grained antibody file can be adjusted according to each aligned rotation matrix and each aligned translation vector to obtain an antibody docking conformation set.
[0070] Step 107, using the coarse-grained antigen file to locally optimize each antibody docking conformation in the antibody docking conformation set to obtain an optimized antibody conformation set.
[0071] In some embodiments, the above execution subject can use the above coarse-grained antigen file to locally optimize each docking conformation in the above antibody docking conformation set to obtain an optimized antibody conformation set.
[0072] In some optional implementations of some embodiments, the above execution subject uses the above coarse-grained antigen file to locally optimize each docking conformation in the above antibody docking conformation set to obtain an optimized antibody conformation set, including: Step S1, selecting a plurality of matching nearest point coordinates from the above coarse-grained antigen file and each epitope positioning point group in the above epitope positioning point group set as an antigen nearest point coordinate sequence to obtain an antigen nearest point coordinate sequence set. Wherein, the matching can be corresponding to the same antigen identifier. Here, the coarse-grained antigen file used can also be a coarse-grained antigen file after deleting non-exposed residues.
[0073] Step S2, according to the above antigen closest point coordinate sequence set and the above antibody binding region residue list, the closest point optimization is carried out on each antibody docking conformation in the above antibody docking conformation set, and the optimized antibody conformation set is obtained. Wherein, the closest point optimization can be carried out on each antibody docking conformation in the above antibody docking conformation set by the above ICP algorithm, and the optimized antibody conformation set is obtained. Here, each antibody docking conformation in the iteration process rotates around its own centroid as the center of rotation, and the rotation angle of the horizontal axis, the vertical axis and the vertical axis of the rotation is limited to 2-60°.
[0074] Step 108, according to the coarse-grained antigen file, each optimized antibody conformation in the optimized antibody conformation set is screened to obtain the target antigen antibody binding conformation.
[0075] In some embodiments, the execution subject can screen each optimized antibody conformation in the optimized antibody conformation set according to the coarse-grained antigen file to obtain the target antigen antibody binding conformation.
[0076] In some optional implementations of some embodiments, the execution subject screens each optimized antibody conformation in the optimized antibody conformation set according to the coarse-grained antigen file to obtain the target antigen antibody binding conformation, comprising: Step S1, according to the coarse-grained antigen file, the interaction potential of each optimized antibody conformation in the optimized antibody conformation set is determined to obtain the interaction potential set. Wherein, the interaction potential of each optimized antibody conformation in the optimized antibody conformation set can be determined by MJ (Miyazawa-Jernigan) potential calculation function to obtain the interaction potential set. Here, Miyazawa-Jernigan potential reflects the preference of different amino acid combinations in the interior of protein based on the statistical interaction frequency between amino acid pairs (such as hydrophobic residues tend to aggregate).
[0077] Step S2, according to the above interaction potential set, the optimized antibody conformation meeting the conformation screening condition is screened from the above optimized antibody conformation set as the target antigen antibody binding conformation. Wherein, the conformation screening condition can be the optimized antibody conformation with the minimum interaction potential.
[0078] Optionally, the screening can also be carried out by the following steps: 1, extract the contact area: for the optimized antibody conformation and the residue pair with a distance less than 5 angstroms in the antigen surface residue information set.
[0079] 2, for each residue pair, the corresponding interaction potential in the interaction potential set is accumulated to obtain the accumulated potential value.
[0080] 3. According to the accumulated potential energy value, the optimized antibody conformation set is sorted from small to large, and the optimized antibody conformation with the minimum interaction potential energy is selected as the target antigen-antibody binding conformation.
[0081] In practice, the present application is optimized for antigen-antibody docking: a space point cloud (i.e. epitope model) is simulated for the CDR region of the antibody, which can completely cover the actual antigen epitope distribution in quantity and distribution; then the ICP algorithm is used to find multiple nearest points for the epitope model and the antigen surface residues, and the ICP optimization is performed again through the determined nearest points and the CDR region of the antibody to adjust the best pose of the antibody. This optimization scheme can significantly improve the hit rate and fault tolerance for antibody binding sites with flexibility. At the same time, because of the coarse-graining of the side chains of the antigen residues and the side chains of the antibody residues, the calculation amount is optimized, enough geometric structure features are retained, and more sufficient elastic adjustment space is provided for the flexible region.
[0082] The above various embodiments of the present disclosure have the following beneficial effects: through the antigen-antibody binding conformation optimization method based on iterative closest point of some embodiments of the present disclosure, the accuracy of the determined antibody conformation can be improved. Specifically, the antigen-antibody binding conformation optimization method based on iterative closest point of some embodiments of the present disclosure, first, the antigen structure file and the antibody structure file are subjected to coarse-grained processing to obtain a coarse-grained antibody file and a coarse-grained antigen file. Here, through the coarse-grained processing, the antigen and antibody structure can be simplified (the CA (i.e. alpha-carbon atom) atom of the side chain center simulation is retained), while reducing the computational complexity, the core spatial information related to antigen-antibody binding is accurately retained, and an efficient calculation basis is provided for the subsequent matching of the CDR region and the epitope. Then, based on the above coarse-grained antibody file and the above coarse-grained antigen file, the above antigen structure file and the above antibody structure file are subjected to surface residue extraction to obtain an antigen surface residue information set and an antibody surface residue information set. Here, by extracting the surface residues, the irrelevant region interference can be reduced, and the subsequent docking specificity can be improved. Then, according to the antibody surface residue information set, an antibody binding region residue list is generated. Here, by generating the antibody binding region residue list, the antibody binding core region can be accurately positioned. Then, based on the antibody binding region residue list, an epitope model is generated. Here, by generating the epitope model, the spatial characteristics of the original epitope can be simulated, the defect of weak complementarity between the CDR region and the antigen can be made up, a more tolerant matching space can be provided for the flexible CDR region, and the fault tolerance rate can be improved. Next, the antigen surface residue information set and the epitope model are subjected to spatial matching to obtain a matched epitope model. Here, through spatial matching, the subsequent search range can be narrowed, the inefficiency of random sampling in the whole space in the conventional method can be avoided, and the docking efficiency can be improved. Then, according to the matched epitope model, an antibody docking conformation set is generated. Here, by generating the docking conformation, the coarse-grained antibody and the epitope conformation can be aligned, the antibody can be preliminarily aligned with the antigen binding site, and the problem of low matching degree between the overall posture of the antibody and the antigen in the conventional method can be solved. In addition, the coarse-grained antigen file is used to locally optimize each antibody docking conformation in the antibody docking conformation set to obtain an optimized antibody conformation set. Here, through local optimization, the flexible characteristics of the CDR region can be adapted, the problem of low fault tolerance rate caused by rigid complementarity can be avoided, and the geometric matching accuracy of the binding interface can be significantly improved. Finally, according to the coarse-grained antigen file, each optimized antibody conformation in the optimized antibody conformation set is screened to obtain a target antigen-antibody binding conformation. Here, through the final screening, the hit rate of the near-native structure can be improved, and the defect that the conventional scoring function ignores the antigen-antibody specific interaction can be solved. Further reference is made to Figure 3As an implementation of the methods shown in the above figures, the present disclosure provides some embodiments of an iterative closest point based antigen-antibody binding conformation optimization apparatus, which implements the method embodiments shown in Figure 1 The iterative closest point based antigen-antibody binding conformation optimization apparatus can be specifically applied in various electronic devices.
[0083] As shown in Figure 3 The iterative closest point based antigen-antibody binding conformation optimization apparatus 300 of some embodiments includes a coarse-grained processing unit 301, an extraction unit 302, a first generation unit 303, a second generation unit 304, a spatial matching unit 305, a third generation unit 306, a local optimization unit 307, and a screening unit 308. The coarse-grained processing unit 301 is configured to perform coarse-grained processing on an antigen structure file and an antibody structure file to obtain a coarse-grained antibody file and a coarse-grained antigen file. The extraction unit 302 is configured to perform surface residue extraction on the antigen structure file and the antibody structure file based on the coarse-grained antibody file and the coarse-grained antigen file to obtain an antigen surface residue information set and an antibody surface residue information set. The first generation unit 303 is configured to generate an antibody binding region residue list according to the antibody surface residue information set. The second generation unit 304 is configured to generate an epitope model based on the antibody binding region residue list. The spatial matching unit 305 is configured to perform spatial matching on the antigen surface residue information set and the epitope model to obtain a matched epitope model. The third generation unit 306 is configured to generate a set of antibody docking conformations according to the matched epitope model. The local optimization unit 307 is configured to perform local optimization on each antibody docking conformation in the set of antibody docking conformations using the coarse-grained antigen file to obtain a set of optimized antibody conformations. The screening unit 308 is configured to perform screening on each optimized antibody conformation in the set of optimized antibody conformations according to the coarse-grained antigen file to obtain a target antigen-antibody binding conformation.
[0084] It can be understood that the units described in the iterative closest point based antigen-antibody binding conformation optimization apparatus 300 correspond to the respective steps in the method described with reference to Figure 1 Thus, the operations, features, and advantages described above for the method also apply to the iterative closest point based antigen-antibody binding conformation optimization apparatus 300 and the units contained therein, which will not be described here again. Reference is made to Figure 4 which shows a structural schematic diagram of an electronic device (such as a computing device) suitable for implementing some embodiments of the present disclosure. Figure 4 The electronic device shown is merely an example and should not impose any limitation on the functions and use range of the embodiments of the present disclosure. AsFigure 4 As shown in the figure, the computer device includes a processor, a memory and a network interface connected through a system bus, wherein the memory can include a non-volatile storage medium and an internal memory. The non-volatile storage medium can store an operating system and a computer program. The computer program includes program instructions which, when executed, can cause the processor to perform any one of the above methods. The processor is used to provide computing and control capabilities to support the operation of the entire computer device. The internal memory provides an environment for the computer program in the non-volatile storage medium to run, which, when executed by the processor, can cause the processor to perform any one of the above methods. The network interface is used for network communication, such as sending assigned tasks, etc. Those skilled in the art can understand that, Figure 4 The structure shown in the figure is only a block diagram of part of the structure related to the present disclosure, and does not constitute a limitation on the computer device to which the present disclosure is applied. The specific computer device can include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0085] It should be understood that the processor can be a central processing unit (CPU), and the processor can also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc.
[0086] In one embodiment, the processor is configured to run a computer program stored in the memory to perform the following steps: performing coarse-grained processing on the antigen structure file and the antibody structure file to obtain a coarse-grained antibody file and a coarse-grained antigen file; performing surface residue extraction on the antigen structure file and the antibody structure file based on the coarse-grained antibody file and the coarse-grained antigen file to obtain an antigen surface residue information set and an antibody surface residue information set; generating an antibody binding region residue list according to the antibody surface residue information set; generating an epitope model based on the antibody binding region residue list; performing spatial matching on the antigen surface residue information set and the epitope model to obtain a matched epitope model; generating a set of antibody docking conformations according to the matched epitope model; performing local optimization on each antibody docking conformation in the set of antibody docking conformations using the coarse-grained antigen file to obtain a set of optimized antibody conformations; and performing screening on each optimized antibody conformation in the set of optimized antibody conformations according to the coarse-grained antigen file to obtain a target antigen-antibody binding conformation.
[0087] The embodiments of the present disclosure further provide a computer readable storage medium, wherein the computer readable storage medium stores a computer program, and the computer program includes program instructions. When the program instructions are executed, a method is implemented, which can refer to each embodiment of the method of the present disclosure.
[0088] The computer readable storage medium can be an internal storage unit of the computer device, for example, a hard disk or a memory of the computer device. The computer readable storage medium can also be an external storage device of the computer device, for example, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc.
[0089] It should be noted that, in this document, the terms "comprising", "including", or any other variant thereof are intended to cover a non-exclusive inclusion, such that processes, methods, articles, or systems that comprise a list of elements do not include only those elements, but also other elements that are not expressly listed, or other elements that are inherent in such processes, methods, articles, or systems. Without more limitations, an element defined by the phrase "comprising a" does not exclude the existence of additional identical elements in the process, method, article, or system that includes the element.
[0090] The above description is merely exemplary of some preferred embodiments of the present disclosure and of the principles thereof. It is to be understood that the present disclosure is not limited to the specific technical features described above, and that the scope of the present disclosure is not limited to the technical solutions formed by the specific combinations of the above technical features, but also covers other technical solutions formed by the combinations of the above technical features or equivalent features thereof without departing from the above inventive concept. For example, the above technical features are replaced with the technical features with similar functions disclosed in the embodiments of the present disclosure (but not limited to) to form technical solutions.
Claims
1. An iterative closest point based method for optimization of antigen-antibody binding conformation, characterized in that, The method comprises the following steps: coarse-grained antibody files and coarse-grained antigen files are obtained by coarse-graining antigen structure files and antibody structure files; antigen surface residue information sets and antibody surface residue information sets are obtained by surface residue extraction of the antigen structure files and the antibody structure files based on the coarse-grained antibody files and the coarse-grained antigen files; antibody binding region residue lists are generated according to the antibody surface residue information sets; epitope models are generated based on the antibody binding region residue lists; matching epitope models are obtained by spatial matching of the antigen surface residue information sets and the epitope models; antibody docking conformation sets are generated according to the matching epitope models; optimized antibody conformation sets are obtained by local optimization of each antibody docking conformation in the antibody docking conformation set using the coarse-grained antigen files; target antigen-antibody binding conformations are obtained by screening each optimized antibody conformation in the optimized antibody conformation set according to the coarse-grained antigen files.
2. The method of claim 1, wherein, The method comprises the following steps: antigen side chain atomic coordinate sequences and antibody side chain atomic coordinate sequences are obtained by extracting antigen side chain atomic coordinates and antibody side chain atomic coordinates in the antigen structure files and the antibody structure files; antigen centroid coordinate groups are obtained by determining centroid coordinates corresponding to the antigen side chain atomic coordinate sequences, and each antigen centroid coordinate in the antigen centroid coordinate groups is combined into a coarse-grained antigen file according to the antigen structure order in the antigen structure files; antibody centroid coordinate groups are obtained by determining centroid coordinates corresponding to the antibody side chain atomic coordinate sequences, and each antibody centroid coordinate in the antibody centroid coordinate groups is combined into a coarse-grained antibody file according to the antibody structure order in the antibody structure files.
3. The method of claim 2, wherein, The method comprises the following steps: antibody solvent solubility surface area sequences are obtained by determining solvent accessibility surface areas of each amino acid residue in the antibody structure files; antibody surface residue information in the antibody solvent solubility surface area sequences that meet preset residue screening conditions and corresponding residue chain information are determined as antibody surface residue information, and are applied to the coarse-grained antibody files to obtain antibody surface residue information sets; antigen solvent solubility surface area sequences are obtained by determining solvent accessibility surface areas of each amino acid residue in the antigen structure files; antigen surface residue information in the antigen solvent solubility surface area sequences that meet preset residue screening conditions and corresponding residue chain information are determined as antigen surface residue information, and are applied to the coarse-grained antigen files to obtain antigen surface residue information sets.
4. The method of claim 3, wherein, The antibody surface residue information in the antibody surface residue information set comprises an amino acid sequence, and the antibody binding region residue list is generated according to the antibody surface residue information set. The antibody surface residue information set is subjected to standardization numbering to obtain an amino acid numbering information sequence, and a complementarity determining region residue numbering sequence group is marked in the amino acid numbering information sequence, wherein the amino acid numbering information comprises a standard heavy chain numbering sequence, a first numbering mapping relationship sequence, a standard light chain numbering sequence, and a second numbering mapping relationship sequence; A solvent solubility surface area corresponding to each complementarity determining region residue numbering in the complementarity determining region residue numbering sequence group is determined to obtain a residue surface area sequence group; Each complementarity determining region residue numbering in the complementarity determining region residue numbering sequence group is sorted according to the residue surface area sequence group to obtain a sorted residue numbering sequence group; A structure curvature corresponding to each sorted residue numbering in the sorted residue numbering sequence group is determined to obtain a structure curvature sequence group; A target number of complementarity determining region residue numberings are screened from each complementarity determining region residue numbering sequence in the complementarity determining region residue numbering sequence group according to the structure curvature sequence group, and the target number of complementarity determining region residue numberings are combined into an antibody binding region residue list.
5. The method of claim 4, wherein, The antibody binding region residue list is used as a template to determine a corresponding residue space distribution range; A candidate point distribution range is set according to the residue space distribution range and a preset candidate point number; An initial candidate point coordinate sequence is generated in the candidate point distribution range, wherein the distance between adjacent initial candidate point coordinates in the initial candidate point coordinate sequence is greater than or equal to a preset distance threshold; Each initial candidate point coordinate in the initial candidate point coordinate sequence is screened based on the residue space distribution range to obtain an epitope model. The antigen surface residue information set and the epitope model are subjected to spatial matching to obtain a matched epitope model, comprising:
6. The method of claim 5, wherein, A global epitope candidate point centroid coordinate of each epitope candidate point in the epitope model is determined; The global epitope candidate point centroid coordinate and each antigen surface residue information in the antigen surface residue information set are subjected to alignment processing to obtain a matched epitope model. The matched epitope model is used to generate an antibody docking conformation set, comprising:
7. The method of claim 6, wherein, The matched epitope model is subjected to global rotation and translation operations according to a preset rotation angle group to obtain a rotation conformation set; The rotation conformation set is subjected to iterative closest point spatial matching with the antigen surface residue information set as an alignment target to obtain an iterative epitope model set; An iterative epitope model is sorted and screened from the iterative epitope model set, and each epitope candidate point in the selected iterative epitope model is used as an epitope positioning point group to obtain an epitope positioning point group set; The epitope model is aligned with each epitope localization point group in the epitope localization point group set to obtain an aligned rotation matrix and an aligned translation vector. Each aligned rotation matrix and each aligned translation vector are then applied to the coarse-grained antibody file to generate an antibody docking conformation, resulting in an antibody docking conformation set.
8. The method of claim 7, wherein, The step involves using the coarse-grained antigen file to locally optimize each antibody docking conformation in the antibody docking conformation set, resulting in an optimized antibody conformation set, including: From the coarse-grained antigen file and the epitope localization point group set, select multiple matching nearest point coordinates from each epitope localization point group to obtain the antigen nearest point coordinate sequence set. Based on the antigen nearest point coordinate sequence set and the antibody binding region residue list, the nearest point optimization is performed on each antibody docking conformation in the antibody docking conformation set to obtain the optimized antibody conformation set.
9. The method of claim 8, wherein, The step of screening each optimized antibody conformation in the optimized antibody conformation set according to the coarse-grained antigen file to obtain the target antigen antibody binding conformation includes: Based on the coarse-grained antigen file, the interaction potential energy of each optimized antibody conformation in the optimized antibody conformation set is determined to obtain the interaction potential energy set. Based on the set of interaction potential energies, optimized antibody conformations that meet the conformation screening conditions are selected from the set of optimized antibody conformations and used as the target antigen-antibody binding conformations.
10. An iterative closest point based antigen-antibody binding conformation optimization apparatus, characterized by, include: The coarse-graining processing unit is configured to coarse-grain the antigen structure file and the antibody structure file to obtain coarse-grained antibody file and coarse-grained antigen file; The extraction unit is configured to extract surface residues from the antigen structure file and the antibody structure file based on the coarse-grained antibody file and the coarse-grained antigen file, to obtain an antigen surface residue information set and an antibody surface residue information set; The first generation unit is configured to generate a list of antibody-binding region residues based on the antibody surface residue information set; The second generation unit is configured to generate an epitope model based on the list of residues in the antibody-binding region. A spatial matching unit is configured to perform spatial matching between the antigen surface residue information set and the epitope model to obtain a matched epitope model. The third generation unit is configured to generate a set of antibody docking conformations based on the matched epitope model; The local optimization unit is configured to use the coarse-grained antigen file to locally optimize each antibody docking conformation in the antibody docking conformation set to obtain an optimized antibody conformation set. The screening unit is configured to screen each optimized antibody conformation in the optimized antibody conformation set according to the coarse-grained antigen file to obtain the target antigen antibody binding conformation.
Citation Information
Patent Citations
Antigen-antibody binding site determination method and device, equipment and storage medium
CN115116543A
Deep learning model-based antibody structure optimization method and device
CN116741260A
Screening method and device for protein-protein binding interface
CN117577166A
Prediction method, device and equipment of antibody and antigen binding epitope and storage medium
CN118380042A
Method, device and equipment for generating antibody based on human antibody skeleton library and storage medium
CN118918981A