Iterative closest point based method and apparatus for optimizing antigen-antibody binding conformation

By using an iterative nearest-point antigen-antibody binding conformation optimization method, a list of antibody-binding region residues is generated through coarse-graining and surface residue extraction. Spatial matching and local optimization are then performed to solve the problem of weak antibody flexibility complementarity in antigen-antibody prediction, thereby improving the accuracy and error tolerance of antibody conformation.

CN120913633BActive Publication Date: 2026-02-10BEIJING ANBAISHENG DIAGNOSTIC TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511429424.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-09
Publication Date
2026-02-10
Estimated Expiration
2045-10-09

AI Technical Summary

Technical Problem

Existing protein docking software is not effective in predicting antigens and antibodies, mainly because antibodies have weak complementarity and the CDR region is flexible, leading to an over-reliance on rigid complementarity and resulting in low error tolerance. Existing methods have failed to effectively simulate the antigen-antibody binding pattern.

Method used

An iterative nearest-point antigen-antibody binding conformation optimization method is adopted. By coarsely processing antigen and antibody structural files, surface residue information is extracted to generate a list of antibody binding region residues and an epitope model. Spatial matching and local optimization are then performed to screen out the target antigen-antibody binding conformation.

Benefits of technology

It improves the accuracy and error tolerance of antibody conformation, enhances the geometric matching precision of antigen-antibody binding and the hit rate of near-native structures, and solves the problem of low matching degree between the overall antibody posture and antigen in existing methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120913633B_ABST
    Figure CN120913633B_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure disclose an antigen-antibody binding conformation optimization method and device based on iterative closest point. A specific embodiment of the method comprises: obtaining a coarse-grained antibody file and a coarse-grained antigen file; performing surface residue extraction on the antigen structure file and the antibody structure file to obtain an antigen surface residue information set and an antibody surface residue information set; generating an antibody binding region residue list; generating an epitope model; performing spatial matching on the antigen surface residue information set and the epitope model to obtain a matched epitope model; generating an antibody docking conformation set; performing local optimization on each antibody docking conformation in the antibody docking conformation set to obtain an optimized antibody conformation set; and screening each optimized antibody conformation in the optimized antibody conformation set to obtain a target antigen-antibody binding conformation. The embodiment can improve the accuracy of the antibody conformation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments disclosed herein relate to the fields of antigen-antibody docking model technology and computer technology, specifically to a method and apparatus for optimizing antigen-antibody binding conformation based on iterative nearest point. Background Technology

[0002] Computational biology methods for protein structure prediction are often referred to as docking techniques. Existing techniques employ various approaches to handle these stages, such as Monte Carlo simulations, Fast Fourier Transform (FFT), spherical harmonics, and energy-based techniques, all of which have been applied to the docking field to pursue better results. Meanwhile, the choice of scoring functions is also quite diverse, including those based on geometric complementarity, force fields, knowledge, machine learning, or combinations thereof, to identify near-natural structures.

[0003] However, currently available protein docking software is less effective than conventional protein docking in predicting antigens and antibodies. This is mainly because antibodies primarily bind to antigenic epitopes via CDRs (Complementarity-Determining Regions). The complementarity between CDRs and antigens is weaker than that of conventional protein-protein interactions. Furthermore, CDRs (especially the H3 (Heavy Chain Complementarity-Determining Region 3)) are highly flexible, and over-reliance on rigid complementarity leads to a lower tolerance for antibody conformational errors. Summary of the Invention

[0004] The summary portion of this disclosure is intended to provide a brief overview of the concepts, which will be described in detail in the detailed description portion. This summary portion is not intended to identify key or essential features of the claimed technical solutions, nor is it intended to limit the scope of the claimed technical solutions.

[0005] Some embodiments of this disclosure propose a method and apparatus for optimizing antigen-antibody binding conformation based on iterative nearest point to solve the technical problems mentioned in the background section above.

[0006] In a first aspect, some embodiments of this disclosure provide an antigen-antibody binding conformation optimization method based on iterative nearest point. The method includes: coarse-graining an antigen structure file and an antibody structure file to obtain a coarse-grained antibody file and a coarse-grained antigen file; extracting surface residues from the antigen structure file and the antibody structure file based on the coarse-grained antibody file and the coarse-grained antigen file to obtain an antigen surface residue information set and an antibody surface residue information set; generating an antibody binding region residue list based on the antibody surface residue information set; generating an epitope model based on the antibody binding region residue list; spatially matching the antigen surface residue information set and the epitope model to obtain a matched epitope model; generating an antibody docking conformation set based on the matched epitope model; locally optimizing each antibody docking conformation in the antibody docking conformation set using the coarse-grained antigen file to obtain an optimized antibody conformation set; and screening each optimized antibody conformation in the optimized antibody conformation set based on the coarse-grained antigen file to obtain a target antigen-antibody binding conformation.

[0007] Secondly, some embodiments of this disclosure provide an antigen-antibody binding conformation optimization apparatus based on iterative nearest point. The apparatus includes: a coarse-graining processing unit configured to coarse-grain an antigen structure file and an antibody structure file to obtain a coarse-grained antibody file and a coarse-grained antigen file; an extraction unit configured to extract surface residues from the antigen structure file and the antibody structure file based on the coarse-grained antibody file and the coarse-grained antigen file to obtain an antigen surface residue information set and an antibody surface residue information set; a first generation unit configured to generate a list of antibody binding region residues based on the antibody surface residue information set; and a second generation unit configured to... Based on the above list of antibody-binding region residues, an epitope model is generated; a spatial matching unit is configured to perform spatial matching between the above set of antigen surface residue information and the above epitope model to obtain a matched epitope model; a third generation unit is configured to generate an antibody docking conformation set based on the above matched epitope model; a local optimization unit is configured to use the above coarse-grained antigen file to locally optimize each antibody docking conformation in the above antibody docking conformation set to obtain an optimized antibody conformation set; a screening unit is configured to screen each optimized antibody conformation in the above optimized antibody conformation set based on the above coarse-grained antigen file to obtain the target antigen antibody-binding conformation.

[0008] Thirdly, some embodiments of this disclosure provide an electronic device, including: one or more processors; and a storage device having one or more programs stored thereon, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any implementation of the first aspect above.

[0009] Fourthly, some embodiments of this disclosure provide a computer-readable medium having a computer program stored thereon, wherein the program, when executed by a processor, implements the method described in any of the implementations of the first aspect above.

[0010] The various embodiments of this disclosure have the following beneficial effects: the accuracy of the determined antibody conformation can be improved by using the antigen-antibody binding conformation optimization method based on iterative nearest point according to some embodiments of this disclosure. Specifically, the antigen-antibody binding conformation optimization method based on iterative nearest point according to some embodiments of this disclosure firstly coarsely processes the antigen structure file and antibody structure file to obtain coarsely processed antibody file and coarsely processed antigen file. Here, coarse processing simplifies the antigen and antibody structures (retaining the CA (i.e., α-carbon atom) atoms simulated in the side chain center), reducing computational complexity while accurately preserving the core spatial information related to antigen-antibody binding, providing an efficient computational basis for subsequent CDR region and epitope matching. Then, based on the coarsely processed antibody file and coarsely processed antigen file, surface residues are extracted from the antigen structure file and antibody structure file to obtain an antigen surface residue information set and an antibody surface residue information set. Here, by extracting surface residues, interference from irrelevant regions can be reduced, improving the targeting of subsequent docking. Afterwards, an antibody binding region residue list is generated based on the antibody surface residue information set. Here, by generating a list of antibody-binding region residues, the core antibody-binding region can be precisely located. Then, based on this list, an epitope model is generated. This epitope model simulates the spatial characteristics of the original epitope, compensating for the weak complementarity between the CDR region and the antigen, providing a more forgiving matching space for the flexible CDR region, and improving fault tolerance. Next, spatial matching is performed between the antigen surface residue information set and the epitope model to obtain a matched epitope model. Spatial matching narrows the subsequent search range, avoiding the inefficiency of random sampling across the entire space in conventional methods, and improving docking efficiency. Then, based on the matched epitope model, an antibody docking conformation set is generated. This generated docking conformation aligns the coarse-grained antibody with the epitope conformation, allowing the antibody to initially align with the antigen-binding site, solving the problem of low overall antibody posture matching with antigen in conventional methods. Furthermore, using the coarse-grained antigen file, each antibody docking conformation in the antibody docking conformation set is locally optimized to obtain an optimized antibody conformation set. Here, through local optimization, the flexible characteristics of the CDR region can be adapted, avoiding the low fault tolerance problem caused by rigid complementarity, and significantly improving the geometric matching accuracy of the binding interface. Finally, based on the above coarse-grained antigen file, each optimized antibody conformation in the optimized antibody conformation set is screened to obtain the target antigen-antibody binding conformation. Here, through final screening, the hit rate of near-native structures can be improved, overcoming the defect of conventional scoring functions that ignore antigen-antibody specific interactions. Attached Figure Description

[0011] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and elements are not necessarily drawn to scale.

[0012] Figure 1 This is a flowchart of some embodiments of the antigen-antibody binding conformation optimization method based on iterative nearest point according to this disclosure;

[0013] Figure 2 This is a schematic diagram showing the positional relationship between the amino acid chain, CDR region, and epitope model in a three-dimensional representation.

[0014] Figure 3 This is a schematic diagram of the structure of some embodiments of the antigen-antibody binding conformation optimization device based on iterative nearest point according to the present disclosure;

[0015] Figure 4 This is a schematic diagram of the structure of an electronic device suitable for implementing some embodiments of the present disclosure. Detailed Implementation

[0016] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0017] It should also be noted that, for ease of description, only the parts relevant to the invention are shown in the accompanying drawings. Unless otherwise specified, the embodiments and features described in this disclosure can be combined with each other.

[0018] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.

[0019] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".

[0020] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.

[0021] In practice, among the mainstream classical algorithms for antibody conformation, the FFT algorithm significantly improves the speed of exhaustive sampling of the search space. However, FFT-based methods have major obstacles: first, they require a correlation function to express the energy; second, they cannot infer pose quality from the value of the energy function. Despite these limitations, this type of method has gained widespread acceptance. In addition to exhaustive methods, docking can also be achieved by exploring the geometric features of the protein surface. For example, tools such as PatchDock and SP-Dock align proteins based on complementary surface patches. Histogram-based descriptors are also popular when finding complementary regions. For example, the Zernike descriptor (ZD) used by LZerD can represent the protein surface shape in the terms of its 3D function series expansion.

[0022] The Iterative Closest Point (ICP) algorithm is a classic 3D shape matching algorithm. Its basic principle is to iteratively calculate and optimize the rotation and translation transformations between the data point cloud and the model point cloud, aiming to minimize the distance between them to achieve point cloud alignment (registration). In recent years, the ICP algorithm has not only been widely used in computer vision, but has also been reported in biology and chemistry, such as for protein structure comparison.

[0023] However, currently available protein docking software is less effective than conventional protein docking in predicting antigens and antibodies. This is mainly because antibodies primarily bind to antigenic epitopes through CDRs (Complementarity-Determining Regions), and the complementarity between CDRs and antigens is weaker than that of conventional protein-protein interactions. Furthermore, CDRs (especially the H3 (Heavy Chain Complementarity-Determining Region 3)) are highly flexible, and over-reliance on rigid complementarity leads to a lower tolerance for errors. For example, SP-Dock protein docking software (SP-Dock: Protein-Protein Docking using Shape and Physicochemical Complementarity) uses the ICP algorithm to achieve shape complementarity by simultaneously matching adjacent surface patches of the receptor and ligand. The specific process is as follows: First, the MSMS (Maximal Speed ​​Molecular Surface) algorithm is used to generate a solvent-excluded surface (SES). Key points are extracted based on local curvature to generate geodesic surface patches (GSPs). Then, the ICP algorithm is used for protein surface alignment to optimize the rigidity transformation of the ligand and receptor patch groups. However, like other protein docking software, due to an overemphasis on the complementarity of the receptor and ligand, and the failure of the number and quality of residues contained in the surface patches to specifically represent antibody binding sites and antigen epitopes, the support effect for antigen-antibody docking is ultimately poor. Here, existing antibody AI generation models have significant defects: they generate antibodies through machine learning of antigen-antibody binding rules, but the success rate is extremely low when the same strategy is applied to antigen-antibody docking. This indicates that machine learning has not truly grasped the underlying binding rules, resulting in the success rate and reliability of antibody generation failing to meet practical needs. Therefore, specialized antigen-antibody docking software is urgently needed to guide antibody AI model generation. Based on this technical problem, this disclosure addresses it through the following methods.

[0024] This disclosure will now be described in detail with reference to the accompanying drawings and embodiments.

[0025] Figure 1 A flowchart 100 is shown, illustrating some embodiments of the antigen-antibody binding conformation optimization method based on iterative nearest point according to this disclosure. This antigen-antibody binding conformation optimization method based on iterative nearest point includes the following steps:

[0026] Step 101: Coarse-grained processing is performed on the antigen structure file and antibody structure file to obtain coarse-grained antibody file and coarse-grained antigen file.

[0027] In some embodiments, the execution entity (e.g., a computing device) of the antigen-antibody binding conformation optimization method based on iterative nearest point can coarsely process the antigen structure file and antibody structure file to obtain coarse-grained antibody file and coarse-grained antigen file. The antigen structure file includes antigen structure information of the free antigen. The antibody structure file includes antibody structure information of the free antibody.

[0028] It should be noted that the aforementioned computing devices can be either hardware or software. When the computing device is hardware, it can be implemented as a distributed cluster consisting of multiple servers or terminal devices, or as a single server or a single terminal device. When the computing device is software, it can be installed on the hardware devices listed above. It can be implemented as, for example, multiple software programs or software modules used to provide distributed services, or as a single software program or software module. No specific limitations are made here.

[0029] Coarse-graining can be achieved by simplifying the data structures of antigen and antibody structure files. For example, the antigen and antibody PDB (Protein Data Bank File, a standardized file format for storing the three-dimensional spatial structure information of biological macromolecules such as proteins, nucleic acids, and carbohydrates) structure files can be coarsened, retaining only the simulated CA atoms at the side chain centers. The side chain atom coordinates of the residues are extracted, i.e., atoms other than the main chain 'N' (nitrogen atom), 'CA' (α-carbon atom), 'C' (carbon atom), and 'O' (oxygen atom). The centroids of the side chain atoms are calculated and represented by virtual CA atoms. Glycine (GLY) has no side chain atoms, so it is replaced by the main chain CA coordinates.

[0030] Specifically, the aforementioned executing entity performs coarse-graining processing on the antigen structure file and antibody structure file to obtain coarse-grained antibody file and coarse-grained antigen file, including the following steps:

[0031] Step S1: Extract the atomic coordinates of the antigen side chain and the antibody side chain from the antigen structure file and the antibody structure file to obtain the atomic coordinate sequences of the antigen side chain and the antibody side chain. The antigen structure file (i.e., the PDB file) includes information such as the atomic coordinates and amino acid sequences. Therefore, the atomic coordinates of the antigen side chain and the antibody side chain can be extracted using an indexing method.

[0032] Step S2 involves determining the centroid coordinates corresponding to the aforementioned antigen side chain atom coordinate sequence, obtaining an antigen centroid coordinate set, and combining the centroid coordinates of each antigen in the aforementioned antigen structure file into a coarse-grained antigen file according to the antigen structure order in the aforementioned antigen structure file. First, the atomic mass corresponding to each antigen side chain atom coordinate is obtained. Then, the average isotopic mass of the atoms can be substituted into the centroid calculation formula to obtain the antigen centroid coordinates. Here, the centroid is calculated using a mass-weighted average of the side chain atoms, rather than a simple geometric average, because it more accurately reflects the spatial distribution center of the residue side chains.

[0033] Step S3 involves determining the centroid coordinates corresponding to the aforementioned antibody side chain atomic coordinate sequences, obtaining an antibody centroid coordinate set, and combining the individual antibody centroid coordinates in the set into a coarse-grained antibody file according to the antibody structure order in the aforementioned antibody structure file. The generation method and achieved technical effects of the antibody centroid coordinates can be referenced in step S2 above and will not be elaborated further. Here, the centroid coordinates can be three-dimensional coordinates.

[0034] Step 102: Based on the coarse-grained antibody file and the coarse-grained antigen file, surface residues are extracted from the antigen structure file and the antibody structure file to obtain the antigen surface residue information set and the antibody surface residue information set.

[0035] In some embodiments, the execution entity may extract surface residues from the antigen structure file and the antibody structure file based on the coarse-grained antibody file and the coarse-grained antigen file to obtain a (coarse-grained) antigen surface residue information set and a (coarse-grained) antibody surface residue information set.

[0036] In some optional implementations of certain embodiments, the execution entity extracts surface residues from the antigen structure file and the antibody structure file based on the coarse-grained antibody file and the coarse-grained antigen file to obtain an antigen surface residue information set and an antibody surface residue information set, including:

[0037] Step S1: Determine the solvent-accessible surface area of ​​each amino acid residue in the antibody structure file to obtain the antibody solvent-soluble surface area sequence. The antibody structure file can be input into the FreeSASA tool to calculate the solvent-accessible surface area (SASA) of each amino acid residue.

[0038] Step S2 involves identifying the antibody solvent-soluble surface area and corresponding residue chain information that meet the preset residue screening conditions in the above antibody solvent-soluble surface area sequence as antibody surface residue information, and applying this information to the coarse-grained antibody file to obtain an antibody surface residue information set. The residue chain information may include: residue chain identifier, residue number, amino acid type, solvent-accessible surface area, and surface residue atomic coordinates. Furthermore, the preset residue screening conditions may be that the antibody solvent-soluble surface area or the antigen solvent-soluble surface area is within a preset area range (e.g., [10-30] square angstroms). Applying the residue chain information to the coarse-grained antibody file may involve deleting non-exposed residues from the coarse-grained antibody file, excluding the individual residue chain information, so that the remaining residue chain information in the coarse-grained antibody file is consistent with the screened residue chain information.

[0039] Step S3: Determine the solvent-accessible surface area of ​​each amino acid residue in the above antigen structural file to obtain the antigen solvent-soluble surface area sequence.

[0040] Step S4: The antigen solvent soluble surface area and corresponding residue chain information that meet the preset residue screening conditions in the above antigen solvent soluble surface area sequence are determined as antigen surface residue information and applied to the coarse-grained antigen file to obtain the antigen surface residue information set.

[0041] Specifically, in the PDB files of antigens and antibodies, each residue (amino acid) has its own original number in the structural record (defined by the "resSeq" field in the ATOM line of the PDB file). For example, in the PDB file of an antibody light chain, residues may be numbered sequentially as "1, 2, 3...", while antigen residues may be numbered as "101, 102..." or independently (residue numbers may be repeated between different chains, requiring differentiation based on chain ID, such as chain L, chain H, chain A, etc.). Here, the residue sequence number directly uses the original numbering from the PDB file and is associated with the chain ID (e.g., L chain residues 1-100, H chain residues 1-120, antigen A chain residues 1-200) to mark exposed residues. Applying residue chain information to coarse-grained antigen files can be achieved by deleting non-exposed residues from the coarse-grained antigen file, excluding the individual residue chain information, so that the remaining residue chain information in the coarse-grained antigen file is consistent with the selected residue chain information.

[0042] Step 103: Generate a list of antibody-binding region residues based on the antibody surface residue information set.

[0043] In some embodiments, the aforementioned executing entity may generate a list of antibody-binding region residues based on the aforementioned antibody surface residue information set.

[0044] In some optional implementations of certain embodiments, the antibody surface residue information in the antibody surface residue information set includes amino acid sequences. The executing entity generates a list of antibody-binding region residues based on the aforementioned antibody surface residue information set, including:

[0045] Step S1 involves standardizing and numbering the amino acid sequences included in the antibody surface residue information set to obtain an amino acid numbering information sequence, and marking the complementarity-determining region (PDR) residue numbering sequence group within the amino acid numbering information sequence. The amino acid numbering information includes a standardized heavy chain numbering sequence, a first numbering mapping sequence, a standardized light chain numbering sequence, and a second numbering mapping sequence. Specifically, the amino acid sequences in each antibody surface residue information set can be standardized and numbered using antibody CDR region partitioning methods (e.g., the KABAT rule) to obtain the amino acid numbering information sequence. Specifically, the amino acid sequences can be numbered using Kabat numbering tools (such as Abnum numbering software, IgBLAST numbering tools, etc.), and the CDR1, CDR2, and CDR3 regions of the heavy and light chains can be marked according to the KABAT rule (i.e., determining the residue numbers contained in each CDR region). Thus, the first and second numbering mapping relationships can represent the mapping relationship between the standardized heavy chain numbering, the Kabat numbering, and the original PDB numbering. Here, at least one complementarity-determining region and the residue number within the complementarity-determining region can be marked in the above amino acid numbering information sequence using the KABAT rule described above, thus obtaining the complementarity-determining region residue numbering sequence.

[0046] Step S2 involves determining the solvent-soluble surface area corresponding to each complementary determinant residue number in the above complementary determinant residue numbering sequence group, thus obtaining the residue surface area sequence group. The solvent-soluble surface area corresponding to the complementary determinant residue number can be determined using the FreeSASA tool.

[0047] In practice, the SASA threshold (10-30) used in the above-mentioned surface residue extraction steps is a general standard applicable to the initial screening of global residues. However, as the functional core of the antibody, the "exposure" of residues in the CDR region has special significance: 1. The SASA value of some CDR region residues may be slightly lower than the threshold (e.g., 8-10 square angstroms), but because they are located in key positions in the CDR region (e.g., the edge of the binding site), they may still participate in antigen binding, and their precise SASA value needs to be calculated separately to confirm whether they are exposed; 2. Conversely, although some CDR region residues are included in the "surface residues", their SASA value needs to be calculated separately to verify whether it truly reflects the degree of exposure (to avoid misjudgment due to errors in global calculations).

[0048] Step S3: According to the above residue surface area sequence group, sort the residue numbers of each complementarity-determining region in each complementarity-determining region residue numbering sequence in the above complementarity-determining region residue numbering sequence group to obtain a sorted residue numbering sequence group. The residues can be sorted in descending order of surface area.

[0049] Step S4: Determine the structural curvature corresponding to each sorted residue number in the above-mentioned sorted residue number sequence group to obtain the structural curvature sequence group. The calculation of structural curvature is based on the three-dimensional spatial structure of the antibody variable region. The core is to quantify the degree of local structural curvature of the region through the atomic / amino acid coordinates surrounding the residue. Specifically, the structural curvature of the sorted residue number can be calculated through the following steps: 1. For each residue in the CDR region, select all atoms (including main chain and side chain atoms) within a certain range (e.g., 5-10 Å) around it and obtain the corresponding three-dimensional spatial coordinates (x, y, z). 2. Using the least squares method, fit the extracted atomic coordinates to an ideal surface (usually a sphere or plane). If it is fitted to a sphere, calculate the radius of the sphere (radius of curvature R), then the curvature K = 1 / R (the smaller the radius, the greater the curvature, indicating that the local structure is more "convex" or "concave"); if it is fitted to a plane, the curvature is 0 (indicating that the local structure is flat and not easily involved in specific binding). Finally, the calculated curvature values ​​are normalized (e.g., scaled to 0-1 according to the curvature range of all residues in the CDR region) to facilitate comparison between different residues and highlight high-curvature residues, thus obtaining the structural curvature sequence group. Each sorted residue number can correspond to a structural curvature.

[0050] Step S5: Based on the above structural curvature sequence group, select a target number (e.g., 1-5) of complementary determinant region (CDR) residue numbers from each CDR residue numbering sequence in the above complementary determinant region (CDR) residue numbering sequence group, and combine the target number of CDR residue numbers into an antibody-binding region residue list. In practice, residue surface area reflects the degree to which residues are exposed in solvent. A residue surface area > 20 square angstroms indicates that most of the residue side chain is exposed, possessing the spatial possibility of binding to the antigen (if the residue is deeply buried inside the protein, even with high curvature, it cannot reach the antigen). Secondly, the side chain structure of residues with high curvature (such as those located at the top of the CDR loop or the edge of the "pocket") is more likely to form specific interactions (such as hydrogen bonds and van der Waals forces) with the complementary structure (such as depressions and protrusions) of the antigen epitope. Therefore, the residue numbering screening conditions can be: preferentially screening the target number of CDR residue numbers with a residue surface area greater than 20 square angstroms and / or those with higher structural curvature.

[0051] Step 104: Generate an epitope model based on the list of antibody-binding region residues.

[0052] In some embodiments, the aforementioned executing entity may generate an epitope model based on the aforementioned list of antibody-binding region residues.

[0053] As an example, a tabletop model can be generated using the following steps:

[0054] 1. Extract relevant information from the antibody surface residue PDB file and CDR file, calculate the coordinate range of the residues in the CDR region, and then generate uniformly distributed coordinate points within this range.

[0055] 2. Translate the coordinate point along the antibody direction, with a random stopping condition of 2-6 angstroms between the coordinate point and the CDR residue. If the distance between the coordinate point and the non-CDR residue is less than the distance between the coordinate point and the CDR residue, delete these conflicting points.

[0056] like Figure 2 As shown, blue represents the amino acid chain, yellow represents the CDR region, white represents the coordinate points generated based on the CDR coordinates, and green represents the generated epitope model.

[0057] In some optional implementations of certain embodiments, the execution entity generates an epitope model based on the aforementioned list of antibody-binding region residues, including:

[0058] Step S1: Using the aforementioned list of antibody-binding region residues as a template, determine the corresponding spatial distribution range of residues. This can be achieved using surface generation tools (e.g., PyMOL molecular visualization system's showsurface, MSMS software) to discretize the surface into grid points (each grid point represents a point on the surface, recording its three-dimensional coordinates and normal vector), thus creating a spatial distribution grid model. Here, each square angstrom can contain 1-2 grid points, thereby determining the spatial distribution range of residues in each antibody-binding region. For example, this includes the minimum circumscribed sphere radius and center coordinates of the residue space.

[0059] Step S2: Set the distribution range of candidate points according to the above-mentioned spatial distribution range of residues and the preset number of candidate points (e.g., 20-150).

[0060] As an example, the candidate point distribution range can be a spherical region with a radius of 5-10 angstroms around the center of the spatial distribution range of the residues, thus obtaining the candidate point distribution range.

[0061] Step S3: Within the aforementioned candidate point distribution range, generate an initial candidate point coordinate sequence. This initial candidate point coordinate sequence can be generated within the aforementioned candidate point distribution range using a random sampling algorithm (e.g., uniform distribution sampling, Gaussian distribution sampling, etc.). The distance between adjacent initial candidate point coordinates in the initial candidate point coordinate sequence is greater than or equal to a preset distance threshold (e.g., 3 angstroms).

[0062] Step S4: Based on the aforementioned spatial distribution range of residues, the coordinates of each initial candidate point in the aforementioned initial candidate point coordinate sequence are screened to obtain the epitope model. This screening can be performed through the following steps:

[0063] 1. Calculate the relative positions of the initial candidate points and the CDR area surface. Specifically, for each of the generated initial candidate point coordinates (20-150 randomly distributed points), calculate the shortest distance from each point to the CDR area surface grid point, as well as the relationship between the normal vector direction of that point and the surface grid point.

[0064] 2. Matching and filtering convex / concave features. If a candidate point is 3-5 angstroms away from a convex region in the CDR area and is located in the opposite direction of the normal vector of that region (i.e., pointing inwards towards the convexity), satisfying the "convex-concave" complementarity, it is retained. If a candidate point is 3-5 angstroms away from a concave region in the CDR area and is located in the direction of the normal vector of that region (i.e., pointing outwards towards the concaveness), satisfying the "concave-convex" complementarity, it is retained. If a candidate point is too close to the surface (e.g., distance < 2 angstroms, potential collision), too far (distance > 6 angstroms, incompatible), or does not match the convex / concave features (e.g., convexity corresponding to convexity), it is discarded.

[0065] 3. Control the number of candidate points. If fewer than 20 candidate points remain after screening, the distance constraint can be relaxed (e.g., 2-7 Å) or a new batch of initial candidate points can be generated for re-screening. If more than 150 remain, they are sorted by "complementary matching degree with the CDR region surface" (e.g., the closer the distance (3-5 Å) and the more consistent the normal vector direction, the higher the matching degree), and the top 150 are retained to obtain the epitope model. That is, the epitope model consists of multiple epitope candidate points.

[0066] Step 105: Spatial matching is performed on the antigen surface residue information set and the epitope model to obtain the matched epitope model.

[0067] In some embodiments, the execution entity may perform spatial matching between the antigen surface residue information set and the epitope model to obtain a matched epitope model. Specifically, the epitope model is spatially matched using the antigen surface residue information set as the alignment target to obtain the matched epitope model.

[0068] In some optional implementations of certain embodiments, the execution entity performs spatial matching between the antigen surface residue information set and the epitope model to obtain a matched epitope model, including:

[0069] Step S1: Determine the centroid coordinates of all candidate epitopes in the above epitope model. The centroid coordinates can be determined using the centroid formula.

[0070] Step S2: Align the centroid coordinates of the above overall epitope candidate points and each antigen surface residue information in the above antigen surface residue information set to obtain a matched epitope model. Among them, the alignment process can be carried out through the following steps:

[0071] First, determine the centroid coordinates of the coordinates of each antigen surface residue in the antigen surface residue information set, and move them to the 0 point as the residue centroid coordinates. Then, translate the centroid coordinates of the overall epitope candidate points to the residue centroid coordinates (i.e., the 0 point), and then translate d (0 < d ≤ 20 Å) in the negative direction of the horizontal axis, so that the centroid coordinates of the overall epitope candidate points are located at the position of (-d, 0, 0). Here, the distance d can be adjusted according to the antigen size. For small molecule antigens, take 5 - 10 Å, and for protein antigens, take 10 - 20 Å (ensure that the initial distance is reasonable to avoid too close collisions or too far without interaction) to obtain a matched epitope model.

[0072] (− d ,0,0)

[0073] Step 106: Generate an antibody docking conformation set according to the matched epitope model.

[0074] In some embodiments, the above execution subject can generate an antibody docking conformation set according to the above matched epitope model.

[0075] In some optional implementation manners of some embodiments, the above execution subject generates an antibody docking conformation set according to the above matched epitope model, including:

[0076] Step S1: Perform overall rotation and translation operations on the above matched epitope model according to a preset rotation angle group to obtain a rotation conformation set. Among them, ICP (Iterative Closest Point) iteration can be carried out through the following steps.

[0077] First, with the residue centroid coordinates as the rotation center, traverse and rotate each matched epitope candidate point in the matched epitope model around the X - Y - Z axes in sequence. At the same time, locate the centroid coordinates of the overall epitope candidate points of the epitope model and rotate for self - rotation compensation around the YZ axis. Thus, facing the antigen in a fixed direction, the angle is set to 2 - 15°, and thus a current epitope model is obtained for each rotation of a preset rotation angle as a rotation conformation, and a rotation conformation set is obtained.

[0078] Here, rotation compensation ensures that the binding surface of the matched epitope candidate site always faces the antigen when rotating around the centroid coordinates of the residue. Specifically, the rotation process is decomposed into two steps: first, rotation around the centroid coordinates of the residue (revolution), and then rotation around the centroid of the epitope model (rotation). When revolving around the Y-axis by an angle θ, the Y-axis simultaneously rotates by an angle θ; when revolving around the Z-axis by an angle φ, the Z-axis simultaneously rotates by an angle φ. This self-centroid rotation (rotation compensation) ensures that the binding surface of the epitope candidate site maintains its relative orientation to the antigen surface throughout the global rotation, preventing binding surface shift due to rotation.

[0079] Step S2: Using the antigen surface residue information set as the alignment target, perform iterative nearest-point spatial matching on the rotation conformation set to obtain the iterative epitope model set.

[0080] The rotating conformation set is iterated using ICP (Iterative Closest Point) through the following steps: The rotating conformation set is translated along the line connecting the centroid coordinates of its overall epitope candidate points and the centroid coordinates of its residues, with the translation direction moving closer to the antigen surface and away from the centroid coordinates of the residues, i.e., with the antigen surface as the target point set. The ICP rotation angle is limited to 2-60° along the XYZ axes. Finally, the iterated epitope model set is obtained through translation and rotation.

[0081] Step S3: Sort and select the iterative epitope models from the above iterative epitope model set, and take each epitope candidate point in the selected iterative epitope model as an epitope positioning point group to obtain an epitope positioning point group set.

[0082] Specifically, from the candidate epitope point group epitope model set after the above iteration, epitope models are sorted and selected as epitope location point groups to obtain the epitope location point group set.

[0083] Optionally, during the sorting and screening process: the RMSD (Root Mean Square Deviation) of the closest point pairs between each candidate epitope in the iterated epitope model and the antigen surface residue information set is used for sorting. This allows for the selection of a predetermined number of iterated epitope models with the smallest RMSD to the antigen surface residue information set. The epitope candidate points included in these models are then used as epitope localization point groups, resulting in a set of epitope localization point groups. For example, 2000-5000 epitope localization point groups can be obtained, forming the epitope localization point set.

[0084] Specifically, the sorting and screening process is as follows: For each iterated epitope model, the n (e.g., 5-15) point pairs that are closest to the antigen surface residue information set are identified, resulting in a point pair group. Then, the RMSD (Root Mean Square Deviation) value of each point pair group is determined. Next, the iterated epitope models are sorted according to their RMSD values, resulting in a sorted epitope model sequence. Finally, the M (e.g., 2000-5000) epitope models with the smallest RMSD values ​​in the sorted epitope model sequence are selected, and the included epitope candidate points are determined as the epitope localization point group, resulting in the epitope localization point group set.

[0085] Step S4 involves aligning the epitope model with each epitope localization point group in the epitope localization point set to obtain an aligned rotation matrix and an aligned translation vector. Each aligned rotation matrix and each aligned translation vector are then applied to the coarse-grained antibody file to generate an antibody docking conformation, resulting in a set of antibody docking conformations. Specifically, the alignment process in step 104 can be followed to align the epitope model with each epitope localization point group in the epitope localization point set, obtaining an aligned rotation matrix and an aligned translation vector. Furthermore, the antibody coordinates in the coarse-grained antibody file can be adjusted according to each aligned rotation matrix and each aligned translation vector to obtain the set of antibody docking conformations.

[0086] Step 107: Using the coarse-grained antigen file, perform local optimization on each antibody docking conformation in the antibody docking conformation set to obtain the optimized antibody conformation set.

[0087] In some embodiments, the execution entity may use the coarse-grained antigen file to locally optimize each docking conformation in the antibody docking conformation set to obtain an optimized antibody conformation set.

[0088] In some optional implementations of certain embodiments, the executing entity utilizes the coarse-grained antigen file to locally optimize each docking conformation in the antibody docking conformation set, obtaining an optimized antibody conformation set, including:

[0089] Step S1: Select multiple matching nearest point coordinates from each epitope localization point group in the coarse-grained antigen file and the epitope localization point group set to obtain the antigen nearest point coordinate sequence set. Matching can refer to corresponding antigen identifiers. Here, the coarse-grained antigen file used can also be a coarse-grained antigen file after deleting non-exposed residues.

[0090] Step S2: Based on the aforementioned antigen nearest-point coordinate sequence set and the aforementioned antibody binding region residue list, perform nearest-point optimization on each antibody docking conformation in the aforementioned antibody docking conformation set to obtain an optimized antibody conformation set. Specifically, the aforementioned ICP algorithm can be used to perform nearest-point optimization on each antibody docking conformation in the aforementioned antibody docking conformation set to obtain the optimized antibody conformation set. Here, during the iteration process, each antibody docking conformation rotates around its own centroid, and the rotation angles of the horizontal, vertical, and axial axes are limited to 2-60°.

[0091] Step 108: Based on the coarse-grained antigen file, screen each optimized antibody conformation in the optimized antibody conformation set to obtain the target antigen antibody binding conformation.

[0092] In some embodiments, the execution entity may, based on the coarse-grained antigen file, screen each optimized antibody conformation in the optimized antibody conformation set to obtain the target antigen-antibody binding conformation.

[0093] In some optional implementations of certain embodiments, the executing entity, based on the coarse-grained antigen file, filters each optimized antibody conformation in the optimized antibody conformation set to obtain the target antigen-antibody binding conformation, including:

[0094] Step S1: Based on the coarse-grained antigen file described above, determine the interaction potential energy of each optimized antibody conformation in the optimized antibody conformation set, thus obtaining the interaction potential energy set. The interaction potential energy set can be obtained by using the MJ (Miyazawa-Jernigan) potential calculation function. Here, the Miyazawa-Jernigan potential is based on the statistical interaction frequency between amino acid pairs, reflecting the preference of different amino acid combinations within the protein (e.g., the tendency of hydrophobic residues to aggregate).

[0095] Step S2: Based on the aforementioned set of interaction potential energies, select optimized antibody conformations that meet the conformation selection criteria from the aforementioned set of optimized antibody conformations, and use them as the target antigen-antibody binding conformations. The conformation selection criteria can be the optimized antibody conformation with the lowest interaction potential energy.

[0096] Alternatively, you can filter using the following steps:

[0097] 1. Extract contact region: Extract residue pairs with a concentration distance of less than 5 angstroms between the optimized antibody conformation and antigen surface residue information.

[0098] 2. For each residue pair, the corresponding interaction potentials in the interaction potential energy set are accumulated to obtain the accumulated potential energy value.

[0099] 3. Sort the optimized antibody conformations from smallest to largest according to the cumulative potential energy pairs, and select the optimized antibody conformation with the smallest interaction potential energy as the target antigen antibody binding conformation.

[0100] In practice, this invention has made targeted optimizations for antigen-antibody docking: A spatial point cloud (i.e., an epitope model) is generated by simulating the CDR region of the antibody. This spatial point cloud completely covers the actual antigen epitope distribution in terms of both quantity and distribution. Subsequently, using antigen surface residues as the target point set, the ICP algorithm is used to find multiple nearest points between the epitope model and the antigen surface residues. ICP optimization is then performed again using the determined nearest points and the antibody's CDR region to adjust the antibody's optimal posture. This optimization scheme significantly improves the hit rate and error tolerance for flexible antibody binding sites. Furthermore, the coarsening of the antigen and antibody residue side chains optimizes computation while preserving sufficient geometric structural features, providing more flexible adjustment space for the flexible region.

[0101] The various embodiments of this disclosure have the following beneficial effects: the accuracy of the determined antibody conformation can be improved by using the antigen-antibody binding conformation optimization method based on iterative nearest point according to some embodiments of this disclosure. Specifically, the antigen-antibody binding conformation optimization method based on iterative nearest point according to some embodiments of this disclosure firstly coarsely processes the antigen structure file and antibody structure file to obtain coarsely processed antibody file and coarsely processed antigen file. Here, coarse processing simplifies the antigen and antibody structures (retaining the CA (i.e., α-carbon atom) atoms simulated in the side chain center), reducing computational complexity while accurately preserving the core spatial information related to antigen-antibody binding, providing an efficient computational basis for subsequent CDR region and epitope matching. Then, based on the coarsely processed antibody file and coarsely processed antigen file, surface residues are extracted from the antigen structure file and antibody structure file to obtain an antigen surface residue information set and an antibody surface residue information set. Here, by extracting surface residues, interference from irrelevant regions can be reduced, improving the targeting of subsequent docking. Afterwards, an antibody binding region residue list is generated based on the antibody surface residue information set. Here, by generating a list of antibody-binding region residues, the core antibody-binding region can be precisely located. Then, based on this list, an epitope model is generated. This epitope model simulates the spatial characteristics of the original epitope, compensating for the weak complementarity between the CDR region and the antigen, providing a more forgiving matching space for the flexible CDR region, and improving fault tolerance. Next, spatial matching is performed between the antigen surface residue information set and the epitope model to obtain a matched epitope model. Spatial matching narrows the subsequent search range, avoiding the inefficiency of random sampling across the entire space in conventional methods, and improving docking efficiency. Then, based on the matched epitope model, an antibody docking conformation set is generated. This generated docking conformation aligns the coarse-grained antibody with the epitope conformation, allowing the antibody to initially align with the antigen-binding site, solving the problem of low overall antibody posture matching with antigen in conventional methods. Furthermore, using the coarse-grained antigen file, each antibody docking conformation in the antibody docking conformation set is locally optimized to obtain an optimized antibody conformation set. Here, through local optimization, the flexible characteristics of the CDR region can be adapted, avoiding the low fault tolerance problem caused by rigid complementarity, and significantly improving the geometric matching accuracy of the binding interface. Finally, based on the above coarse-grained antigen file, each optimized antibody conformation in the optimized antibody conformation set is screened to obtain the target antigen-antibody binding conformation. Here, through final screening, the hit rate of near-native structures can be improved, overcoming the defect of conventional scoring functions that ignore antigen-antibody specific interactions.

[0102] Further reference Figure 3As an implementation of the methods shown in the above figures, this disclosure provides some embodiments of an antigen-antibody binding conformation optimization device based on iterative nearest point. These device embodiments are similar to... Figure 1 Corresponding to the method embodiments shown, this antigen-antibody binding conformation optimization device based on iterative nearest point can be specifically applied to various electronic devices.

[0103] like Figure 3 As shown, an antigen-antibody binding conformation optimization device 300 based on iterative nearest point in some embodiments includes: a coarse-graining processing unit 301, an extraction unit 302, a first generation unit 303, a second generation unit 304, a spatial matching unit 305, a third generation unit 306, a local optimization unit 307, and a screening unit 308. The coarse-graining processing unit 301 is configured to coarse-grain the antigen structure file and the antibody structure file to obtain a coarse-grained antibody file and a coarse-grained antigen file; the extraction unit 302 is configured to extract surface residues from the antigen structure file and the antibody structure file based on the coarse-grained antibody file and the coarse-grained antigen file to obtain an antigen surface residue information set and an antibody surface residue information set; the first generation unit 303 is configured to generate an antibody binding region residue list based on the antibody surface residue information set; and the second generation unit 304 is configured to generate an epitope model based on the antibody binding region residue list. Spatial matching unit 305 is configured to perform spatial matching between the antigen surface residue information set and the epitope model to obtain a matched epitope model; third generation unit 306 is configured to generate an antibody docking conformation set based on the matched epitope model; local optimization unit 307 is configured to use the coarse-grained antigen file to locally optimize each antibody docking conformation in the antibody docking conformation set to obtain an optimized antibody conformation set; screening unit 308 is configured to screen each optimized antibody conformation in the optimized antibody conformation set based on the coarse-grained antigen file to obtain a target antigen antibody binding conformation.

[0104] It is understandable that the units described in this iterative nearest point-based antigen-antibody binding conformation optimization device 300 are related to the reference... Figure 1 The steps in the described method correspond accordingly. Therefore, the operations, features, and beneficial effects described above for the method also apply to the antigen-antibody binding conformation optimization device 300 based on iterative nearest point and the units contained therein, and will not be repeated here.

[0105] The following is for reference. Figure 4 It illustrates a schematic diagram of the structure of an electronic device (such as a computing device) suitable for implementing some embodiments of the present disclosure. Figure 4The electronic device shown is merely an example and should not be construed as limiting the functionality or scope of the embodiments of this disclosure. Figure 4 As shown, the computer device includes a processor, memory, and a network interface connected via a system bus. The memory may include a non-volatile storage medium and internal memory. The non-volatile storage medium may store an operating system and a computer program. The computer program includes program instructions that, when executed, cause the processor to perform any of the methods described above. The processor provides computational and control capabilities to support the operation of the entire computer device. The internal memory provides an environment for the execution of the computer program in the non-volatile storage medium; when executed by the processor, the computer program causes the processor to perform any of the methods described above. The network interface is used for network communication, such as sending assigned tasks. Those skilled in the art will understand that... Figure 4 The structure shown is merely a block diagram of a portion of the structure related to the present disclosure and does not constitute a limitation on the computer device to which the present disclosure is applied. A specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0106] It should be understood that the processor can be a Central Processing Unit (CPU), but it can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among these, a general-purpose processor can be a microprocessor or any conventional processor.

[0107] In one embodiment, the processor is configured to run a computer program stored in a memory to perform the following steps: coarse-graining processing of an antigen structure file and an antibody structure file to obtain a coarse-grained antibody file and a coarse-grained antigen file; extracting surface residues from the antigen structure file and the antibody structure file based on the coarse-grained antibody file and the coarse-grained antigen file to obtain an antigen surface residue information set and an antibody surface residue information set; generating an antibody binding region residue list based on the antibody binding region residue list; generating an epitope model based on the antibody binding region residue list; spatially matching the antigen surface residue information set and the epitope model to obtain a matched epitope model; generating an antibody docking conformation set based on the matched epitope model; locally optimizing each antibody docking conformation in the antibody docking conformation set using the coarse-grained antigen file to obtain an optimized antibody conformation set; and screening each optimized antibody conformation in the optimized antibody conformation set based on the coarse-grained antigen file to obtain a target antigen-antibody binding conformation.

[0108] This disclosure also provides a computer-readable storage medium storing a computer program, the computer program including program instructions, and the method implemented when the program instructions are executed can be referred to the various embodiments of the methods described above.

[0109] The aforementioned computer-readable storage medium may be an internal storage unit of the computer device described in the foregoing embodiments, such as the hard disk or memory of the computer device. Alternatively, the aforementioned computer-readable storage medium may be an external storage device of the computer device, such as a plug-in hard disk, SmartMedia Card (SMC), Secure Digital (SD) card, or Flash Card equipped on the computer device.

[0110] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or system that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or system. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or system that includes that element.

[0111] The above description is merely a selection of preferred embodiments of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of the invention involved in the embodiments of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described inventive concept. For example, technical solutions formed by substituting the above-described features with (but not limited to) technical features with similar functions disclosed in the embodiments of this disclosure.

Claims

1. A method for optimizing antigen-antibody binding conformation based on iterative nearest point, characterized in that, include: The antigen structure file and antibody structure file are coarse-grained to obtain coarse-grained antibody file and coarse-grained antigen file; Based on the coarse-grained antibody file and the coarse-grained antigen file, surface residues are extracted from the antigen structure file and the antibody structure file to obtain an antigen surface residue information set and an antibody surface residue information set; Based on the antibody surface residue information set, a list of antibody binding region residues is generated; Based on the list of antibody-binding region residues, an epitope model is generated; Spatial matching is performed between the antigen surface residue information set and the epitope model to obtain the matched epitope model; Based on the matched epitope model, an antibody docking conformation set is generated; Using the coarse-grained antigen file, each antibody docking conformation in the antibody docking conformation set is locally optimized to obtain an optimized antibody conformation set. Based on the coarse-grained antigen file, each optimized antibody conformation in the optimized antibody conformation set is screened to obtain the target antigen antibody binding conformation; The step of generating an antibody docking conformation set based on the matched epitope model includes: According to the preset rotation angle group, the matched episodic model is rotated and translated as a whole to obtain a set of rotated conformations; Using the antigen surface residue information set as the alignment target, the set of rotational conformations is iteratively matched to the nearest point space to obtain the set of epitope models after iteration. The iterative epitope models are sorted and selected from the set of iterative epitope models, and each epitope candidate point in the selected iterative epitope model is used as an epitope positioning point group to obtain a set of epitope positioning point groups. The epitope model is aligned with each epitope localization point group in the epitope localization point group set to obtain an aligned rotation matrix and an aligned translation vector. Each aligned rotation matrix and each aligned translation vector are then applied to the coarse-grained antibody file to generate an antibody docking conformation, resulting in an antibody docking conformation set.

2. The method according to claim 1, characterized in that, The process of coarse-graining the antigen structure file and antibody structure file to obtain coarse-grained antibody file and coarse-grained antigen file includes: Extract the antigen side chain atomic coordinates and antibody side chain atomic coordinates from the antigen structure file and the antibody structure file to obtain the antigen side chain atomic coordinate sequence and the antibody side chain atomic coordinate sequence; Determine the centroid coordinates corresponding to the atomic coordinate sequence of the antigen side chain to obtain the antigen centroid coordinate group, and combine the individual antigen centroid coordinates in the antigen centroid coordinate group into a coarse-grained antigen file according to the antigen structure order in the antigen structure file. The centroid coordinates corresponding to the atomic coordinate sequences of the antibody side chains are determined to obtain the antibody centroid coordinate set. Then, according to the antibody structure order in the antibody structure file, the centroid coordinates of each antibody in the antibody centroid coordinate set are combined into a coarse-grained antibody file.

3. The method according to claim 2, characterized in that, The step of extracting surface residues from the antigen structure file and the antibody structure file based on the coarse-grained antibody file and the coarse-grained antigen file to obtain an antigen surface residue information set and an antibody surface residue information set includes: The solvent-accessible surface area of ​​each amino acid residue in the antibody structural file is determined to obtain the antibody solvent-soluble surface area sequence; The antibody solvent soluble surface area and corresponding residue chain information that meet the preset residue screening conditions in the antibody solvent soluble surface area sequence are determined as antibody surface residue information and applied to the coarse-grained antibody file to obtain the antibody surface residue information set. The solvent-accessible surface area of ​​each amino acid residue in the antigen structural file is determined to obtain the antigen solvent-soluble surface area sequence. The antigen solvent soluble surface area and corresponding residue chain information that meet the preset residue screening conditions in the antigen solvent soluble surface area sequence are determined as antigen surface residue information and applied to the coarse-grained antigen file to obtain an antigen surface residue information set.

4. The method according to claim 3, characterized in that, The antibody surface residue information in the antibody surface residue information set includes amino acid sequences, wherein generating the antibody binding region residue list based on the antibody surface residue information set includes: The amino acid sequences included in the antibody surface residue information set are standardized and numbered to obtain an amino acid numbering information sequence, and a complementarity-determining region residue numbering sequence group is marked in the amino acid numbering information sequence, wherein the amino acid numbering information includes a standardized heavy chain numbering sequence, a first numbering mapping sequence, a standardized light chain numbering sequence, and a second numbering mapping sequence; The solvent-soluble surface area corresponding to each complementary region residue number in the complementary region residue number sequence group is determined to obtain the residue surface area sequence group; According to the residue surface area sequence group, sort the residue numbers of each complementary determination region in each complementary determination region residue numbering sequence in the complementary determination region residue numbering sequence group to obtain the sorted residue numbering sequence group. Determine the structural curvature corresponding to each sorted residue number in the sorted residue number sequence group to obtain the structural curvature sequence group; Based on the structural curvature sequence set, a target number of complementary decision region (CDR) residue numbers are selected from each CDR residue number sequence in the complementary decision region (CDR) residue number sequence set, and the target number of CDR residue numbers are combined into an antibody binding region residue list.

5. The method according to claim 4, characterized in that, The step of generating an epitope model based on the list of residues in the antibody-binding region includes: Using the list of residues in the antibody binding region as a template, the corresponding spatial distribution range of residues is determined; The distribution range of candidate points is set according to the spatial distribution range of the residues and the preset number of candidate points; Within the distribution range of the candidate points, an initial candidate point coordinate sequence is generated, wherein the distance between adjacent initial candidate point coordinates in the initial candidate point coordinate sequence is greater than or equal to a preset distance threshold. Based on the spatial distribution range of the residues, the coordinates of each initial candidate point in the initial candidate point coordinate sequence are screened to obtain the epitope model.

6. The method according to claim 5, characterized in that, The step of spatially matching the antigen surface residue information set and the epitope model to obtain the matched epitope model includes: Determine the global centroid coordinates of each candidate epitope point in the epitope model; The centroid coordinates of the overall epitope candidate points and the information of each antigen surface residue in the antigen surface residue information set are aligned to obtain the matched epitope model.

7. The method according to claim 6, characterized in that, The step involves using the coarse-grained antigen file to locally optimize each antibody docking conformation in the antibody docking conformation set, resulting in an optimized antibody conformation set, including: From the coarse-grained antigen file and the epitope localization point group set, select multiple matching nearest point coordinates from each epitope localization point group to obtain the antigen nearest point coordinate sequence set. Based on the antigen nearest point coordinate sequence set and the antibody binding region residue list, the nearest point optimization is performed on each antibody docking conformation in the antibody docking conformation set to obtain the optimized antibody conformation set.

8. The method according to claim 7, characterized in that, The step of screening each optimized antibody conformation in the optimized antibody conformation set according to the coarse-grained antigen file to obtain the target antigen antibody binding conformation includes: Based on the coarse-grained antigen file, the interaction potential energy of each optimized antibody conformation in the optimized antibody conformation set is determined to obtain the interaction potential energy set. Based on the set of interaction potential energies, optimized antibody conformations that meet the conformation screening conditions are selected from the set of optimized antibody conformations and used as the target antigen-antibody binding conformations.

9. A device for optimizing antigen-antibody binding conformation based on iterative nearest point, characterized in that, include: The coarse-graining processing unit is configured to coarse-grain the antigen structure file and the antibody structure file to obtain coarse-grained antibody file and coarse-grained antigen file; The extraction unit is configured to extract surface residues from the antigen structure file and the antibody structure file based on the coarse-grained antibody file and the coarse-grained antigen file, to obtain an antigen surface residue information set and an antibody surface residue information set; The first generation unit is configured to generate a list of antibody-binding region residues based on the antibody surface residue information set; The second generation unit is configured to generate an epitope model based on the list of residues in the antibody-binding region. A spatial matching unit is configured to perform spatial matching between the antigen surface residue information set and the epitope model to obtain a matched epitope model. The third generation unit is configured to generate a set of antibody docking conformations based on the matched epitope model; The local optimization unit is configured to use the coarse-grained antigen file to locally optimize each antibody docking conformation in the antibody docking conformation set to obtain an optimized antibody conformation set. The screening unit is configured to screen each optimized antibody conformation in the optimized antibody conformation set according to the coarse-grained antigen file to obtain the target antigen antibody binding conformation. The step of generating an antibody docking conformation set based on the matched epitope model includes: According to the preset rotation angle group, the matched episodic model is rotated and translated as a whole to obtain a set of rotated conformations; Using the antigen surface residue information set as the alignment target, the set of rotational conformations is iteratively matched to the nearest point space to obtain the set of epitope models after iteration. The iterative epitope models are sorted and selected from the set of iterative epitope models, and each epitope candidate point in the selected iterative epitope model is used as an epitope positioning point group to obtain a set of epitope positioning point groups. The epitope model is aligned with each epitope localization point group in the epitope localization point group set to obtain an aligned rotation matrix and an aligned translation vector. Each aligned rotation matrix and each aligned translation vector are then applied to the coarse-grained antibody file to generate an antibody docking conformation, resulting in an antibody docking conformation set.

Citation Information

Patent Citations

  • Deep learning model-based antibody structure optimization method and device

    CN116741260A

  • Screening method and device for protein-protein binding interface

    CN117577166A