Protein molecule fingerprint calculation method based on geometric model and application thereof

By constructing a protein molecular fingerprint based on a geometric model and combining it with machine learning tools, the problems of low accuracy in antibody epitope prediction and screening were solved, achieving efficient antigen epitope prediction and antibody screening, and improving the specificity of immune responses and the effectiveness of infectious disease prevention and control.

CN115482878BActive Publication Date: 2025-11-21FUDAN UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202211177489.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-26
Publication Date
2025-11-21
Estimated Expiration
2042-09-26

AI Technical Summary

Technical Problem

Existing antibody epitope prediction methods lack consideration of the three-dimensional interaction between antigens and antibodies, resulting in low prediction accuracy and difficulty in targeting specific epitopes during antibody screening, leading to unsatisfactory neutralization and blocking effects.

Method used

We construct protein molecular fingerprints based on geometric models, describe the antigen-antibody complex recognition interface on both sides, and combine machine learning tools to achieve antigen epitope prediction and high-throughput virtual antibody screening.

Benefits of technology

It improves the accuracy of antigen epitope prediction, enables specific antibody screening, enhances the neutralization blocking effect, and is applicable to immunomolecular design and infectious disease prevention and control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0003865230220000057
    Figure BDA0003865230220000057
  • Figure BDA0003865230220000111
    Figure BDA0003865230220000111
  • Figure BDA0003865230220000181
    Figure BDA0003865230220000181
Patent Text Reader

Abstract

The application provides a set of methods for constructing antigen-antibody complex mutual recognition interface descriptors (i.e., protein molecular fingerprints), which maximally present the structure and physicochemical characteristics of specific recognition by describing the binding interface of the three-dimensional structure of the antibody-antigen from both sides. Based on the protein molecular fingerprints generated by the set of methods, based on the specific interaction recognition rules of the antigen-antibody, the epitope prediction algorithm specific to the antibody and the virtual screening model of the antibody are designed, and the existing machine learning or deep learning tools are docked, so that the antigen epitope prediction based on the antibody and the high-throughput virtual screening of the antibody based on the specific epitope can be quickly realized.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to bioinformatics and immuno-informatics. Specifically, it relates to a method for constructing a set of descriptors for the interface of antigen-antibody complex based on the rules of antigen-antibody specific recognition, and a corresponding bioinformatics computational model. BACKGROUND

[0002] In the process of molecular recognition of immune system, immune cells specifically recognize a certain region on the surface of antigen molecule, which is called epitope or antigenic determinant. The antigen molecule binds to the specific antigen receptor on the surface of lymphocyte or free antibody through epitope, thereby triggering the corresponding immune response process. Therefore, epitope represents an immunologically active region on the antigen molecule, which determines the specificity of the antigen. Based on whether the epitope is continuous on the primary sequence of antigen protein, the epitope can be divided into linear epitope and conformational epitope; linear epitope refers to the antigen epitope composed of short peptides arranged linearly and continuously; while some amino acids, although not continuous on the primary sequence, fold into a specific conformation that gathers together in space, which is called conformational epitope. In cellular immunity, T lymphocytes usually do not directly recognize antigen molecules, but recognize antigen peptides presented by MHC molecules after being processed by APCs, which are T cell epitopes. In this process, the antigen peptides generated by APC processing need to be combined with the MHC molecules on the cell surface to form a peptide-MHC complex (pMHC), which can be recognized and combined by the TCR on the surface of the corresponding T lymphocyte, and then form a three-molecule complex composed of TCR, antigen peptide segment and MHC. Different types of T lymphocytes recognize antigen peptides presented by different MHC molecules. Among them, CD8+ T cells recognize 9-amino acid peptides presented by MHC class I molecules; CD4+ T cells recognize 13-17 amino acid peptides presented by MHC class II molecules. In humoral immunity, B lymphocytes recognize and bind to antigen epitope regions through surface BCR or antibodies. Correspondingly, the region of antibody molecule that interacts with antigen epitope is called antibody paratope. Since the interaction region between antigen and antibody is usually located in the CDR region of antibody, most of the amino acids of antibody paratope fall within the CDR region. Unlike T cell epitopes, which are linear peptide fragments, B cell epitopes exist in both linear and conformational forms, but about 90% of B cell epitopes are conformational epitopes. Antigen and antibody recognize and interact with each other through epitope and paratope, and one antibody corresponds to a specific epitope region on the antigen. However, for complex protein antigens, there are usually multiple B cell epitopes, some of which can play the role of immunodominant epitopes. The above factors all lead to the complexity and high challenge of B cell epitope recognition.

[0003] Accurate identification of B cell epitopes is essential for understanding the molecular basis of immune responses, and epitope determination also plays an important role in the development of immunogens, therapeutic antibodies and immunodiagnostic tools. For example, the determination of the epitope corresponding to a certain functional antibody can guide the design of a vaccine that can trigger the body to produce similar antibodies. Currently, various experimental methods have been successfully applied to the identification of epitopes, which can be roughly divided into two categories: structural epitope determination methods and functional epitope determination methods. In recent years, with the development of proteomics and the accumulation of antigen-antibody complex data, more and more B cell conformational epitope prediction algorithms have been developed. Most of the existing prediction algorithms are based on the existing epitope data to induce the characteristics that can distinguish epitopes and non-epitopes, and integrate these characteristics through a suitable machine learning model to build an epitope prediction model. The known epitope prediction methods are summarized as immunogenicity prediction algorithms, because the essence of these algorithms is to scan all the surface amino acids of the antigen protein and predict the amino acids that have the potential to become epitopes. With the development of antibody discovery technology, it is relatively easy to obtain functional antibodies corresponding to the target antigen, but the specific epitope region targeted by the antibody is usually unknown. In order to better reflect this biological reality and further improve the performance of epitope prediction methods, several B cell epitope prediction algorithms based on antibody information have appeared in recent years, which generally use the interaction preference of antigen-antibody amino acid pairs for design, such as ASEP, Bepar, ABepar and PEASE, and representative methods include the SEPPA series, the Discotope series, CBTOPE and Epitopia, etc. Since an antigen protein will display multiple epitope regions and be recognized by multiple different antibodies, and the above methods lack the information of the corresponding antibodies, the predicted regions are all immunogenic regions, not the specific epitope region of a certain antibody, so the above methods have not been widely used in the field of antibody-specific epitope prediction. In addition, in practical applications, researchers often have access to antibodies. However, due to the complex three-dimensional spatial structure of antibody-antigen mutual recognition, it is difficult to know the epitope region targeted by the antibody. The epitope prediction algorithm based on antibody sequence does not consider the protein interaction rules in three-dimensional space, so the overall prediction accuracy is low.

[0004] On the other hand, antibody screening is currently based on antibody libraries, using antigens to fish for potential antibodies. However, such obtained antibodies cannot guarantee that they target specific epitopes, and even if high-affinity antibodies are obtained, the later neutralization blocking effect is still not ideal. There is currently no computational technology for large-scale virtual screening of antibodies targeting specific epitopes.

[0005] Antigen-antibody specific recognition and interaction is the basic biochemical reaction of humoral immunity. In response to thousands of antigenic substances and constantly evolving viral antigens in nature, antibody molecules (or BCRs) in the body specifically recognize and interact with corresponding antigens to trigger subsequent immune responses. Infectious disease pathogens such as the new coronavirus, influenza virus, and HIV have always been a major threat to human health. A large number of antibodies can be isolated from the serum of infected animals / humans, but not all antibodies have neutralizing effects. Only monoclonal antibodies targeting specific antigen epitopes can block infection. In this process, it is difficult to determine the isolated antibodies and target antigen epitopes, and it is also a major challenge to screen target antibodies in the antibody library.

[0006] With the development of antibody omics technology, a large number of antibody sequences can be quickly obtained for a specific antigen protein. However, the antigen surface often expresses multiple dominant epitopes, and each dominant epitope can be recognized by multiple antibodies; on the other hand, different antibodies can recognize and bind to different epitope regions of the same antigen protein. The present application designs a set of methods for constructing protein molecular fingerprints of antigen-antibody complex mutual recognition interface (i.e., protein molecular fingerprints) based on the specific interaction characteristics of antigen-antibody, which has important significance in immune molecular design and modification and infectious disease prevention and control. SUMMARY

[0007] The present application provides a set of methods for constructing protein molecular fingerprints of antigen-antibody complex mutual recognition interface (i.e., protein molecular fingerprints), which presents the structure and physicochemical characteristics of specific recognition to the greatest extent by describing the binding interface of antibody-antigen three-dimensional structure from both sides. The protein molecular fingerprints generated based on this set of methods can quickly realize antigen epitope prediction based on antibodies and high-throughput antibody virtual screening based on specific epitopes by docking existing machine learning or deep learning tools.

[0008] In a first aspect, the present application provides a molecular fingerprint construction method of protein recognition interface based on geometric model, comprising the following steps:

[0009] S01 determining the geometric center point of the protein molecule and the recognition interface

[0010] The geometric center point of the protein molecule is obtained by taking the average of the spatial coordinates of all amino acids of the protein molecule.

[0011] The geometric center point of the recognition interface is obtained by taking the average of the spatial coordinates of the specific amino acids of the recognition interface.

[0012] S02 determining the description method of the geometric model of the protein molecule

[0013] According to the protein spatial conformation, different geometric models are selected for description, which can be selected from: sphere, cylinder, ellipsoid, cube, cone, or other geometric models. Further, the geometric models can describe structures including but not limited to antigen proteins, antigen epitope regions, antibody CDR regions, and antigen or antibody protein surface patches.

[0014] Further, when selected from the sphere model, the sphere model can describe structures including but not limited to antigen proteins, antigen epitope regions, and antibody CDR regions.

[0015] Further, when selected from the cylinder model, the cylinder model can describe structures including but not limited to structural features of antigen or antibody protein surface patches.

[0016] S03 Protein molecule geometric model sub-region division

[0017] According to the selected geometric model, it is divided into several sub-regions, and the principle of sub-region division is to ensure that each sub-region covers a reasonable number of amino acids. It is necessary to combine the size of different proteins, the size of amino acid molecules, and the spatial distance to determine the size range of the sub-region.

[0018] The effect achieved by the division is to ensure that each sub-region contains a reasonable number of amino acids, for example, each sub-region contains 0-10 amino acids, further 0-5 amino acids, and further 0-3 amino acids.

[0019] For the sphere model, the division methods include but are not limited to: taking a suitable step length from the center point to divide a multi-layer hollow sphere (sphere shell) sub-region; or taking a suitable step length from the center point to divide the equal angle of the sphere center, which divides the entire sphere into multiple equal spherical angle body sub-regions.

[0020] For the cylinder model, the division methods include but are not limited to: determining the starting point and the center of the surface patch, constructing a cylinder according to it, taking a suitable step length to cut the cylinder into equal-height sub-cylinders along the longitudinal direction, and dividing each sub-cylinder into multiple concentric circular ring sub-regions along the radial direction; or taking a suitable step length to cut the cylinder into equal-height sub-cylinders along the longitudinal direction, and dividing each sub-cylinder into multiple equal-division sector sub-regions.

[0021] The step length refers to the distance between the centers of adjacent sub-regions.

[0022] In one embodiment, the sphere model is selected for sub-region division; the sphere model sub-region division method can adopt the multi-layer sphere shell sub-region division method, and each sub-region is described by the following formula:

[0023]

[0024] C - the coordinate of the center point of the structure to be described;

[0025] R - the size of the sphere (the present application takes as an example);

[0026] S - the step size (the present application takes as an example);

[0027] i∈[1,R / S],j - the number of amino acids in the structure to be described,

[0028] Res - each amino acid in the structure to be described.

[0029] In another embodiment, a cylindrical model is selected to divide the molecular region;

[0030] The amino acids on the protein are divided into surface amino acids and internal amino acids, and the antigen-antibody interaction often occurs in the surface amino acid region. The surface amino acids can be determined according to the solvent accessibility of the amino acids. In the present application, the surface amino acid is defined as the amino acid with a surface solvent accessible area (SASA) greater than , and the surface solvent accessible area is calculated by the program Naccess V2.1.1.

[0031] For each surface amino acid Resr, all neighbor amino acids Resi with a distance less than from it are extracted as its corresponding surface patch Patch r .

[0032] Res i ∈Patch r ,if D min (Res i ,Res r )≤10 (2)

[0033] D min (Res i ,Res r ) - the minimum atomic distance between any amino acid Res i on the antigen protein and a specific surface amino acid Res r .

[0034] The spatial surface patch involved in the present application can define a set of amino acids, that is, a sphere with a radius of d generated by taking any surface amino acid r as the center, all amino acids on the protein surface amino acid r n set R, that is:

[0035] r n ∈R,if D(r,r n)≤d (3)

[0036] Any surface amino acid r n If the distance to r is less than d, it belongs to the surface patch set R.

[0037] For each surface patch on the antigen or antibody protein, a cylindrical structure is constructed based on its geometric center, and the cylinder has the geometric center C ag of the antigen protein and the geometric center C patch of the surface patch as the axis, with a width of and a height of The midpoint of the rotation axis segment is C patch , and the two endpoints of the rotation axis segment are C rat and C rab The surface is rotated to form a cylinder. The radius of the cylinder is and the height is At the same time, the cylinder is divided into N sub-regions according to a radius step of and a height step of , where:

[0038] N = R / S R × H / S H (3)

[0039] Where N represents the number of sub-regions, R represents the radius of the cylinder, H represents the height of the cylinder, S R represents the radial step of the cylinder, and S H represents the longitudinal step of the cylinder.

[0040] For any amino acid Res j in the cylindrical model space, its spatial coordinates are C j , and a line segment is drawn from point C j perpendicular to line segment C rat C rab , with foot coordinates C jdf . The sub-regions of the cylindrical model are described as follows:

[0041]

[0042] Where Area represents the overall cylindrical model area, composed of N sub-regions Area x,y , x is the vertical dimension number of sub-region Area x,y , and y is the horizontal dimension number of sub-region Area x,y . Each sub-region is composed of a set of amino acids [Res1, Res2, …, Res j] that meet the relevant conditions.

[0043] Selection of S04 protein molecule amino acid physicochemical property index

[0044] The physicochemical properties of amino acids include: acidity, basicity, number of hydrogen bond donors, hydrophobic index, van der Waals force, alpha-helix, beta-turn, and amino acid composition; and can also include other physicochemical properties with a correlation coefficient of 0.8 or more.

[0045] The physicochemical property index of each amino acid determines the amino acid factor, which is selected from an amino acid factor database (for example, the AAindex database). For example

[0046] Acidity: ZIMJ680104; basicity: ZIMJ680104; number of hydrogen bond donors: FAUJ880109; hydrophobic index: ARGP820101; van der Waals force: FAUJ880103; alpha-helix: GEIM800101; beta-turn: CHOP780101; amino acid composition: DAYM780101.

[0047] S05 Constructing the molecular fingerprint of the protein recognition interface

[0048] Based on the selected amino acid physicochemical property index, the characteristic value of each amino acid physicochemical property of each sub-region is calculated, and all the characteristic values of all the amino acids of all the sub-regions constitute the molecular fingerprint of the protein recognition interface.

[0049] Further, the fingerprint of the protein molecule is based on the sub-region Area formed by the geometric model, and the amino acid physicochemical property information AAindexData, the amino acid physicochemical property characteristic value of each sub-region is calculated, and all the characteristic values of all the sub-regions are combined to form the molecular fingerprint FP of the protein recognition interface.

[0050] Further, the calculation of the characteristic value of each amino acid physicochemical property of each sub-region refers to the arithmetic operation of the numerical values of the same type of physicochemical property index of all amino acids in the sub-region, and finally the characteristic value of the amino acid physicochemical property of the sub-region is obtained. The arithmetic operation can be addition operation, multiplication operation, arithmetic mean operation or geometric mean operation, etc. Preferably, the arithmetic operation is addition operation.

[0051] As a preferred example,

[0052] FP = {[AaindexData1Area1, …, AaindexData1Area n , …, AaindexDatamArean]}

[0053] As a preferred example,

[0054] FP={ [AaindexData1Area11, …, AaindexData1AreaXnYn, …, AaindexDatamAreaXnYn]}

[0055] FP - molecular fingerprint of the protein recognition interface to be described;

[0056] AaindexData - the numerical value of the selected amino acid physicochemical property index, if the sub-region contains multiple amino acids, the sum of the numerical values of all amino acid property indexes is taken;

[0057] Area - each sub-region of the structure to be described after division.

[0058] The molecular fingerprint dimension D of the final generated protein recognition interface m = M x N, where M is the number of amino acid property indexes, and N is the number of sub-regions of the geometric model.

[0059] In a second aspect, the present application provides a method for predicting an antigen epitope based on a corresponding antibody CDR sequence, the method comprising the following steps:

[0060] S01 Construction of antigen-antibody complex data set

[0061] Construction of a paired antigen-antibody complex data set with non-redundant epitopes; the antigen-antibody complex data set is a data set with 100% similarity for antigen epitope de-redundancy.

[0062] S02 Extraction of the recognition interface of the antigen-antibody complex data set

[0063] From the antigen-antibody complex structure data set, extract the amino acids with a minimum atomic distance between the antigen side and the antibody side within to constitute the recognition interface.

[0064] S03 Construction of an antigen epitope prediction model based on antibody CDR sequence, the construction method comprising the following steps:

[0065] 1. Around each surface amino acid on the antigen protein in the antigen-antibody complex data set, a surface patch is generated;

[0066] 2. Select a geometric model such as a cylindrical model to construct a molecular fingerprint of the protein recognition interface for each surface patch of the antigen and the antibody CDR structure, respectively.

[0067] 3. Training a classification model based on a machine learning algorithm

[0068] The molecular fingerprints of the antigen surface patches and the CDR structures of the antibodies are combined as training data, and a classification model for predicting antigen epitopes is trained by a machine learning algorithm. During the training of the classification model, the surface patches in the training set data that contain more than 30% of the amino acids of the real epitopes are set as positive samples, and the surface patches that do not contain the amino acids of the real epitopes are set as negative samples.

[0069] Preferably, during the training of the classification model, the surface patches in the training set data that contain more than 50% of the amino acids of the real epitopes are set as positive samples.

[0070] Further, the machine learning algorithm can be selected from: extreme gradient boosting algorithm (XGBoost, XGB), support vector machine algorithm (support vector machine, SVM), random forest algorithm (random forrest, RF), decision tree algorithm (decision tree, DT), multi-layer perception algorithm (multi-layer perception, MLP), gradient descent algorithm (gradient descent, GD), Gaussian Bayes algorithm (gaussian bayes, GNB) and linear regression algorithm (linear regression, LR).

[0071] Preferably, the extreme gradient boosting algorithm (XGBoost, XGB) is used.

[0072] S04 using the classification model to obtain the predicted antigen epitopes

[0073] The protein structure information of the antigen to be predicted and the CDR structure information of the corresponding antibody are input into the model obtained in the previous step to obtain the predicted antigen epitope information.

[0074] In a third aspect, the present application provides an antibody virtual screening method based on antigen epitopes. The screening method is based on the specific interaction characteristics of antigen-antibody, and constructs an antibody rapid screening algorithm ISEAP for the specific conformational epitopes of the target antigen. The steps of the method are as follows:

[0075] S01 construction of antigen-antibody complex structure data set

[0076] The antigen-antibody complex structure data set is constructed, and the antigen-antibody complex structure is selected to have complete light and heavy chains and the primary sequence length of the antigen protein is greater than 50 amino acids.

[0077] Further, the selection of the data set is based on the case that when a group of complex antigens, antibodies, epitopes and counter amino acids are completely the same, only one structure data is retained.

[0078] Construction of S02 antibody library

[0079] The antibody library includes an antibody structure library and an antibody sequence library. Antibody structure information and antibody sequence information can be collected from various databases, such as the protein structure database PDB and the sequence information database NCBI. After collection, redundant sequences with a CDR similarity of 100% are removed.

[0080] The antibody library can also be randomly and virtually constructed. Known antibodies are used as initial templates, and the known antibodies are constructed by randomly introducing different numbers of mutations in the CDR region. For example, 1-12, or 1-15, or 1-10 random mutations are introduced into the CDR of a certain known antibody, respectively. Each type of mutation level generates a certain number, such as 1000, of random mutation sequences, and an analog antibody library is constructed.

[0081] S03 extraction of recognition interface from antigen-antibody complex data set

[0082] From the antigen-antibody complex structure data set, the amino acids with a minimum atomic distance between the antigen side and the antibody side of within the recognition interface are extracted.

[0083] S04 Construction of an antibody virtual screening model based on an antigen epitope, which includes the following steps:

[0084] 1. Select a geometric model such as a spherical model to construct a molecular fingerprint of the protein recognition interface for the antigen epitope and the antibody CDR region in the antigen-antibody complex structure data set.

[0085] 2. Model training based on machine learning algorithm

[0086] After combining the molecular fingerprints of the antigen epitope and the antibody CDR structure, the combination is used as training data to train a prediction model of the antigen-antibody binding probability by a machine learning algorithm. During the model training process, the known real binding antigen-antibody pairs in the training set data are set as positive samples, and the antibodies known to bind to irrelevant antigens are set as negative samples.

[0087] Further, the machine learning algorithm can be selected from: extreme gradient boosting algorithm (XGBoost, XGB), support vector machine algorithm (support vector machine, SVM), random forest algorithm (random forrest, RF), LightGBM algorithm (LGB), decision tree algorithm (decision tree, DT), and linear regression algorithm (linear regression, LR).

[0088] Preferably, the random forest algorithm (random forrest, RF) is used.

[0089] S05 Virtual screening of antibodies using the trained model

[0090] Input the known epitope, use the model trained in S04 to carry out virtual screening in the antibody library, and obtain potential binding antibodies. BRIEF DESCRIPTION OF DRAWINGS

[0091] Figure 1 . Workflow diagram of SEPPA-mAb. (a) Input files required by SEPPA-mAb: antigen structure and antibody structure: (b) SEPPA-mAb generates surface patches for each amino acid of the antigen protein and antibody CDR after receiving the input files: (c) SEPPA-mAb generates fingerprint descriptors for surface patches based on physicochemical and structural features, and combines YGBoost classifier to predict scores for each surface patch: (d) represents the initial epitope score of the amino acid by considering the scores of all surface patches involved by the specific amino acid: (e) represents the final epitope score of the amino acid by considering the antigenicity trend of the neighbor amino acid.

[0092] Figure 2 . 5-fold cross-validation results of SEPPA-mAb on 8 machine learning algorithms.

[0093] Figure 3 . ROC of SEPPA-mAb on the independent test set

[0094] Figure 4 . Comparison of the prediction performance of SEPPA-mAb with other similar algorithms. (a) represents the ROC curve of different algorithms on the independent test set, the X axis is the false positive rate, and the Y axis is the true positive rate: (b) represents the balanced accuracy of different epitope prediction algorithms on the independent test set, the X axis is different epitope prediction algorithms, and the Y axis is the balanced accuracy value.

[0095] Figure 5 . Epitope prediction results of SEPPA-mAb and other algorithms for HIV envelope glycoprotein. (a) The red spherical region represents the real epitope of gp120 glycoprotein (PDBID: 6cm3 Chain E); (b-i) The spherical regions marked in the figure correspond to the predicted epitopes of SEPPA-mAb, SEPPA 3.0, Bepipred 2.0, CBTOPE, Discotope 2.0, Epitopia, ZDOCK and ClusPro for gp120 glycoprotein (PDBID: 6cm3 Chain E), respectively, wherein the predicted scores of the epitope amino acids in (b) from high to low are red, salmon and white, respectively, and the scores of the epitope amino acids in (c) from high to low are red, salmon, pink, white and gray, respectively.

[0096] Figure 6 Workflow of ISAEP algorithm. (a) represents the construction of a spherical shell model for the epitope region of the target antigen; (b) represents the construction of a spherical shell model for the CDR region of any antibody to be screened in the antibody library; (c) generates 80-bit fingerprint descriptors for the antigen epitope and antibody CDR respectively based on physicochemical and structural properties; (d) predicts the binding probability of the target antigen epitope to different antibody CDRs and ranks the antibodies based on the fingerprint descriptors.

[0097] Figure 7 ROC curves and AUC values of ISAEP model and human-mouse sub-model in independent validation set.

[0098] Figure 8 Performance of ISAEP in antibody structure library. (a) scatter plot represents the specific ranking position of the true binding antibodies in the antibody structure library. (b) the proportion of antigen epitopes in each ranking interval in figure a. (c) the number and proportion of antigen epitopes ranked in the top 5, 10, 15, 20, 25 and 29 (top 1%) by the true binding antibodies in the antibody structure library. X-axis represents different ranking ranges, left Y-axis represents the number of antigen epitopes, and right Y-axis represents the proportion of antigen epitopes.

[0099] Figure 9 Distribution of antibody CDR length. Histograms represent the CDR length distribution of heavy chain (a) and light chain (b) respectively, X-axis represents the number of CDR amino acids, Y-axis represents the number of antibodies, and the red line divides the CDR into three length intervals.

[0100] Figure 10 Practicality evaluation of ISAEP for antibody sequences. (a-c) ROC curves represent the prediction performance of ISAEP based on antibody sequences on all independent tests, human data, and mouse data respectively, X-axis represents specificity, and Y-axis represents sensitivity.

[0101] Figure 11 Performance of ISAEP for screening antibody structure library of three antigen proteins; boxplot represents the binding probability of the target antigen to the corresponding antibodies of different antigen similarity antigen proteins.

[0102] Figure 12 Performance of ISAEP for screening antibody sequence library of three antigen proteins. X-axis is the CDR similarity of antibodies, and Y-axis is the interval of the prediction score of ISAEP.

[0103] Figure 13 Prediction performance of ISAEP for screening mutant simulation antibody library of three antigen proteins. X-axis represents the number of mutations of the simulation antibodies, and Y-axis represents the binding probability of the target antigen to the simulation antibodies predicted by ISAEP. DETAILED DESCRIPTION

[0104] DEFINITIONS

[0105] For easier understanding of the present application, we set forth additional definitions of some terms throughout the application.

[0106] As used in this specification and claim, the words "comprise", "comprising", "inclusive" or "comprising" set forth in the context of possibilities. The words "comprise", "comprising", "inclusive" or "comprising" set forth in the context of possibilities.

[0107] The term "for example" should be used as an example, not as a limitation, and should not be understood as referring only to those items explicitly listed in the specification.

[0108] Throughout the specification, the word "comprise" or "comprise" will be understood to include the stated elements or steps, but not exclude any other elements or steps, which include the meaning of "consist of" and / or "consist essentially of".

[0109] The present application focuses on the specific recognition of antigen antibodies in humoral immunity. Unless otherwise specified, the epitopes mentioned in this paper refer to B cell epitopes.

[0110] The term "epitope" in the present application refers to the specific binding of antigen antibodies, which forms a binding interface. The set of amino acids on the antigen side of the binding interface is called the antigen epitope (epitope for short).

[0111] The term "counter" in the present application refers to the set of amino acids on the antibody side of the binding interface, which is called the antibody counter (counter for short).

[0112] The "recognition interface" in the present application is composed of "epitope" amino acids and "counter" amino acids.

[0113] The term "surface amino acid" used in the present application refers to an amino acid with a solvent accessible surface area (SASA) greater than 2, calculated by the program Naccess V2.1.1. 2.

[0114] The term "surface patch" used in the present application refers to extracting all neighboring surface amino acids Resi with a distance less than 5.0 A from each surface amino acid Resr as its corresponding surface patch Patchr.

[0115]

[0116] Dmin(Resi,Resr) - represents the minimum atomic distance between any amino acid Resi and a specific surface amino acid Resr on the antigen protein. ​

[0117] As used herein, the term "amino acid fingerprint descriptor" is used to describe the physico-chemical micro-environment and spatial arrangement of the CDRs of an antibody and the surface patches of an antigen. The fingerprint descriptors are constructed based on the structural and physico-chemical properties of the antigen-antibody specific binding, such as the structural composition, electrostatic interaction, hydrophobic interaction, hydrogen bond, and van der Waals force. The isoelectric point, hydrogen bond donor, hydrophobic force, van der Waals force, alpha-helix, beta-sheet, and amino acid composition are collected from the AAindex database. For the isoelectric point property, it is split into two features, acidic and basic. The selection of the AAindex scales follows the following principles: 1) if there is only one set of scale in the AAindex database, it is directly selected, for example, the three properties of isoelectric point, hydrogen bond donor, and van der Waals force. 2) for the properties with multiple sets of scales, the scale describing the general protein, rather than the specific sub-class of protein, is selected first. In one embodiment, for the alpha-helix, there are four related scales (GEIM800101-GEIM800104), among which GEIM800102-4 are the descriptors for alpha-protein or beta-protein, and GEIM800101 is the general descriptor without distinguishing the protein type, so GEIM800101 is selected. 3) among the multiple sets of scales describing the same property, the scale with the highest correlation to the other scales, i.e., the scale best representing the property, is selected. In another embodiment, for the hydrophobic force, there are five sets of scales, among which ARGP820101 has the highest correlation to the other four sets, so ARGP820101 is selected. In the present application, to eliminate the dimension effect, the values of the 20 amino acids corresponding to each amino acid factor are normalized for subsequent descriptor calculation. The values of the 20 amino acids corresponding to the eight amino acid factors after processing are shown in Table 1.

[0118] Table 1 Values of the 20 amino acids corresponding to the eight amino acid factors

[0119]

[0120] As used herein, the term "disulfide bond" refers to the two cysteine residues in the antibody that form a stable "Y" shaped structure of the immunoglobulin. The two cysteine residues form a disulfide bond that is essential for the stability of the antibody. In the hypervariable region (Fv) of the antibody, there are two intra-chain disulfide bonds (light chain: A23-A88; heavy chain: B22-B92).

[0121] The term "correction" used in this article refers to the correction step of SEPPA-mAb. For an amino acid with a low initial epitope score, if most of its neighbors have high epitope scores, the correction step will increase its epitope score, meaning its probability of becoming an epitope increases. Similarly, for an amino acid with a high initial epitope score, if its surrounding amino acids are all low-scoring, the correction step will decrease its probability of becoming an epitope.

[0122] The “interacting amino acid pair” described in this invention refers to the interacting amino acid pair between the antigen side and the antibody side at the binding interface of an antigen-antibody complex protein structure; wherein, the amino acid on the antigen side is denoted as the epitope amino acid, and the amino acid on the antibody side is denoted as the para amino acid.

[0123] The "epitope" mentioned in this invention refers to the epitope located at a distance of [missing information - likely a distance from the para amino acid]. The antigen-side amino acids within the set of antigenic amino acids; in its expression, for any amino acid agr on the set of antigenic amino acids i If any atom A j The minimum spatial distance to the corresponding antibody is less than This amino acid belongs to the epitope EpR.

[0124] The term "para" in this invention refers to the position of the amino acid at the para site being at a distance from the epitope amino acid. The expression for the antibody-side amino acids within the antibody amino acid set is as follows: for any amino acid abr i If any atom A_j satisfies the minimum spatial distance to the corresponding antigen is less than 1 / 2. This amino acid belongs to the para-PaR position.

[0125] The method for numbering the antibody CDR region in this invention can be IMGT, Kabat, Chothia, Martin, AHo method, or other known CDR region numbering methods.

[0126] The IMGT numbering method is as follows: CDR1 is 27-38, CDR2 is 56-65, and CDR3 is 105-117.

[0127] The Kabat numbering method is as follows: for antibody heavy chains, CDR1 is 31–35, CDR2 is 50–65, and CDR3 is 95–102; for antibody light chains, CDR1 is 24–34, CDR2 is 50–56, and CDR3 is 89–97.

[0128] Example 1. Constructing protein molecular fingerprints

[0129] This embodiment describes the construction of protein molecular fingerprints using antigen epitopes and antibody CDRs.

[0130] The data of this example is from Protein Data Bank (PDB) with ID 7q9g, which is the complex structure of SARS-CoV-2 Spike protein antigen and neutralizing antibody COVOX-222; wherein the Spike protein is A / B / C chain, the antibody COVOX-222 heavy chain is E / H chain, and the antibody COVOX-222 light chain is F / L chain; chain E, chain F, and chain B form an antigen-antibody complex, which is named 7q9g_E_F_B. Here, E and F represent the light chain and heavy chain of the antibody, and B represents the antigen chain.

[0131] 1. Antigen protein surface patch molecule fingerprint construction

[0132] To achieve epitope prediction, the surface amino acids of the antigen protein need to be divided into surface patches, and the molecular fingerprint of each surface patch needs to be calculated to describe it in detail.

[0133] First, the surface amino acids of the antigen protein need to be extracted. The antigen chain B of the complex 7q9g_E_F_B contains a total of 1060 amino acids, and the surface amino acids are calculated using Naccess V2.1.1, a total of 932 surface amino acids are obtained as follows:

[0134] {13 SER, 14 GLN, 15 CYS, 16 VAL, 17 ASN, 18 PHE, 19 THR, 20 THR, 21 ARG, 22 THR, 23 GLN, 24 LEU, 25 PRO, 26 PRO, 27 ALA, 28 TYR, 29 THR, 30 ASN, 32 PHE, 33 THR, 34 ARG, 38 TYR, 39 PRO, 40 ASP, 41 LYS, 42 VAL, 43 PHE, 44 ARG, 45 SER, 46 SER, 47 VAL, 48 LEU, 49 HIS, 50 SER, 51 THR, 52 GLN, 53 ASP, 54 LEU, 57 PRO, 58 PHE, 59 PHE, 60 SER, 61 ASN, 63 THR, 64 TRP, 65 PHE, 66 HIS, 68 ILE, 69 HIS, 70 VAL, 71 SER, 72 GLY, 75 GLY, 76 THR, 77 LYS, 78 ARG, 79 PHE, 80 ALA, 81 ASN, 82 PRO, 83 VAL, 84 LEU, 85 PRO, 86 PHE, 87 ASN, 88 ASP, 94 SER, 95 THR, 96 GLU, 97 LYS, 98 SER, 99 ASN, 100 ILE, 101 ILE, 102 ARG, 104 TRP, 105 ILE, 108 THR, 109 THR, 110 LEU, 111 ASP, 112 SER, 113 LYS, 114 THR, 115 GLN, 118 LEU, 119 ILE, 120 VAL, 121 ASN, 122 ASN, 123 ALA, 124 THR, 125 ASN, 126 VAL, 127 VAL, 128 ILE, 129 LYS, 132 GLU, 134 GLN, 135 PHE, 136 CYS, 137 ASN, 138 ASP, 139 PRO, 140 PHE, 141 LEU, 144 TYR, 145 TYR, 146 HIS, 147 LYS, 148 ASN, 149 ASN, 150 LYS, 151 SER, 152 TRP, 153 MET, 154 GLU, 155 SER, 156 GLU, 157 PHE, 158 ARG, 160 TYR, 161 SER, 162 SER, 163 ALA, 164 ASN, 165 ASN, 166 CYS, 167 THR, 168 PHE, 169 GLU, 170 TYR, 171 VAL, 172 SER,173_GLN, 174_PRO, 175_PHE, 176_LEU, 177_MET, 178_ASP, 179_LEU, 186_PHE, 187_LYS, 188_ASN, 190_ARG, 192_PHE, 194_PHE, 195_LYS, 196_ASN, 197_ILE, 198_ASP, 199_GLY, 200_TYR, 202_LYS, 203_ILE, 204_TYR, 205_SER, 206_LYS, 207_HIS, 208_THR, 209_PRO, 210_ILE, 211_ASN, 212_LEU, 213_VAL, 214_ARG, 215_GLY, 216_LEU, 217_PRO, 218_GLN, 219_GLY, 220_PHE, 221_SER, 222_ALA, 224_GLU, 225_PRO, 226_LEU, 227_VAL, 228_ASP, 229_LEU, 230_PRO, 231_ILE, 232_GLY, 233_ILE, 234_ASN, 235_ILE, 236_THR, 237_ARG, 239_GLN, 240_THR, 241_LEU, 242_HIS, 244_SER, 245_TYR, 246_LEU, 247_THR, 248_PRO, 249_GLY, 257_ALA, 258_GLY, 259_ALA, 261_ALA, 264_VAL, 265_GLY, 266_TYR, 267_LEU, 268_GLN, 269_PRO, 270_ARG, 271_THR, 275_LYS, 276_TYR, 277_ASN, 278_GLU, 279_ASN, 280_GLY, 281_THR, 282_ILE, 283_THR, 284_ASP, 285_ALA, 286_VAL, 287_ASP, 288_CYS, 289_ALA, 290_LEU, 291_ASP, 292_PRO, 293_LEU, 294_SER, 295_GLU, 297_LYS, 298_CYS, 299_THR, 300_LEU, 301_LYS, 302_SER, 303_PHE, 304_THR, 306_GLU, 307_LYS, 308_GLY, 309_ILE, 310_TYR, 311_GLN, 312_THR, 313_SER, 314_ASN, 315_PHE, 316_ARG, 317_VAL, 318_GLN, 319_PRO, 320_THR, 321_GLU, 322_SER, 323_ILE, 324_VAL, 325_ARG,326 PHE, 327 PRO, 328 ASN, 329 ILE, 330 THR, 331 ASN, 332 LEU, 333 CYS, 334 PRO, 335 PHE, 336 GLY, 337 GLU, 338 VAL, 339 PHE, 340 ASN, 341 ALA, 342 THR, 343 ARG, 344 PHE, 345 ALA, 346 SER, 348 TYR, 349 ALA, 350 TRP, 351 ASN, 352 ARG, 353 LYS, 354 ARG, 355 ILE, 356 SER, 357 ASN, 358 CYS, 359 VAL, 361 ASP, 363 SER, 364 VAL, 365 LEU, 366 TYR, 367 ASN, 368 SER, 369 ALA, 370 SER, 371 PHE, 372 SER, 373 THR, 374 PHE, 375 LYS, 376 CYS, 377 TYR, 378 GLY, 379 VAL, 380 SER, 381 PRO, 382 THR, 383 LYS, 385 ASN, 386 ASP, 387 LEU, 388 CYS, 389 PHE, 390 THR, 391 ASN, 393 TYR, 396 SER, 400 ARG, 401 GLY, 402 ASP, 403 GLU, 404 VAL, 405 ARG, 406 GLN, 407 ILE, 408 ALA, 409 PRO, 410 GLY, 411 GLN, 412 THR, 413 GLY, 414 ASN, 417 ASP, 418 TYR, 420 TYR, 421 LYS, 422 LEU, 423 PRO, 424 ASP, 425 ASP, 426 PHE, 427 THR, 430 VAL, 433 TRP, 434 ASN, 435 SER, 436 ASN, 437 ASN, 438 LEU, 439 ASP, 440 SER, 441 LYS, 442 VAL, 443 GLY, 444 GLY, 445 ASN, 446 TYR, 447 ASN, 448 TYR, 449 LEU, 450 TYR, 451 ARG, 452 LEU, 453 PHE, 454 ARG, 455 LYS, 456 SER, 457 ASN, 458 LEU, 459 LYS, 460 PRO, 461 PHE, 462 GLU, 463 ARG, 464 ASP, 465 ILE, 466 SER, 467 THR,468 GLU, 469 ILE, 470 TYR, 471 GLN, 472 ALA, 473 GLY, 474 SER, 475 THR, 476 PRO, 477 CYS, 478 ASN, 479 GLY, 480 VAL, 481 LYS, 482 GLY, 483 PHE, 484 ASN, 485 CYS, 486 TYR, 487 PHE, 488 PRO, 489 LEU, 490 GLN, 491 SER, 492 TYR, 493 GLY, 494 PHE, 495 GLN, 496 PRO, 497 THR, 498 TYR, 499 GLY, 500 VAL, 501 GLY, 502 TYR, 503 GLN, 505 TYR, 506 ARG, 511 SER, 512 PHE, 513 GLU, 514 LEU, 515 LEU, 516 HIS, 517 ALA, 518 PRO, 519 ALA, 520 THR, 522 CYS, 524 PRO, 525 LYS, 526 LYS, 527 SER, 528 THR, 529 ASN, 530 LEU, 531 VAL, 532 LYS, 533 ASN, 534 LYS, 535 CYS, 537 ASN, 538 PHE, 539 ASN, 540 PHE, 541 ASN, 542 GLY, 543 LEU, 544 THR, 545 GLY, 546 THR, 548 VAL, 550 THR, 551 GLU, 552 SER, 553 ASN, 554 LYS, 555 LYS, 556 PHE, 557 LEU, 558 PRO, 559 PHE, 560 GLN, 561 GLN, 562 PHE, 563 GLY, 564 ARG, 565 ASP, 566 ILE, 567 ALA, 568 ASP, 569 THR, 570 THR, 571 ASP, 574 ARG, 576 PRO, 577 GLN, 578 THR, 579 LEU, 580 GLU, 581 ILE, 582 LEU, 583 ASP, 584 ILE, 585 THR, 586 PRO, 587 CYS, 588 SER, 589 PHE, 590 GLY, 591 GLY, 592 VAL, 593 SER, 597 PRO, 598 GLY, 599 THR, 600 ASN, 601 THR, 602 SER, 603 ASN, 604 GLN, 605 VAL, 607 VAL, 608 LEU, 609 TYR,610 GLN, 61 1 GLY, 612 VAL, 613 ASN, 614 CYS, 615 THR, 616 GLU, 617 VAL, 618 PRO, 619 VAL, 638 ASN, 639 VAL, 640 PHE, 641 GLN, 642 THR, 643 ARG, 644 ALA, 645 GLY, 648 ILE, 649 GLY, 650 ALA, 651 GLU, 652 HIS, 653 VAL, 654 ASN, 655 ASN, 656 SER, 657 TYR, 658 GLU, 659 CYS, 660 ASP, 661 ILE, 662 PRO, 663 ILE, 664 GLY, 665 ALA, 666 GLY, 667 ILE, 668 CYS, 670 SER, 671 TYR, 672 GLN, 673 THR, 686 SER, 687 GLN, 688 SER, 690 ILE, 692 TYR, 693 THR, 694 MET, 695 SER, 696 LEU, 697 GLY, 698 VAL, 699 GLU, 700 ASN, 701 SER, 702 VAL, 703 ALA, 704 TYR, 705 SER, 706 ASN, 707 ASN, 708 SER, 709 ILE, 710 ALA, 711 ILE, 712 PRO, 713 THR, 714 ASN, 715 PHE, 716 THR, 717 ILE, 718 SER, 719 VAL, 720 THR, 721 THR, 722 GLU, 723 ILE, 724 LEU, 725 PRO, 726 VAL, 727 SER, 728 MET, 729 THR, 730 LYS, 731 THR, 732 SER, 733 VAL, 734 ASP, 735 CYS, 736 THR, 737 MET, 738 TYR, 740 CYS, 741 GLY, 742 ASP, 743 SER, 744 THR, 745 GLU, 747 SER, 748 ASN, 749 LEU, 750 LEU, 751 LEU, 752 GLN, 753 TYR, 754 GLY, 755 SER, 756 PHE, 757 CYS, 758 THR, 759 GLN, 760 LEU, 761 ASN, 762 ARG, 763 ALA, 764 LEU, 765 THR, 766 GLY, 767 ILE, 768 ALA, 769 VAL, 770 GLU, 771 GLN,772_ASP, 773_LYS, 774_ASN, 775_THR, 776_GLN, 777_GLU, 779_PHE, 780_ALA, 781_GLN, 782_VAL, 783_LYS, 784_GLN, 785_ILE, 786_TYR, 787_LYS, 788_THR, 789_PRO, 790_PRO, 791_ILE, 792_LYS, 793_ASP, 794_PHE, 795_GLY, 796_GLY, 797_PHE, 798_ASN, 799_PHE, 800_SER, 801_GLN, 802_ILE, 803_LEU, 804_PRO, 805_ASP, 806_PRO, 807_SER, 808_LYS, 809_PRO, 810_SER, 811_LYS, 812_ARG, 813_SER, 814_PHE, 816_GLU, 817_ASP, 818_LEU, 820_PHE, 821_ASN, 822_LYS, 823_VAL, 824_THR, 852_PHE, 853_ASN, 854_GLY, 855_LEU, 856_THR, 857_VAL, 858_LEU, 859_PRO, 860_PRO, 861_LEU, 862_LEU, 863_THR, 864_ASP, 865_GLU, 866_MET, 867_ILE, 869_GLN, 870_TYR, 872_SER, 875_LEU, 876_ALA, 880_THR, 881_SER, 882_GLY, 883_TRP, 884_THR, 885_PHE, 886_GLY, 887_ALA, 888_GLY, 889_ALA, 890_ALA, 891_LEU, 892_GLN, 893_ILE, 894_PRO, 895_PHE, 896_ALA, 897_MET, 898_GLN, 901_TYR, 904_ASN, 905_GLY, 906_ILE, 907_GLY, 909_THR, 910_GLN, 911_ASN, 912_VAL, 914_TYR, 915_GLU, 916_ASN, 917_GLN, 918_LYS, 919_LEU, 921_ALA, 922_ASN, 923_GLN, 925_ASN, 926_SER, 927_ALA, 928_ILE, 929_GLY, 930_LYS, 931_ILE, 932_GLN, 933_ASP, 934_SER, 935_LEU, 936_SER, 937_SER, 938_THR, 939_ALA, 940_SER, 941_ALA,942 LEU, 943 GLY, 944 LYS, 945 LEU, 946 GLN, 947 ASP, 948 VAL, 949 VAL, 950 ASN, 951 GLN, 952 ASN, 953 ALA, 954 GLN, 955 ALA, 956 LEU, 957 ASN, 958 THR, 959 LEU, 960 VAL, 961 LYS, 962 GLN, 963 LEU, 964 SER, 965 SER, 966 ASN, 967 PHE, 968 GLY, 969 ALA, 970 ILE, 971 SER, 972 SER, 973 VAL, 974 LEU, 975 ASN, 976 ASP, 978 LEU, 979 SER, 980 ARG, 981 LEU, 982 ASP, 983 PRO, 984 PRO, 985 GLU, 987 GLU, 988 VAL, 989 GLN, 991 ASP, 992 ARG, 995 THR, 996 GLY, 997 ARG, 998 LEU, 999 GLN, 1000 SER, 1002 GLN, 1003 THR, 1004 TYR, 1005 VAL, 1006 THR, 1007 GLN, 1008 GLN, 1009 LEU, 1010 ILE, 1011 ARG, 1012 ALA, 1013 ALA, 1014 GLU, 1015 ILE, 1016 ARG, 1017 ALA, 1018 SER, 1020 ASN, 1021 LEU, 1023 ALA, 1024 THR, 1025 LYS, 1027 SER, 1028 GLU, 1031 LEU, 1032 GLY, 1033 GLN, 1034 SER, 1035 LYS, 1036 ARG, 1037 VAL, 1038 ASP, 1039 PHE, 1040 CYS, 1041 GLY, 1042 LYS, 1043 GLY, 1044 TYR, 1045 HIS, 1052 SER, 1053 ALA, 1054 PRO, 1055 HIS, 1056 GLY, 1061 HIS, 1065 VAL, 1066 PRO, 1067 ALA, 1068 GLN, 1069 GLU, 1070 LYS, 1071 ASN, 1072 PHE, 1073 THR, 1074 THR, 1075 ALA, 1076 PRO, 1079 CYS, 1080 HIS, 1081 ASP, 1082 GLY, 1083 LYS, 1085 HIS,1086_PHE,1087_PRO,1088_ARG,1089_GLU,1090_GLY,1091_VAL,1094_SER,1095_ASN,1096_GLY,1097_THR,1098_HIS,1099_TRP,1100_PHE,1101_VAL, 1103_GLN,1104_ARG,1105_ASN,1106_PHE,1107_TYR,1108_GLU,1109_PRO,1110_GLN,1111_ILE,1112_ILE,1113_THR,1114_THR,1115_ASP,1116_ASN, 1117_THR,1118_PHE,1119_VAL,1120_SER,1121_GLY,1122_ASN,1123_CYS,1124_ASP,1125_VAL,1126_VAL,1127_ILE,1128_GLY,1129_ILE,1130_VAL, 1131_ASN,1132_ASN,1133_THR,1134_VAL,1135_TYR,1136_ASP,1137_PRO,1138_LEU,1139_GLN,1140_PRO,1141_GLU,1142_LEU,1143_ASP,1144_SER},

[0135] The following section uses the surface amino acid 455_LYS as an example to calculate the corresponding surface patch, then constructs a cylindrical model, and finally calculates the molecular fingerprint.

[0136] 1.1 Surface sheet generation. For surface amino acid 455_LYS, extract the amino acid at a distance of less than [missing information]. Each neighboring amino acid is used as its corresponding surface plate, thus obtaining its surface plate as follows:

[0137] {418_TYR,453_PHE,454_ARG,455_LYS,456_SER,457_ASN,468_GLU,470_TYR,471_GLN,472_ALA}

[0138] 1.2 Determine the geometric center point. The geometric center of the surface sheet is the spatial coordinate of surface amino acid 455_LYS (i.e., the geometric center point of the coordinates of all atoms of surface amino acid 455_LYS, X = 146.875, Y = 191.715, Z = 224.799); the geometric center of the antigen protein is the mean value of the spatial coordinates of all atoms of all amino acids on the antigen protein (PDB: 7q9g_B) (coordinates are X = 183.052, Y = 169.794, Z = 159.206).

[0139] 1.3 Determine the geometric model description method: describe the object as a surface patch corresponding to each surface amino acid on the antigen. In this embodiment, the surface amino acid 455_LYS is taken as an example, and the geometric model is selected as a cylindrical model.

[0140] 1.4 Model sub-region division: in the cylindrical model, the cylindrical axis passes through the geometric center point of the surface patch and the geometric center point of the antigen protein. In this embodiment, the radius of the cylinder is taken as and the height is The radius step is taken as and the axis step is The cylindrical model is divided into 25 (10 / 2*20 / 4) sub-regions.

[0141] 1.5 Selection of amino acid physicochemical property indexes of protein molecules. Eight representative amino acid physicochemical property indexes are selected. In this embodiment, the eight properties selected are:

[0142] {Acidic: ZIMJ680104; Basic: ZIMJ680104; Hydrogen bond donor number: FAUJ880109; Hydrophobic index: ARGP820101; Van der Waals force: FAUJ880103; Alpha-helix: GEIM800101; Beta-turn: CHOP780101; Amino acid composition: DAYM780101}

[0143] 1.6 Construction of antigen surface patch molecular fingerprint. Based on the 25 sub-regions divided in 1.4 and the 8 amino acid property indexes determined in 1.5, the 200-dimensional antigen surface patch molecular fingerprint is generated as:

[0144] {0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, -0.89, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.5, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.25, 0.0, 0.25, 0.5, 0.0, 0.0, 0.0, 0.0, 0.25, 1.0, 0.0, 0.25, 0.0, 0.0, 0.5, 0.0, 0.0, 0.0, 0.0, 0.0, 0.2302, 0.0, 0.0, 0.0, 0.0, 0.7094, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.7623, 0.0, 0.0, 0.0, 0.7094, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.4889, 0.0, 0.1238, 0.0, 0.0, 0.0, 0.0, 0.8007, 0.0, 0.4678, 0.5903, 0.0, 0.0, 0.0, 0.729, 0.198, 0.7587, 0.0, 0.8007, 0.0, 0.0, 0.3651, 0.0, 0.0, 0.0, 0.4835, 0.0, 0.7253, 0.0, 0.0, 0.0, 0.0, 0.0879, 0.0, 0.9451, 0.7692, 0.0, 0.0, 0.0, 0.5495, 0.1648, 0.4066, 0.0, 0.0879, 0.0, 0.0, 0.1978, 0.0, 0.0, 0.0, 0.4679, 0.0, 0.1743, 0.0, 0.0, 0.0, 0.0, 0.6147, 0.0, 0.2477, 0.4954, 0.0, 0.0, 0.0, 0.1193, 0.8807, 0.4404, 0.0, 0.6147, 0.0, 0.0, 1.0, 0.0, 0.0, 0.0, 0.3562, 0.0, 1.0, 0.0, 0.0, 0.0, 0.0, 0.2877, 0.0, 0.6438, 0.726, 0.0, 0.0, 0.0, 0.3151, 0.7808, 0.4932, 0.0, 0.2877, 0.0, 0.0, 0.411, 0.0, 0.0, 0.0}

[0145] 2. Antibody CDR protein molecule fingerprint construction

[0146] Firstly, the CDR regions of the antibody were determined: the determination of the antibody CDRs was based on the IMGT antibody numbering rule (CDR1 : 27-38, CDR2: 56-65, CDR3: 105-117). Thus, according to the amino acid sequence number in the PDB file, the CDRs of the heavy chain E chain and the light chain F chain were extracted respectively as shown in Table 2:

[0147] Table 2: List of CDR extraction of heavy chain E chain and light chain F chain

[0148]

[0149] 2.1 Determination of the geometric center point: the center point was determined as the average of the spatial coordinates of all atoms belonging to all amino acids in the CDR regions of the antibody light and heavy chains. The coordinates of the center point are (X = 161.380, Y = 194.867, Z = 233.203).

[0150] 2.2 Determination of the geometric model description method: the object to be described is the CDR region of the antibody. In this embodiment, the geometric model is a sphere model.

[0151] 2.3 Model sub-region division: divide the multi-layer hollow sphere (shell) sub-region from the center point with an appropriate step size. Referring to formula (1), the size (radius) of the sphere is taken as The step size is It is divided into 10 hollow spherical shell sub-regions, and the amino acids contained in each region are shown in Table 3:

[0152] Table 3: Amino acids in sub-regions of heavy chain E chain and light chain F chain

[0153]

[0154] 2.4 Selection of amino acid physicochemical property indices of protein molecules. Eight representative amino acid physicochemical property indices were selected. In this embodiment, the eight properties selected are:

[0155] {Acidic: ZIMJ680104; Basic: ZIMJ680104; Number of hydrogen bond donors: FAUJ880109; Hydrophobic index: ARGP820101; Van der Waals force: FAUJ880103; Alpha-helix: GEIM800101; Beta-turn: CHOP780101; Amino acid composition: DAYM780101}

[0156] 2.5 Construction of antibody CDR molecular fingerprint. Based on the 10 sub-regions divided in 2.3 and the 8 amino acid property indices determined in 2.4, an 80-dimensional antibody CDR molecular fingerprint is generated:

[0157] {0.0000, 0.0000, 0.0000, 0.0000, 1.0000, 1.2100, 0.0000, 0.0000, 0.0000, 0.0000, 0.0000, -0.8900, 0.0000, 0.0000, -1.0000, -1.0000, 0.0000, 0.0000, 0.0000, 0.0000, 0.0000, 1.0000, 0.2500, 0.5000, 2.5000, 2.5000, 1.0000, 0.7500, 0.2500, 0.7500, 0.0000, 0.0000, 1.7623, 2.8904, 1.6490, 1.5735, 0.0000, 2.5396, 0.0000, 0.0000, 0.0000, 1.0309, 1.7290, 2.2746, 3.5149, 3.3960, 0.9158, 2.4158, 0.1980, 0.6869, 0.0000, 1.3077, 1.1539, 0.1758, 2.4725, 2.6702, 0.6482, 2.3516, 0.1648, 0.6483, 0.0000, 3.1284, 1.5688, 4.1560, 4.1010, 5.0826, 4.0916, 5.5136, 1.8807, 1.3486, 0.0000, 2.8082, 1.2877, 2.6164, 3.7809, 3.9863, 3.9725, 7.9861, 1.7534, 1.1370}

[0158] Example 2, Method for predicting epitopes based on antibody CDR sequences (SEPPA-mAb)

[0159] 1. Dataset construction

[0160] The test dataset in this example was derived from antigen-antibody complex structures in the protein data bank (PDB) up to May 2019. First, a search in the PDB database was performed using keywords related to antigens and antibodies. Then, a series of filtering steps were applied to the search results to obtain a non-redundant set of paired antigen-antibody data. The specific filtering steps were as follows:

[0161] 1) Only complexes with antibodies having both light and heavy chains were retained;

[0162] 2) Complex structures with protein antigen primary sequences longer than 50 amino acids were retained;

[0163] 3) Complex structures with less than 10 epitope amino acids were removed;

[0164] 4) Remove redundancy of antigen epitopes according to similarity 100%.

[0165] After the above filtering steps, 1051 non-redundant antigen-antibody complex structures were finally obtained.

[0166] Then, based on time division training set and test set, 850 complex structures before July 2017 were taken as the training set, and the remaining 201 were removed 8, and the remaining 193 were taken as the independent test set. The PDB ID, antigen chain, antibody chain and antigen species information of the above data set are shown in Table A1 and Table A2.

[0167] 2. Identification interface extraction

[0168] The identification interface is composed of epitope amino acids and counter amino acids. The epitope amino acids and counter amino acids are defined based on the spatial distance method, and the threshold is set to 4 angstroms, that is, the epitope region is the amino acid with a distance less than 4 angstroms between any amino acid on the antigen side and the corresponding antibody side, and the counter amino acid on the antibody side is defined in the same way. The determination of antibody CDR is based on IMGT antibody numbering rule (CDR1: 27-38, CDR2: 56-65, CDR3: 105-117).

[0169] 3. Surface patch generation

[0170] For the input target antigen protein, a surface patch is generated for each surface amino acid on the antigen protein.

[0171] First, the surface amino acid is defined as the amino acid with a surface solvent accessible surface area (SASA) greater than The surface solvent accessible surface area is calculated by the program NaccessV2.1.1. Then, for each surface amino acid Resr, all neighbor amino acids Resi with a distance less than are extracted as its corresponding surface patch Patchr.

[0172]

[0173] where D min (Res i ,Res r ) represents the minimum atomic distance between any amino acid Res i and a specific surface amino acid Res r on the antigen protein.

[0174] 4. Generation of identification interface molecular fingerprint

[0175] Eight quantitatively described physicochemical properties that play a key role in antigen-antibody interaction were collected from the amino acid index database (AAindex database), including isoelectric point, hydrogen bond donor, hydrophobic interaction, van der Waals force, alpha helix, beta sheet and amino acid composition (ZIMJ680104_acid ZIMJ680104_base FAUJ880109 ARGP820101 FAUJ880103 GEIM800101 CHOP780101 and DAYM780101).

[0176] Based on the surface patch with a radius of 10 angstroms, a cylindrical model with a radius of 10 angstroms and a height of 20 angstroms was established to describe the spatial arrangement characteristics of the surface patch. The central axis of the cylinder was set as the vector passing through the center of the antigen and the center of the surface patch to be described, and the length of 20 angstroms from the center point of the surface patch in front and behind was set as the height of the cylinder; as a rotating plane with a radius of 10 angstroms rotates around the central axis, each amino acid in the surface patch will be mapped to a specific area of the cylinder; when the step length of the rotating radius is set to 2 angstroms and the step length of the central axis is set to 4 angstroms, a three-dimensional grid containing 25 positions (10 / 2*20 / 4) will be established; finally, for the above-mentioned eight descriptors, a string of molecular fingerprints containing 200 positions (25*8) is obtained to describe the surface patch.

[0177] Based on the above-mentioned cylindrical model of the antigen surface patch, a few modifications were made to construct the fingerprint descriptors of the antibody CDR. There are two intra-chain disulfide bonds in the hyper-variable region of the heavy chain and the light chain of the antibody (light chain A23-A88; heavy chain B22-B92), and the geometric center of the two disulfide bonds is taken as the center of the antigen in the above-mentioned antigen model, and the geometric center of the CDR is similar to the center of the surface patch. A similar cylindrical model is established to generate molecular fingerprints for CDR.

[0178] 5. Training a classification model based on machine learning algorithm

[0179] Based on the machine learning algorithm, the interaction probability of each surface patch on the antigen with the antibody CDR is predicted as the score of the surface patch. Subsequently, in order to obtain the epitope score of a single amino acid, all the surface patch scores involved in the target amino acid are mapped to the target amino acid based on the distance from the target amino acid to the center of the surface patch to obtain the initial epitope score. Finally, based on the overall antigenic trend of the neighbor amino acids, the initial score of the target amino acid is corrected, and after normalization, the final potential epitope score is output. The amino acid with high score is the predicted epitope.

[0180] The specific method is as follows:

[0181] 1) Surface patch score calculation based on machine learning algorithm

[0182] For each surface patch on the antigen protein, the fingerprint descriptors of the surface patch and the CDRs of the antibody are combined, resulting in 400 descriptors to predict the potential interaction probability of the surface patch and the CDRs. To divide the surface patches into positive and negative samples, first, the target value T of each surface patch is calculated, which is defined as the proportion of the amino acids of the real epitope in the surface patch. Then, for each antigen protein in the training dataset, the top 10 surface patches with a target value T greater than 0.5 are taken as positive samples, and 10 surface patches with a target value T equal to 0 are randomly taken as negative samples. In addition, based on the descriptors created above, eight commonly used machine learning algorithms are introduced to predict the scores of the surface patches, including extreme gradient boosting algorithm (XGBoost, XGB), support vector machine (SVM), random forrest (RF), decision tree (DT), multi-layer perceptron (MLP), gradient descent (GD), gaussian bayes (GNB), and linear regression (LR).

[0183] 2) Initial epitope score calculation

[0184] To determine whether the amino acid on the antigen protein is an epitope amino acid, the predicted surface patch score needs to be mapped to a single amino acid to obtain the initial epitope score. For any amino acid r on the antigen, consider all the surface patches involved, and take the distance d of the amino acid r to the center of each surface patch as a coefficient to integrate the scores of the surface patches to calculate the initial score of the amino acid r, raw_residue_score r As shown in equation 2.4:

[0185]

[0186] patch_score i is the score of a certain surface patch i in which the amino acid r is located, d is the distance between the amino acid r and the geometric center of the surface patch i, and M is the number of surface patches involved in the amino acid r.

[0187] 3) Classification and identification of epitope amino acids

[0188] A correction and standardization process is introduced to adjust the initial score to obtain the final epitope score of the amino acid.

[0189] First, the purpose of introducing the correction step is to correct the initial score of a single amino acid by the overall epitope tendency of its neighbor environment. For any amino acid r on the antigen, the correction score of r is defined as the average of the initial epitope scores of all surface amino acids within 5 angstroms distance from r, as shown in Equation 2.5:

[0190]

[0191] where raw_residue_score j represents the initial epitope score of any neighbor amino acid j with a minimum atomic distance less than 5 angstroms from amino acid r, and Σraw_residue_score j represents the sum of the initial epitope scores of the neighbor amino acids of amino acid r, and N represents the number of neighbor amino acids of amino acid r.

[0192] Subsequently, to make the scores of different protein antigens comparable, a normalization step is introduced to obtain the final epitope score. The score is normalized to the range interval of 0-1 using Equation 2.6.

[0193]

[0194] where min(adjust_residue_score) represents the minimum value of the corrected scores of all amino acids on the antigen, and max(adjust_residue_score) represents the maximum value.

[0195] The workflow diagram of the method for predicting antigen epitopes based on antibody CDR sequences (SEPPA-mAb) is shown in Figure 1 .

[0196] Example 3 Performance verification of SEPPA-mAb

[0197] Based on the training data set, 8 commonly used machine learning methods were tested by 5-fold cross-validation, including extreme gradient boosting algorithm (XGBoost, XGB), support vector machine (SVM), random forest (RF), decision tree (DT), multi-layer perceptron (MLP), gradient descent (GD), Gaussian Bayes algorithm (GNB), and linear regression (LR), to test the prediction surface patch scores of SEPPA-mAb under different machine learning algorithms.

[0198] As shown in Figure 2 ​As shown, the results of 5-fold cross-validation show that for the 850 antigen structures in the training dataset, the extreme gradient boosting algorithm XGBoost has the best effect, and the area under the curve (AUC) of the ROC receiver operating characteristic curve reaches 0.776, so this algorithm is used to construct the prediction model of SEPPA-mAb. For the independent test set containing 201 protein antigens, the AUC value of SEPPA-mAb for epitope prediction is 0.775 Figure 3 ) on this dataset is 0.682 when the threshold value of epitope division is set to 0.392.

[0199] Comparison of prediction effect with other algorithms

[0200] Comparison of prediction effect of the present application with five traditional epitope prediction algorithms.

[0201] For the 201 antigen proteins in the independent test set, only 193 proteins can be run and obtain prediction results in all algorithms, so the performance is compared based on the above-mentioned 193 antigen proteins and other algorithms.

[0202] First, for the first type of traditional epitope prediction algorithm, five online available B cell epitope prediction tools commonly used in the past decade are selected, including Epitopia, CBTOPE, Discotope 2.0, BepiPred 2.0 and SEPPA 3.0. In order to evaluate the prediction accuracy of each algorithm objectively and fairly, first of all, the area under the ROC curve (AUC) commonly used in binary classification models is used to measure the overall prediction effect of different algorithms, and then the balanced accuracy (BA) and F1 score are introduced to further compare the advantages and disadvantages of each method. For the above five different algorithms, the balanced accuracy of each algorithm is calculated based on the default threshold value, and the performance comparison results and the threshold value of each algorithm are shown in Table 4.

[0203] Table 4 Comparison of SEPPA-mAb with other similar algorithms

[0204] Method Threshold AUC Mean Accuracy F1 Score CBTOPE 4.00 0.549 0.539 0.114 Discotope2.0 -3.700 0.670 0.561 0.146 Epitopia 0.133 0.672 0.631 0.151 BcpiPred2.0 0.500 0.678 0.624 0.150 Seppa3.0 0.089 0.730 0.635 0.188 SEPPA-mAb 0.392 0.774 0.681 0.223

[0205] In combination with Table 4 and Figure 4As shown, the method of the present application can achieve the highest AUC value of 0.774, while the AUC value of the best SEPPA3.0

[45] among the other five prediction algorithms is 0.730. In addition, under the default threshold, the balance accuracy BA and F1 score of the method of the present application are 0.681 and 0.223, respectively, both of which are higher than those of other epitope prediction algorithms. The results show that the algorithm of the present application can significantly improve the epitope prediction performance by considering the recognition information of antigen-antibody interaction.

[0206] Example 4. Epitope prediction of SEPPA-mAb for HIV envelope glycoprotein

[0207] This example demonstrates the epitope prediction effect of the method of the present application and other algorithms on HIV glycoprotein antigen based on the method described in Example 2, taking HIV gp120 protein (PDB ID: 6cm3, Chain: E) as an example.

[0208] As shown in Figure 5 , based on the antigen-antibody complex corresponding to the target antigen in the PDB database, it can be known that the antibody bound to the antigen protein is 17b. Based on the 4 angstrom distance threshold, the antigen epitope amino acids targeted by the antibody 17b are extracted, and 13 real epitope amino acids are obtained, which correspond to the red spherical region of Figure 5 a. The epitope targeted by the antibody 17b predicted by SEPPA-mAb is shown in the spherical region of Figure 5 b, in which the potential epitopes (threshold value is 0.392) predicted by SEPPA-mAb are marked with three color gradients of red salmon and white, and the scores change from high to low. Among the 13 real epitope amino acids of the antigen protein, 11 of them are predicted as potential epitopes by SEPPA-mAb. Compared with other methods, the epitope amino acids predicted by the present application are closer to the real epitopes, indicating that SEPPA-mAb has good epitope recognition potential on pathogen antigen proteins.

[0209] Example 5. Antibody virtual screening method based on antigen epitope (ISAEP method)

[0210] The overall workflow of the antibody virtual screening model based on antigen epitope constructed in this example is shown in Figure 6 , and the input of the algorithm is the antigen epitope structure and the antibody library to be screened.

[0211] 1. Construction of antigen-antibody complex structure dataset

[0212] The antigen-antibody complex structure data used in this embodiment is from the antibody structure database SAbDab (The structural Antibody Database). The collection time is February 2022. First, search for antigen-antibody complex structures labeled by SabDab; then, select complexes with complete light and heavy chains for the antibody and a primary sequence length of the antigen protein greater than 50 amino acids, and complexes with epitopes greater than 8 amino acids; finally, a total of 3609 antigen-antibody complex structures are collected. For this data set, randomly divide the training and test sets in a ratio of 3:2, with 60% as the training set and 40% as the independent test set, for the construction and testing of the antibody virtual screening model.

[0213] The antibody complex structure data set includes antigen-antibody complex structure data set D, which is defined as follows:

[0214] D = {d i |i = 1 ~ N, d i = [Ag i1 Ab i1 , Ag i2 Ab i2 , Ag i3 Ab i3 … Ag in Ab in ]} (1)

[0215] Where I represents from 1 to N; Ag represents the antigen in each complex i; Ab represents the antibody in each complex i, a total of N complex data.

[0216] The de-redundant data set D uni is defined as follows:

[0217] D uni = {d i |i = 1 ~ N, if Ag i & Ab i & EpR i & PaR i = Ag i+m & Ab i+k & EpR i+m & PaR i+k , removeAg i+m & Ab i+k & EpR i+m & PaR i+k} (2)

[0218] 2. Construction of antibody library

[0219] Where NH h is the total number of human antibodies. lThe number of mouse-derived total antibodies is NM h *NM l Bar.

[0220] The antibody library used in the antibody screening of this embodiment includes an antibody structure library and an antibody sequence library. Antibody structures are collected from the antibody structure database SAbDab, and 2867 non-redundant structures are obtained as the antibody structure library based on 100% CDR similarity de-redundancy. Antibody sequences are collected from the NCBI protein database. The selected species are antibody sequences of Homo sapiens and Mus musculus, and de-redundancy is also based on CDRs. Therefore, for "Homo sapiens", 20892 antibody heavy chain sequences and 15505 antibody light chain sequences are obtained; for "Mus musculus", 6846 heavy chain sequences and 2333 light chain sequences are obtained. By combining the light and heavy chains of the antibodies, a total of 323,930,460 human-derived antibody sequences and 15,971,718 mouse-derived antibody sequences are obtained as the antibody sequence library.

[0221] The simulated antibody library is constructed by introducing different numbers of mutations randomly in the CDR region with a specific reference antibody as the initial template, and a total of 1.2*10 4 Bar simulated antibody sequences.

[0222] 3. Extraction of interface data set of antigen-antibody complex

[0223] For any antigen-antibody complex structure in the data set, the antigen epitope and antibody paratope are extracted as its recognition interface.

[0224] 4. Construction of antibody virtual screening model based on antigen epitope

[0225] By combining the molecular fingerprints on both sides of the antigen-antibody, machine learning algorithms such as support vector machine SVM, random forest RF, LightGBM LGB, XGBoost XGB, linear regression LR and decision tree DT are introduced, and the appropriate algorithm is selected according to the effect. Predict the binding probability of antigen-antibody; sort the antibodies based on the binding probability, and the top-ranked antibodies are the potential binding antibodies of the target antigen epitope.

[0226] The specific steps are as follows:

[0227] 1) Generation of molecular fingerprints of binding interface

[0228] For the target epitope structure, a spherical shell model with a total radius of 20 angstroms and 10 layers with a step of 2 angstroms is constructed based on its geometric center P; the amino acids in the epitope will be projected into different layers of the spherical shell based on their geometric distance to the center P; then the above 8 amino acid property indices are assigned to each layer of the spherical shell, and the eigenvalue of all amino acids in the layer of the spherical shell under a specific property is calculated; finally, for the 8 descriptors and the 10-layer spherical shell model, a total of 80-dimensional binding interface molecular fingerprints are calculated to describe the spatial and physicochemical microenvironment of the epitope.

[0229] Similarly, for the antibody CDRs, a similar spherical shell model is constructed with the geometric center of the CDRs as the center point P to generate the interaction-related feature descriptors of the antibody CDRs.

[0230] If the antibody only has sequence information, it needs to be mapped to a specific template antibody structure. First, based on the length of the CDRs of the antibody structure in the antibody structure library, each of the light and heavy chains of the antibody is divided into 3 intervals, so that 9 combinations of antibody light and heavy chain intervals can be obtained. In each combination interval, an antibody structure that meets the CDR length standard is selected as a template structure (Table 5) and a spherical shell model of its CDRs is constructed; based on the length of the CDRs of its light and heavy chains, a template antibody is matched. Then the CDRs of the antibody can be mapped to the spherical shell model of the template antibody based on the antibody Kabat numbering to obtain the layer of the spherical shell where each amino acid in the CDRs is located. Finally, combined with the 8 groups of amino acid property indices, the fingerprint descriptors of the antibody CDRs are calculated.

[0231] Table 5 Template antibody information

[0232] PDB ID Antibody Chain HCDRs Length LCDRs Length 6NMV HL 23 17 5NHR HL 24 23 6K68 CD 22 24 2EZ0 CD 30 17 7Q9G EF 28 18 7A5S HL 28 24 5WOB IJ 38 16 2NXZ DC 37 20 7CZX HK 37 24

[0233] 2) Training a classification model based on a machine learning algorithm

[0234] By combining the antigen epitope and antibody CDR descriptors, a total of 160-dimensional molecular fingerprints are obtained to predict the binding potential of antigen epitopes and antibody CDRs. For each target antigen protein, the positive sample is defined as the corresponding true antibody in the antigen-antibody complex. For the negative sample, 10 background antibodies are randomly selected as their corresponding negative samples. The selection rules of the negative sample are as follows: first, based on the antigenicity comparison tool CE-BLAST

[72] to calculate the antigenicity score of any two antigen epitopes in the dataset; then, for any target antigen, the antigen with a CE-BLAST score lower than 0.7 is defined as a dissimilar antigen, and the antibody corresponding to the dissimilar antigen is the potential non-binding antibody of the target antigen; finally, 10 non-binding antibodies of the target antigen are randomly selected as negative samples. Thus, based on the feature descriptors of the above positive and negative samples, six commonly used machine learning algorithms are introduced to predict the binding potential of the target antigen and any antibody to be screened in the antibody library. The machine learning algorithms tested in this application include: support vector machine SVM, random forest RF, LightGBM LGB, XGBoost XGB, linear regression LR and decision tree DT.

[0235] After the above model training is completed, the antigen epitope structure and virtual antibody library can be input to carry out screening.

[0236] Example 6. Performance verification of ISAEP method

[0237] The practicability of ISAEP is first verified by the divided 1444 independent spatial epitope test data sets. For the 1444 epitopes from the independent test data set, ISAEP can achieve an AUC value of 0.864 ( Figure 7 a).

[0238] In addition, we constructed the corresponding sub-models for two most widely studied immune hosts: mice and humans, respectively, based on the same method. The results show that the AUC value of the mouse sub-model can reach 0.890, and the AUC value of the human sub-model can reach 0.860 ( Figure 7 b-c).

[0239] In addition, based on the antibody structure library of 2867 non-redundant antibodies derived from SabDab, we analyzed it combined with the independent verification set data. For each epitope in the test data set, the binding potential of any antibody in the antibody structure library is predicted, and the antibodies are sorted in descending order based on the binding potential. Figure 8 a-c show the sorting of the true binding antibodies corresponding to the 1444 antigen epitopes in the independent test set in the antibody structure library. Figure 8 a The sorting position of the true antibody corresponding to the antigen is marked in orange, Figure 8b. Statistics of the proportion of antigen epitopes in each ranking interval, the results show that ISAEP can rank most of the real binding antibodies corresponding to the antigen in the front position. By looking at the proportion of the top 1% of antigens, it is found that ISAEP can rank 40% (574) of the real binding antibodies corresponding to the antigen in the top 5, 48% (691) of the antigens in the top 15, 51% (732) of the antigens in the top 25, and finally 52% (749) of the real binding antibodies corresponding to the antigen epitopes in the top 1% of the antigens. Figure 8 c). The above results show that ISAEP has good performance in predicting the potential binding CDR of the target epitope.

[0240] For the antibody level, since the antibody structure is usually difficult to obtain, with the development of antibody sequencing technology, antibody sequence information is relatively easy to obtain. Therefore, as an alternative, the antibody sequence is mapped to the template antibody structure to obtain its feature descriptor, and the prediction effect of ISAEP under this scheme is tested. For the CDR regions of heavy and light chains, according to the statistical analysis of the residue length of the antibody structure library of 2867 non-redundant sequences obtained in Example 5, further, the light and heavy chains are divided into three length intervals, respectively, as shown in Figure 9 Combining light and heavy chains, a total of 9 CDR length combinations can be obtained, and one template antibody is selected for each combination. Then, all sequences of the 1444 antibodies in our test data set are derived and mapped into different templates, and further tested by the model.

[0241] The results show that by mapping the antibody sequence to the template antibody, ISAEP can achieve similar prediction performance to the real antibody structure, with an AUC value of 0.884( Figure 10 a). For different immune hosts, sub-models are constructed, and the AUC value of ISAEP based on antibody sequence prediction on the human source model is 0.880( Figure 10 b), and the AUC value of the model on the mouse source model can reach 0.907( Figure 10 c). The above results show that for antibodies with only sequence information, by mapping them to the template antibody structure for application, ISAEP can still achieve good prediction effect, thereby further improving the practicability of the model.

[0242] In summary, whether starting from the antigen level, i.e. testing the influence of antigen epitope deletion on the prediction performance of the model, or starting from the antibody level, i.e. testing the prediction performance of the model based on antibody sequence information for antibody virtual screening, ISAEP shows good prediction effect, indicating the high robustness and practicability of the model for antibody screening.

[0243] Example 7. Application of antibody screening method based on protein molecular fingerprint descriptors (ISAEP method) in screening antibodies against epidemic pathogens

[0244] To test the practical application effect of antibody screening method ISAEP based on protein molecular fingerprint descriptors of the present application, it was applied to three important pathogens, including human immunodeficiency virus (HIV), influenza virus, and coronavirus. The main antigen proteins including glycoprotein GP120 of HIV, hemagglutinin protein HA of influenza A (IVA), and spike protein Spike of SARS-CoV-2 were strictly evaluated.

[0245] For each kind of antigen protein, the antigen is the epitope-CDR complex of the protein in the 1444 independent spatial epitope test data sets; which includes 158 GP120 protein epitopes, 91 HA protein epitopes and 239 Spike protein epitopes, as the specific epitope-CDR data set of each antigen.

[0246] Select one case epitope from each antigen type, and calculate its epitope similarity with other epitopes in the same antigen data set. The formula for calculating epitope similarity is as follows:

[0247]

[0248] Where Epi a , Epi b represent epitope a, epitope b (in the form of amino acid set) respectively. Epitope similarity sim(Epi a , Epi b ) is the number of intersection amino acids N(Epi a ∩Epi b ) divided by the number of union amino acids N(Epi a ∪Epi b ) between epitope a and epitope b.

[0249] According to the epitope similarity score, the antigen epitopes in the same antigen data set are clustered into highly similar epitopes (>0.8), similar epitopes (0.6-0.8) and dissimilar epitopes (<0.6). Then, the binding ability of all antibodies in the same antigen data set to the epitope is predicted using ISAEP, and the predicted binding score of each antibody to the case epitope is obtained.

[0250] The results show that the case epitope of GP120 (PDB ID: 2nxz, Chain: A) corresponding to the original neutralizing antibody 17b (PDB ID: 2nxz, Chain: D&C) ranks the 4th among the 158 queries of the GP120 antibody library, and the top 3 are the same antibodies with subtle spatial structure differences. The case epitope of HA (PDB ID: 4o5i, Chain: C) corresponding to the original neutralizing antibody F045-092 (PDB ID: 4o5i, Chain: O&P) ranks the 1st among the 91 queries of the HA antibody library. The case epitope of Spike (PDB ID: 6xcn, Chain: C) corresponding to the original neutralizing antibody C105 (PDB ID: 6xcn, Chain: H&L) ranks the 1st among the 239 queries of the Spike antibody library. It is illustrated that ISAEP has practical neutralizing antibody screening capability in the three pathogens.

[0251] In combination with the epitope similarity clustering results, it is found that in the three infectious diseases, the antibodies with higher binding scores are the antibodies with similar corresponding epitopes to the case epitopes (PDB ID: 2nxz, Chain: A; PDB ID: 4o5i, Chain: C; PDB ID: 6xcn, Chain: C), which confirms that the antibodies with similar epitopes have higher cross-reactivity potential. Figure 11

[0252] Subsequently, the practical value of ISAEP is tested in the high-throughput antibody sequence library, and the antigen cases are selected to perform antibody virtual screening in the non-redundant antibody sequence library of the immune host corresponding to the real antibodies bound to the antigen themselves. The application example constructs a potential antibody sequence library of human and mouse origin, respectively containing 3.239*10 8 7 The selected cases are shown in the following table:

[0253] Table 6 Information of antibody sequence library test cases

[0254] Antigen Type PDB ID Antigen Chain Antibody Chain Antibody Name HIV GP120 1n6q B H&L Fab28 IVA HA 4mhh E H&L H5M9 SARS2 Spike 7eam A H&L 7D6

[0255] Specifically, the binding potential of the target antigen epitope and any antibody in the antibody library is predicted using ISAEP, and the antibodies are sorted based on the binding probability. The results show that ISAEP can rank the antibodies that are actually bound to the case and antigen at the 1st position in the three pathogens, indicating the high accuracy of ISAEP for antibody screening of HIV, influenza and new crown.

[0256] Subsequently, the antibodies in the antibody library are divided into different score intervals based on the prediction scores of ISAEP from high to low. For the intervals with more than 1000 antibodies, 1000 sequences are randomly selected to view their CDR similarity with the reference antibodies.

[0257] The results are as follows:​​Figure 12 As shown, in the first interval, more than half of the antibodies have more than 80% similarity with the reference antibody, and in intervals 2-4, more than half of the antibodies have more than 60% similarity with the reference antibody; the antibodies in the first four intervals have higher similarity with the reference antibody, indicating that antibodies with higher ISAEP prediction scores are more similar to the reference antibody, which may contain potential binding antibodies of the target antigen.

[0258] In addition, for the three case antigens selected this time, we constructed simulated antibody libraries with different mutation degrees 1-10, 15, and 20 based on their real binding antibodies, and further tested the ability of ISAEP to distinguish antibodies with different similarities. As shown in Figure 13 As shown, with the increase of the number of mutations, the binding probability of the antibody to the target antigen gradually decreases.

[0259] In summary, the above results show that for the hemagglutinin protein HA of influenza A (IVA), the glycoprotein GP120 of HIV, and the spike protein Spike of SARS-CoV-2, ISAEP can screen their real binding antibodies, and the antibodies with high prediction ranking are more similar to the real binding antibodies. This shows that ISAEP can provide computational support for the rapid screening of potential binding antibodies of the three infectious disease main pathogenic proteins.

[0260] APPENDIX A SEPPA-mAb-Antibody Specific Antigen Epitope Prediction Algorithm

[0261] Table A1 Training data set of SEPPA-mAb

[0262]

[0263] Continued

[0264]

[0265]

[0266] Continued

[0267]

[0268] Continued

[0269]

[0270] Continued

[0271]

[0272] Continued

[0273]

[0274] Continued

[0275]

[0276] Continued

[0277]

[0278] PDB ID a PDB number, antigen chain and antibody chain representing the antigen-antibody complex structure; Species b H for human origin, B for bacterial, V for viral, O for other host species representing the antigen species origin.

[0279] Table A2 Test data set for SEPPA-mAbs

[0280]

[0281] Continued

[0282]

[0283] Continued

[0284]

[0285] PDB ID a PDB number, antigen chain and antibody chain representing the antigen-antibody complex structure;

[0286] Species b H for human origin, B for bacterial, V for viral, O for other host species representing the antigen species origin.

[0287] * Antigen structures for which epitope prediction software Epitopia, DiscoTope2 could not be run.

Claims

1. A method for constructing molecular fingerprints of protein recognition interfaces based on geometric models, characterized in that, Includes the following steps: S01 Determine the geometric center point of the protein molecule and its recognition interface: The geometric center point of the protein molecule is obtained by taking the average of the spatial coordinates of a specific amino acid at the recognition interface of all amino acids in the protein molecule, which is the coordinate of its geometric center point. The geometric center point of the recognition interface is obtained by taking the average of the spatial coordinates of a specific amino acid on the recognition interface. S02 Determines the method for describing the geometric model of a protein molecule: The method for describing the protein molecular geometry model is based on the protein's spatial conformation, using different geometric models for description. The geometric models are selected from: spheres, cylinders, ellipsoids, cubes, or cones. When the sphere model is selected, the structure described by the sphere model includes antigen protein, antigen epitope region, and antibody CDR region; when the cylinder model is selected, the structure described by the cylinder model includes the structural features of the surface sheet of the antigen or antibody protein. S03 Protein Molecular Geometry Model Subregion Division: The sub-region division of the protein molecule geometry model is based on the selected geometry model, which is divided into several sub-regions. The principle of sub-region division is to ensure that each sub-region covers an appropriate number of amino acids. The size range of the sub-region needs to be determined by combining the size of different proteins, the size of amino acid molecules, and spatial distance. S04 Selection of amino acid physicochemical property indices for protein molecules; S05 Constructing the molecular fingerprint of the protein recognition interface: The molecular fingerprint of the protein recognition interface is constructed based on the selected amino acid physicochemical property index. The characteristic value of each amino acid physicochemical property in each sub-region is calculated, and the characteristic values ​​of all amino acids in all sub-regions constitute the molecular fingerprint of the protein recognition interface.

2. The method for constructing molecular fingerprints of protein recognition interfaces based on geometric models as described in claim 1, characterized in that, The structure described by the geometric model includes antigen protein, antigen epitope region, antibody CDR region, and antigen or antibody protein surface sheet; when selected from the sphere model, the structure described by the sphere model includes antigen protein, antigen epitope region, and antibody CDR region; when selected from the cylinder model, the structure described by the cylinder model includes the structural features of antigen or antibody protein surface sheet.

3. The method for constructing molecular fingerprints of protein recognition interfaces based on geometric models as described in claim 1, characterized in that, The selection of the physicochemical property indices of the amino acids in the protein molecules is based on the following amino acid physicochemical properties: acidity, basicity, number of hydrogen bond donors, hydrophobicity index, van der Waals forces, α-helix, β-turn angle, and amino acid composition.

4. A method for predicting antigenic epitopes based on corresponding antibody CDR sequences, the method comprising the following steps: Construction of the S01 antigen-antibody complex dataset: Construct a dataset of paired antigen-antibody complexes with non-redundant epitopes; Extraction of the identification interface for the S02 antigen-antibody complex dataset: From the antigen-antibody complex structure dataset, amino acids with a minimum atomic distance of less than 4 Å between the antigen and antibody sides are extracted to form the recognition interface; S03 Construction of an antigenic epitope prediction model based on antibody CDR sequences: (1) Generate a surface sheet around each surface amino acid on the antigen protein in the antigen-antibody complex dataset; (2) Using a cylindrical model, molecular fingerprints of the protein recognition interface are constructed for each surface sheet of the antigen and the CDR structure of the antibody; the molecular fingerprints are constructed according to the method described in claim 1. (3) Training a classification model based on machine learning algorithm: The molecular fingerprint of antigen surface sheet and antibody CDR structure is combined as training data. The classification model for antigen epitope prediction is trained by machine learning algorithm. During the training process of the classification model, surface sheets containing more than 30% of the real epitope amino acids in the training set data are set as positive samples, and surface sheets that do not contain real epitope amino acids are set as negative samples. S04 Use a classification model to obtain the antigenic epitope to be predicted: Input the protein structure information of the antigen to be predicted and the CDR structure information of the corresponding antibody into the model constructed in the previous step to obtain the epitope information of the antigen to be predicted.

5. A virtual antibody screening method based on antigenic epitopes, characterized in that, The screening method is based on antigen-antibody specific interaction characteristics and constructs a rapid antibody screening algorithm to perform rapid virtual antibody screening for specified conformational epitopes of the target antigen, including the following steps: Construction of the S01 antigen-antibody complex structure dataset; Construction of the S02 antibody library; Extraction of the identification interface of the S03 antigen-antibody complex dataset; S04 Construction of a virtual antibody screening model based on antigenic epitopes: (1) Molecular fingerprints of protein recognition interfaces are constructed by using the spherical model as the antigen epitope and antibody CDR region in the antigen-antibody complex structure dataset; the molecular fingerprints are constructed according to the method described in claim 1. (2) Model training based on machine learning algorithm: The molecular fingerprint of antigen epitope and antibody CDR structure is combined as training data. The prediction model of antigen-antibody binding probability is trained by machine learning algorithm. During the model training process, known real antigen-antibody pairs in the training set data are set as positive samples, and known antibodies that bind to irrelevant antigens are set as negative samples. S05 uses the trained model to perform virtual antibody screening.

6. The antibody virtual screening method based on antigenic epitopes as described in claim 5, characterized in that, The antigen-antibody complex structure dataset was constructed by selecting complex structures with complete light and heavy chains and an antigen protein primary sequence length greater than 50 consecutive amino acids.

7. The antibody virtual screening method based on antigenic epitopes as described in claim 5, characterized in that, The extraction of the recognition interface of the antigen-antibody complex dataset involves extracting amino acids with a minimum atomic distance of 4 Å between the antigen and antibody sides from the antigen-antibody complex structure dataset to form the recognition interface.

8. The antibody virtual screening method based on antigenic epitopes as described in claim 5, characterized in that, The process of using the trained model for virtual antibody screening involves inputting a specified epitope of a known antigen structure and using the model trained in S04 to perform virtual screening in the antibody library to obtain potential binding antibodies.