Systems and methods for predicting b-cell epitopes

A graph-based machine learning approach for predicting B-cell epitopes and paratopes addresses the inefficiencies of traditional methods by using protein and binder representations, achieving accurate and efficient predictions.

WO2025240561A1PCT designated stage Publication Date: 2025-11-20SEISMIC THERAPEUTICS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/029274
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-05-24
Filing Date
2025-05-14
Publication Date
2025-11-20

AI Technical Summary

Technical Problem

Traditional experimental methods for identifying B-cell epitopes are time-consuming and resource-intensive, and existing in silico approaches for predicting B-cell epitopes are lacking due to the conformational complexity and scarcity of structural data.

Method used

A method using graph representations of proteins and binders, combined with machine learning algorithms trained on loop-mediated protein-protein interaction interfaces, to predict B-cell epitopes and paratopes through bipartite graph construction and collaborative filtering.

Benefits of technology

The method achieves accurate and efficient prediction of B-cell epitopes and paratopes, demonstrated by consistent performance against experimental data, improving the accuracy of antibody research and biologics development.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025029274_20112025_PF_FP_ABST
    Figure US2025029274_20112025_PF_FP_ABST
Patent Text Reader

Abstract

Predictive models are deployed to generate epitope predictions (e.g., B-cell epitopes). Predictive models analyze structural features of a protein, such as an antibody, or an antibody-antigen pair, and can identify epitopes and paratopes.
Need to check novelty before this filing date? Find Prior Art

Description

Attorney Docket No.: SES-014WO PATENT SYSTEMS AND METHODS FOR PREDICTING B-CELL EPITOPES CROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims priority to U.S. Provisional Application No. 63 / 647,407, filed May 14, 2024, and U.S. Provisional Application No. 63 / 651,590, filed May 24, 2024, each of which is hereby incorporated by reference in its entirety. BACKGROUND

[0002] The accurate identification of B-cell epitopes is crucial to the development of antibodies and biologics, but traditional experimental methods for epitope identification are time- consuming and resource-intensive. While there exist robust methods to predict T-cell epitopes in silico using machine learning, reliable in silico approaches have yet to be developed for the prediction of B-cell epitopes, due largely to their conformational complexity and the sparsity of publicly available structural data. The embodiments provided for herein fulfill this need as well as others. SUMMARY

[0003] In some embodiments, provided herein is a method for identifying a B-cell epitope of a protein, the method comprising: (a) obtaining or having obtained a graph representation of the protein; (b) providing the graph representation of the protein to a generalized machine learning algorithm model to generate a plurality of predicted probabilities, wherein the generalized machine learning algorithm model is trained using a training data set comprising at least 1,000 loop‑mediated protein‑protein interaction interfaces; (c) classifying each of one or more amino acids of the protein as an epitope-member residue or not an epitope-member residue based on the plurality of predicted probabilities; and (d) generating a B-cell epitope prediction comprising the classified epitope-member residues.

[0004] In some embodiments, provided herein is a method for identifying a B-cell epitope of a protein, and a paratope of a binder of the epitope, the method comprising: (a) obtaining or having obtained a graph representation of the protein and a graph representation of the binder; (b) transforming the obtained graph representation of the protein and the graph representation of the binder to construct a bipartite graph comprising: a plurality of protein embeddings; a plurality of binder embeddings; and a plurality of edges connecting protein embeddings and binder 1 IPTS / 128951704.1Attorney Docket No.: SES-014WO PATENT embeddings, wherein the plurality of edges comprise predicted weights determined through graph-based collaborative filtering; (c) providing the bipartite graph to a machine learning model to generate a plurality of predicted probabilities; (d) classifying each amino acid of the protein as an epitope-member residue or not an epitope-member residue, and each amino acid of the binder as a paratope-member or not a paratope-member based on the plurality of predicted probabilities; and (e) generating a B-cell epitope prediction comprising the classified epitope-member residues.

[0005] In some embodiments, provided herein is a non-transitory computer-readable storage medium, the computer-readable storage medium comprising instructions that when executed by a processor, cause the processor to: (a) obtain or having obtained a graph representation of a protein; (b) providing the graph representation of the protein to a generalized machine learning algorithm model to generate a plurality of predicted probabilities, wherein the generalized machine learning algorithm model is trained using a training data set comprising at least 1,000 loop‑mediated protein‑protein interaction interfaces; (c) classify each of one or more amino acids of the protein as an epitope-member residue or not an epitope-member residue based on the plurality of predicted probabilities; and (d) generate a B-cell epitope prediction comprising the classified epitope-member residues.

[0006] In some embodiments, provided herein is a non-transitory computer-readable storage medium, the computer-readable storage medium comprising instructions that when executed by a processor, cause the processor to: (a) obtain or having obtained a graph representation of a protein and a graph representation of a binder; (b) transform the obtained the graph representation of the protein and the graph representation of the binder to construct a bipartite graph comprising: a plurality of protein embeddings; a plurality of binder embeddings; and a plurality of edges connecting protein embeddings and binder embeddings, wherein the plurality of edges comprise predicted weights determined through graph-based collaborative filtering; (c) provide the bipartite graph to a machine learning model to generate a plurality of predicted probabilities; (d) classify each amino acid of the protein as an epitope-member residue or not an epitope-member residue, and each amino acid of the binder as a paratope-member or not a paratope-member based on the plurality of predicted probabilities; and (e) generate a B-cell epitope prediction comprising the classified epitope-member residues.

[0007] In some embodiments, provided herein is a system comprising a non-transitory computer-readable storage medium and a processor, wherein the non-transitory computer- 2 IPTS / 128951704.1Attorney Docket No.: SES-014WO PATENT readable storage medium comprises: a) a bipartite graph encoded on the non-transitory computer-readable storage medium; and b) instructions for generating a B-cell epitope prediction comprising a classified epitope-member residues from bipartite graph, wherein the bipartite graph is a transformation of a graph representation of a protein and a graph representation of a binder. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] These and other features, aspects, and advantages of the present disclosure will become better understood with regard to the following description and accompanying drawings.

[0009] Figure (FIG.) 1A is an example system overview for predicting an epitope of a protein (such as a protein sequence, or a protein structure).

[0010] FIG. 1B depicts an example block diagram for an predicting an epitope a protein sequence.

[0011] FIG. 1C depicts an example block diagram for an predicting an epitope an antigen and a pratope of an antibody.

[0012] FIGs. 2A-2D depict a process for predicting an epitope a protein sequence.

[0013] FIGs. 3A-3B depict a process for predicting an epitope an antigen and a pratope of an antibody.

[0014] FIG. 4 shows illustrates an example computer for implementing the predictive models, methods, and systems disclosed herein.

[0015] FIG. 5 illustrates schematic of the EpiGraph.

[0016] FIG. 6 panel (A) shows selection of sequence similarity thresholds (yellow X) to partition the training and test sets. Panel (B) shows that node and edge features capture geometric and biochemical properties of the antigen and antibody. Panel (C) shows the antibody-agnostic epitope prediction model comprising message-passing operations on the antigen graph alone, followed by a residue-level classification layer. Panel (D) shows that the antibody-specific epitope prediction model employs the collaborative filtering technique together with a graph attention network applied to learned representations of both the antigen and antibody feature graphs.

[0017] FIG. 7 shows, in panel (A), example ground truth comparisons of predictions from the antibody-agnostic EpiGraph model for antigen monomer structures in the testing set. In each case, the experimentally mapped epitope is a subset of the epitopes predicted by the model. Panel 3 IPTS / 128951704.1Attorney Docket No.: SES-014WO PATENT (B) shows example ground truth comparisons of epitope predictions from the antibody-specific model for antigen-antibody complex structures in the test set. Panel (C) shows schematic of the twin graph convolutional network architecture. Panel (D) shows that recommender-style GAT models consistently outperform twin GCN models trained in parallel on the same dataset.

[0018] FIG. 8 shows, in panel (A), EpiGraph’s prediction for the SM201 epitope on FcγRIIb is confirmed by experimental mapping of the epitope by hydrogen-deuterium exchange mass spectrometry (HDX-MS). Here, the strongly predicted epitope consists of residues assigned a probability >0.4 by EpiGraph, while the weakly predicted epitope consists of residues with probability between 0.25 and 0.4. The structure of FcγRIIb is from PDB ID 3WJJ. Panel (B) is the AlphaFold2 prediction for the SM201:FcγRIIb complex. AF2 incorrectly places the antibody (green and yellow chains) at the endogenous Fc interaction site instead of the experimentally mapped epitope (red residues). Panel (C) shows the two anti-FcγRIIb antibodies with publicly available structures (Top: PDB 6YAX, Bottom: PDB 5OCC) bind FcγRIIb at sites distant from the SM201 epitope (red residues). Panel (D) shows EpiGraph predictions superimposed on the cryoEM structure of the T-cell activation surface protein marker PD-1 in complex with a novel antibody whose epitope has not been previously reported.

[0019] FIG. 9 shows that EpiGraph predicts the structural regions that pre-existing ADAs are most likely to target on the IgG-cleaving protease IdeS from Streptococcus pyogenes. Panel (A) shows effects of mutagenesis in eight surface regions on binding of pre-existing antibodies from IVIG, measured experimentally by ELISA. Panel (B) shows that EpiGraph correctly ranks region A as the immunodominant hotspot (cyan loops are the IgG1 Fc pose from PDB 8A47). Panel (C) shows that a state-of-the-art publicly available model (DiscoTope3) failed to correctly identify the immunodominant hotspots, instead prioritizing the IgG1 Fc interface region.

[0020] FIG. 10 is a cryoEM structure of SM201:FcgRIIb complex with the dotted outline showing previously predicted epitope site. DETAILED DESCRIPTION

[0021] All technical and scientific terms used herein, unless otherwise defined below, are intended to have the same meaning as commonly understood by one of ordinary skill in the art. The mention of techniques employed herein are intended to refer to the techniques as commonly understood in the art, including variations on those techniques or substitutions of equivalent techniques that would be apparent to one of skill in the art. While the following terms are 4 IPTS / 128951704.1Attorney Docket No.: SES-014WO PATENT believed to be well understood by one of ordinary skill in the art, the following definitions are set forth to facilitate explanation of the presently disclosed subject matter.

[0022] As used herein and in the appended claims, the singular forms “a”, “an” and “the” include plural reference unless the context clearly dictates otherwise.

[0023] As used herein, the term “about” means that the numerical value is approximate and small variations would not significantly affect the practice of the disclosed embodiments. Where a numerical limitation is used, unless indicated otherwise by the context, “about” means the numerical value can vary by ±5% and remain within the scope of the disclosed embodiments. Thus, about 100 means 95 to 105.

[0024] It should be understood that the term “at least one of” includes individually each of the recited objects after the expression and the various combinations of two or more of the recited objects unless otherwise understood from the context and use. The term “and / or” in connection with three or more recited objects should be understood to have the same meaning unless otherwise understood from the context.

[0025] As used herein, the terms “comprising” (and any form of comprising, such as “comprise”, “comprises”, and “comprised”), “having” (and any form of having, such as “have” and “has”), “including” (and any form of including, such as “includes” and “include”), or “containing” (and any form of containing, such as “contains” and “contain”), are inclusive or open-ended and do not exclude additional, unrecited elements or method steps. Any composition or method that recites the term “comprising” should also be understood to also describe such compositions as consisting, consisting of, or consisting essentially of the recited components or elements.

[0026] It should be understood that the order of steps or order for performing certain actions is immaterial, unless explicitly stated that the order of the steps is required or that the order is necessary for the embodiments to remain operable. Moreover, two or more steps or actions may be conducted simultaneously.

[0027] At various places in the present specification, variable or parameters are disclosed in groups or in ranges. It is specifically intended that the description include each and every individual subcombination of the members of such groups and ranges. For example, an integer in the range of 0 to 5 is specifically intended to individually disclose 0, 1, 2, 3, 4, 5, and an integer in the range of 1 to 3 is specifically intended to individually disclose 1, 2, and 3. 5 IPTS / 128951704.1Attorney Docket No.: SES-014WO PATENT

[0028] The use of any and all examples, or exemplary language herein, for example, “such as” or “including,” is intended merely to illustrate better the present disclosure and does not pose a limitation on the scope of any embodimet(s) unless claimed. No language in the specification should be construed as indicating any non-claimed element as essential to the practice of that provided by the present disclosure.

[0029] As used herein, the terms “protein” and “polypeptide” are used interchangeably and generally refer to a macromolecule that includes one or more linked chains of amino acid residues, which can be natural amino acids, unnatural amino acids or both.

[0030] As used herein, the term “epitope” refers to a region of a protein that is specifically recognized by a binding partner, such as an antibody or another binding protein. The epitope may generally span a portion of the protein. Often, proteins may have multiple such regions where binding partners can attach. Epitopes typically fall into two classes: continuous epitopes (also known as linear epitopes), which are epitopes defined by linear sequences of consecutive amino acids, and discontinuous epitopes (also known as conformational epitopes), which are epitopes defined by discontinuous amino acids that are brought together into spatial proximity when a protein is in its folded state.

[0031] As used herein, the term "paratope" refers to the specific region of a binding molecule (e.g., binding protein) that recognizes and binds an epitope of a target molecule. The paratope typically comprises 5-20 amino acids.

[0032] As used herein, the phrase “structural feature(s)” is generally used in the context of amino acids, e.g., an amino acid residue present in a polypeptide.

[0033] The term “subject” encompasses a cell, tissue, or organism, human or non-human, whether in vivo, ex vivo, or in vitro, male or female.

[0034] The term “mammal” encompasses both humans and non-humans and includes but is not limited to humans, non-human primates, canines, felines, murines, bovines, equines, and porcines.

[0035] The term “sample” can include a single cell or multiple cells or fragments of cells or an aliquot of body fluid, such as a blood sample, taken from a subject, by means including venipuncture, excretion, ejaculation, massage, biopsy, needle aspirate, lavage sample, scraping, surgical incision, or intervention or other means known in the art. Examples of an aliquot of 6 IPTS / 128951704.1Attorney Docket No.: SES-014WO PATENT body fluid include amniotic fluid, aqueous humor, bile, lymph, breast milk, interstitial fluid, blood, blood plasma, cerumen (earwax), Cowper’s fluid (pre-ejaculatory fluid), chyle, chyme, female ejaculate, menses, mucus, saliva, urine, vomit, tears, vaginal lubrication, sweat, serum, semen, sebum, pus, pleural fluid, cerebrospinal fluid, synovial fluid, intracellular fluid, and vitreous humour.

[0036] The terms “marker,” “markers,” “biomarker,” and “biomarkers” encompass, without limitation, lipids, lipoproteins, proteins, cytokines, chemokines, growth factors, peptides, nucleic acids, genes, and oligonucleotides, together with their related complexes, metabolites, mutations, variants, polymorphisms, modifications, fragments, subunits, degradation products, elements, and other analytes or sample-derived measures. A marker can also include mutated proteins, mutated nucleic acids, variations in copy numbers, and / or transcript variants, in circumstances in which such mutations, variations in copy number and / or transcript variants are useful for generating a predictive model, or are useful in predictive models developed using related markers (e.g., non-mutated versions of the proteins or nucleic acids, alternative transcripts, etc.).

[0037] The term "antibody" is used in the broadest sense and specifically covers monoclonal antibodies (including full length monoclonal antibodies), polyclonal antibodies, multispecific antibodies (e.g., bispecific antibodies), and antibody fragments that are antigen-binding so long as they exhibit the desired biological activity, e.g., an antibody or an antigen-binding fragment thereof.

[0038] "Antibody fragment", and all grammatical variants thereof, as used herein are defined as a portion of an intact antibody comprising the antigen binding site or variable region of the intact antibody, wherein the portion is free of the constant heavy chain domains (i.e. CH2, CH3, and CH4, depending on antibody isotype) of the Fc region of the intact antibody. Examples of antibody fragments include Fab, Fab', Fab'-SH, F(ab')2, and Fv fragments; diabodies; any antibody fragment that is a polypeptide having a primary structure consisting of one uninterrupted sequence of contiguous amino acid residues (referred to herein as a "single-chain antibody fragment" or "single chain polypeptide").

[0039] Predictive models, as disclosed herein, are useful for identifying epitopes, such as B-cell epitopes, T-cell epitopes, and others. In some embodiments, predictive models disclosed herein are useful for identifying paratopes. In some embodiments, predictive models disclosed herein 7 IPTS / 128951704.1Attorney Docket No.: SES-014WO PATENT are useful for identifying epitopes, such as B-cell epitopes, T-cell epitopes, and others, and paratopes.

[0040] Disclosed herein are methods, systems, and non-transitory computer readable media for predicting one or more epitopes, such as, without limitation, B-cell epitopes, T-cell epitopes, and others, of a protein (such as a protein sequence, or a protein structure).

[0041] FIG. 1A is an exemplary system overview for predicting one or more epitopes, such as, without limitation, B-cell epitopes, T-cell epitopes, and others, of a protein (such as a protein sequence, or a protein structure). The system overview includes an epitope prediction system 130 and one or more third party entities 110A and / or 110B in communication with one another through a network 120. FIG. 1A depicts one embodiment of the overall system environment. In other embodiments, additional or fewer third party entities 110 in communication with the protein analysis system 130 can be included. Generally, the epitope prediction system 130 implements methods disclosed herein for predicting one or more epitopes, such as, without limitation, B-cell epitopes, T-cell epitopes, and others, of a protein (such as a protein sequence, or a protein structure). As referenced herein, methods for predicting one or more epitopes, such as, without limitation, B-cell epitopes, T-cell epitopes, and others, of a protein (such as a protein sequence, or a protein structure) include one or more of: obtaining or having obtained a graph representation of the protein, providing the graph representation to a generalized machine learning algorithm model to generate a plurality of predicted probabilities, wherein the generalized machine learning algorithm model has been trained using a non-curated training data set, classifying each of one or more amino acids of the protein as an epitope-member residue or not an epitope-member residue based on the plurality of predicted probabilities, and generating a B-cell epitope prediction comprising the classified epitope-member residues. In some embodiments, methods for predicting one or more epitopes, such as, without limitation, B- cell epitopes, T-cell epitopes, and others, of a protein (such as a protein sequence, or a protein structure)include one or more of: obtaining or having obtained a graph representation of the protein and a graph representation of the binder, transforming the obtained graph representation of the protein and graph representation of the binder to construct a bipartite graph, providing the bipartite graph to a machine learning model to generate a plurality of predicted probabilities, classifying each amino acid of the protein as an epitope-member residue or not an epitope- member residue, and each amino acid of the binder as a paratope-member or not a paratope- 8 IPTS / 128951704.1Attorney Docket No.: SES-014WO PATENT member based on the plurality of predicted probabilities, and generating a B-cell epitope prediction comprising the classified epitope-member residues. The third party entities 110 communicate with the epitope prediction system 130 for purposes associated with using the characterized one or more structural features of the protein sequence e.g., for rational drug discovery.

[0042] In various embodiments, the third party entity 110 represents a partner entity of the epitope prediction system 130 that operates either upstream or downstream of the epitope prediction system 130. As one example, the third party entity 110 operates upstream of the epitope prediction system 130 and provide information to the epitope prediction system 130 to enable the characterization of one or more structural features of a protein (such as a protein sequence, or a protein structure). In this scenario, the epitope prediction system 130 receives data from the third party entity 110, examples of which protein sequence(s) and / or characteristics of protein sequence(s) (e.g., a protein interface of a protein (such as a protein sequence, or a protein structure)). The epitope prediction system 130 performs methods disclosed herein to characterize one or more structural features of the protein sequence.

[0043] Referring to the network 120, any suitable network 120 may be implemented that enables connection between the epitope prediction system 130 and third party entities 110. The network 120 may comprise any combination of local area and / or wide area networks, using both wired and / or wireless communication systems. In one embodiment, the network 120 uses standard communications technologies and / or protocols. For example, the network 120 includes communication links using technologies such as Ethernet, 802.11, worldwide interoperability for microwave access (WiMAX), 3G, 4G, code division multiple access (CDMA), digital subscriber line (DSL), etc. Examples of networking protocols used for communicating via the network 704 include multiprotocol label switching (MPLS), transmission control protocol / Internet protocol (TCP / IP), hypertext transport protocol (HTTP), simple mail transfer protocol (SMTP), and file transfer protocol (FTP). Data exchanged over the network 704 may be represented using any suitable format, such as hypertext markup language (HTML) or extensible markup language (XML). In some embodiments, all or some of the communication links of the network 120 may be encrypted using any suitable technique or techniques. 9 IPTS / 128951704.1Attorney Docket No.: SES-014WO PATENT Predictive Model

[0044] Since the FDA approved Humulin as the first recombinant, protein-based therapeutic in 1982, development of protein-based drugs has accelerated rapidly, with hundreds of candidates approved and in clinical trials. Anti-drug antibodies (ADAs) elicited against a therapeutic protein can neutralize its activity, impair its pharmacokinetic properties, induce toxicity, and in some cases lead to autoimmune disease. The surface regions on proteins to which antibodies bind are known as B-cell epitopes. They can take the form of short stretches of contiguous amino acid residues or larger conformational epitopes comprising distal residues in the protein’s linear amino acid sequence. Such epitopes are challenging to map experimentally, requiring time- consuming and expensive experiments capable of probing 3D protein structure such as hydrogen-deuterium mass spectrometry or X-ray crystallography. Therefore, the task of predicting and modulating these epitopes in silico is an important and challenging problem for antibody research and biologics development. A number of computational methods aimed at performing such predictions have arisen in recent years. Such methods must confront the significant challenge of leveraging 3D structural information rather than dealing only with linear protein sequence, leading to a wide variety of proposed model architectures. However, there has been little to no experimental validation of these models, and benchmarking studies have demonstrated that the majority perform poorly in general, typically no better than random guessing.

[0045] Current epitope predictors, such as B-cell epitope predictors, have highly varied output predictions and do not consistently agree with experimental data. For example, on a rigorous benchmarking task of nine leading epitope prediction tools, the ROC-AUC values for predicted epitopes were shown to be almost equivalent to random predictions. In particular, many methods did not have better Matthews correlation coefficients (MCC) with ground truth than generating random surface residues or random patches. Furthermore, only one of these methods (EpiPred, deprecated as of May 2022) predicts the epitope specific to a given antibody, rather than predicting all potential epitopes on a protein’s surface. As monoclonal antibodies are the largest class of approved biologic drugs, this limits the applicability of these methods in drug development settings. Finally, while these methods train and test on existing epitope annotations derived from experimental data, they do not further validate their models by experimentally 10 IPTS / 128951704.1Attorney Docket No.: SES-014WO PATENT mapping previously unobserved epitopes, leaving open the question of their applicability to real- world prediction tasks.

[0046] A variety of tools to predict epitopes, such as B cell epitopes, based on linear sequences alone or via explicit geometric information derived from protein structure. Methods that only exploit linear sequence enjoy the advantage of significantly larger datasets of antibody antigen interactions, but are highly limited in the sense that over 90% of known B-cell epitopes are non- linear and exhibit complex 3D topologies. Methods incorporating structural information tend to score individual residues or small surface patches based on their solvent accessibility and sidechain orientation, or compute antibody-antigen complementarity measures via computationally intensive docking procedures. While such methods have been available for almost twenty years, they have performed poorly both in silico and on de novo prediction tasks.

[0047] Given the shortcomings of methods based on manually engineered protein surface features or computationally intensive docking studies, recent approaches have turned to deep learning as a means of integrating structural data directly into the training of epitope prediction models.

[0048] Disclosed herein are predictive models that are trained and / or deployed to identify an epitope, such as a B-cell epitope, a T-cell epitope, or others. In some embodiments, the predictive models disclosed herein and trained and / or deployed to identify a paratope. In some embodiments, the predictive models disclosed herein and trained and / or deployed to identify an epitope, such as a B-cell epitope, a T-cell epitope, or others, and a paratope.

[0049] In some embodiments, the predictive models disclosed herein analyze interactions between a protein, such as an antibody, and its antigen. In some embodiments, the predictive models disclosed herein treat interactions between antibody and antigen amino acid residues analogously to user-item interactions in recommender systems based on collaborative filtering.In some embodiments, the task of predicting interactions between ^ users may be described as^ = {^^, ^^, … , ^^} and ^ items ^ = {^^, ^^, … , ^^}. In some embodiments, known userpreferences are encoded by an interaction matrix ^ ∈ ^^×^ where ^^^ = 1 if user ^^ interactswith item ^^ , and ^^^ = 0 otherwise. In some embodiments, the interaction matrix is interpretedas the adjacency matrix of a bipartite graph ^ = {^ ∪ ^, ^}, where the edges ^ = {(^, ^) ∶ ^^^ =1} correspond to user-item interactions. In some embodiments, graph-basedfiltering may be described as the problem of inferring this graph topology for unknown user-item 11 IPTS / 128951704.1Attorney Docket No.: SES-014WO PATENTinteractions based on observed data: ^^^ = (^, Θ), where is a learnable function and Θrepresents the model parameters.

[0050] In some embiments, the models disclosed herein predict epitopes, such as a B-cell epitope, a T-cell epitope, or others, using the antibody-agnostic model, in which all potential surface patches on a target protein that may be susceptible to binding by antibodies are identified. In some embiments, the predictive models disclosed herein predict epitopes, such as a B-cell epitope, a T-cell epitope, or others, using an antibody-specific model, in which the specific epitope targeted by an antibody that is known to bind the protein is identified. In some embodiments, all involved proteins are represented by their structure, which may be experimentally determined or predicted using an algorithm.

[0051] In some embodiments, the bias arising from conformational changes in the antibody and / or antigen upon complex formation is avoided by first separating the antibody and antigen into monomeric structures and moving the sidechains into a conformation with minimized free energy. These relaxed structures may be cast as graphs whose nodes correspond to amino acid residues, while edges correspond to peptide bonds or close contact between residues. In some embodiments, a node feature tensor is constructed by aggregating the atomic information for each residue’s sidechain. In some embodiments, the aggregation of the atomic information foreach residue’s sidechain is achieved by, for each atom " belonging to the sidechain of a residueρ, constructing a feature vector ℎ% encoding the atom type, distance from the residue’s alphacarbon, and curvature (calculated as the discrete mean curvature of a sphere centered at the atom). In some embodiments, a residue-level sidechain representation is then constructed by aggregating these feature vectors over all atoms belonging to the residue’s sidechain &': *1= / where φ is a learned multi-layer some the remaining residue node features encode its 3D coordinates, secondary structure, and standard amino acid properties derived from reduced dimensional embeddings of physicochemical measurements of amino acids. In some embodiments, the edge features encode the distances and angles between residue alpha-carbons, as well as characterization of the residue interaction type as a peptide bond, 12 IPTS / 128951704.1Attorney Docket No.: SES-014WO PATENT backbone carbonyl interaction, salt bridge, pi-pi interaction, hydrogen bond, or ionic interaction (FIG. 6, panel B).

[0052] In some embodiments, the disclosed herein representations of antigen and antibody structures are used to derive feature embeddings for the residue nodes using message-passing graph neural networks. In some embodiments, the basic operation of the GNNs is the graphattentional operator. In some embodiments, the node feature vectors are descibred as0 , 2^ 0^, … , 01 ∈ ^ , where 3 is the number of nodes and 4 is the number of node features, and aweight matrix Θ ∈ ^5×2, where ^ is the number of edge features, which is learned by a multi-layer perceptron. In some embodiments, a message passing via node and edge updates identifies key elements within input structure data that are predictive of epitope residues, and is performed as following: 0′^ = α^,^Θ80^ 9 . α^,^ Θ;0^

[0053] In some embodiments,of neighbors of node ^, and the attention weights are computed by

[0054] In some embodiments, the attention mechanism is a single-layer feedforward neural network parametrized by the weight vector and <^,^are edge features.

[0055] In some embodiments, theoperations are applied iteratively and independently to the antibody and antigen feature graphs (or antigen alone in the case ofantibody-agnostic epitope prediction) to produce embeddings ^ = {^ >^, ^^, … , ^^} ⊂ ^ and^ = {^ , >^ ^^, … , ^^} ⊂ ^ of the antibody and antigen nodes, respectively.

[0056] In some embodiments, for antibody-specific epitope predictions, once representations have been obtained, a second graph attention network (GAT) is trained to predict the edge weights of a complete bipartite graph with disjoint vertex sets ^ and ^, akin to a user-item graph in typical recommender systems as described above. In some embodiments, following the final iteration of message passing over the antibody-antigen bipartite graph with predicted edge weights (or antigen graph alone for antibody-agnostic predictions), a single-layer perceptron with 13 IPTS / 128951704.1Attorney Docket No.: SES-014WO PATENT one-dimensional output and hyperbolic tangent activation is applied to each fully updated node feature tensor to obtain the final model output in the form of node-level probabilities for theantigen ρ@ @ @^, ρ^, … , ρ^ ∈ A0,1B and antibody CD D D^ , C^ , … , ρ^ ∈ A0,1B. In some embodiments, ρ@^ =EH is interpreted as the probability that ^ in the antigen belongs to the epitope atantibody binds, and ρD = EHJ as the that residue ^ in the antibody belongs toparatope of this interaction.

[0057] In some the epitope prediction task (either antibody agnostic or antibody specific) results in a small number of annotated epitope residues out of hundreds to thousands of residues constituting a given antigen. In some embodiments, not annotated residues may indeed belong to an epitope for an antibody whose interaction with the antigen has yet to be characterized. Thus, the antibody-agnostic model is primarily has positive predictive power. Training the Predictive Model

[0058] In some embodiments, the predictive model is trained on annotated epitopes, such as B- cell epitopes, T-cell epitopes, or others, derived from a database comprising protein, antibody, and / or antigen structures. In some embodiments, the predictive model is constructed using the PyTorch Geometric library for graph neural networks. In some embodiments, Bayesian hyperparameter optimization is applied to identify the ideal values for the training batch size, number of heads for the multi-headed attention operations, number of hidden dimensions, number of graph convolutional layers, and learning rate using the aforementioned validation set. In some embodiments, the identification of the ideal values for training is conducted automatically using the Sweeps functionality within the Weights & Biases platform.

[0059] In some embodiments, the machine learning module can be trained using techniques such as unsupervised, supervised, semi-supervised, reinforcement learning, transfer learning, incremental learning, curriculum learning techniques, and / or learning to learn. Training typically occurs after selection and development of a machine learning module and before the machine learning module is operably in use. In one aspect, the training data used to teach the machine learning module can comprise input data such as omics images and the respective target output data such as gene expression profiles.

[0060] In some embodiments, unsupervised learning is implemented. Unsupervised learning can involve providing all or a portion of unlabeled training data to a machine learning module. 14 IPTS / 128951704.1Attorney Docket No.: SES-014WO PATENT The machine learning module can then determine one or more outputs implicitly based on the provided unlabeled training data. In some embodiments, supervised learning is implemented. Supervised learning can involve providing all or a portion of labeled training data to a machine learning module, with the machine learning module determining one or more outputs based on the provided labeled training data, and the outputs are either accepted or corrected depending on the agreement to the actual outcome of the training data. In some examples, supervised learning of machine learning system(s) can be governed by a set of rules and / or a set of labels for the training input, and the set of rules and / or set of labels may be used to correct inferences of a machine learning module.

[0061] In some embodiments, semi-supervised learning is implemented. Semi-supervised learning can involve providing all or a portion of training data that is partially label to a machine learning module. During semi-supervised learning, supervised learning is used for a portion of labeled training data, and unsupervised learning is used for a portion of unlabeled training data. In some embodiments, reinforcement learning is implemented. Reinforcement learning can involve first providing all or a portion of the training data to a machine learning module and as the machine learning module produces an output, the machine learning module receives a “reward” signal in response to a correct output. Typically, the reward signal is a numerical value and the machine learning module is developed to maximize the numerical value of the reward signal. In addition, reinforcement learning can adopt a value function that provides a numerical value representing an expected total of the numerical values provided by the reward signal over time.

[0062] In some embodiments, transfer learning is implemented. Transfer learning techniques can involve providing all or a portion of a first training data to a machine learning module, then, after training on the first training data, providing all or a portion of a second training data. In some embodiments, a first machine learning module can be pre-trained on data from one or more computing devices. The first trained machine learning module is then provided to a computing device, where the computing device is intended to execute the first trained machine learning model to produce an output. Then, during the second training phase, the first trained machine learning model can be additionally trained using additional training data, where the training data can be derived from kernel and non-kernel data of one or more computing devices. This second training of the machine learning module and / or the first trained machine learning model using 15 IPTS / 128951704.1Attorney Docket No.: SES-014WO PATENT the training data can be performed using either supervised, unsupervised, or semi-supervised learning. In addition, it is understood transfer learning techniques can involve one, two, three, or more training attempts. Once the machine learning module has been trained on at least the training data, the training phase can be completed. The resulting trained machine learning model can be utilized as at least one of trained machine learning module.

[0063] In some embodiments, incremental learning is implemented. Incremental learning techniques can involve providing a trained machine learning module with input data that is used to continuously extend the knowledge of the trained machine learning module. Another machine learning training technique is curriculum learning, which can involve training the machine learning module with training data arranged in a particular order, such as providing relatively easy training examples first, then proceeding with progressively more difficult training examples. As the name suggests, difficulty of training data is analogous to a curriculum or course of study at a school.

[0064] In some embodiments, learning to learn is implemented. Learning to learn, or meta- learning, comprises, in general, two levels of learning: quick learning of a single task and slower learning across many tasks. For example, a machine learning module is first trained and comprises of a first set of parameters or weights. During or after operation of the first trained machine learning module, the parameters or weights are adjusted by the machine learning module. This process occurs iteratively on the success of the machine learning module. In another example, an optimizer, or another machine learning module, is used wherein the output of a first trained machine learning module is fed to an optimizer that constantly learns and returns the final results. Other techniques for training the machine learning module and / or trained machine learning module are possible as well.

[0065] In some examples, after the training phase has been completed but before producing predictions expressed as outputs, a trained machine learning module can be provided to a computing device where a trained machine learning module is not already resident, in other words, after training phase has been completed, the trained machine learning module can be downloaded to a computing device. For example, a first computing device storing a trained machine learning module can provide the trained machine learning module to a second computing device. Providing a trained machine learning module to the second computing device may comprise one or more of communicating a copy of trained machine learning module to the 16 IPTS / 128951704.1Attorney Docket No.: SES-014WO PATENT second computing device, making a copy of trained machine learning module for the second computing device, providing access to trained machine learning module to the second computing device, and / or otherwise providing the trained machine learning system to the second computing device. In some embodiments, a trained machine learning module can be used by the second computing device immediately after being provided by the first computing device. In some examples, after a trained machine learning module is provided to the second computing device, the trained machine learning module can be installed and / or otherwise prepared for use before the trained machine learning module can be used by the second computing device.

[0066] After a machine learning model has been trained it can be used to output, estimate, infer, predict, or determine a result. A trained machine learning module can receive input data and operably generate a result, such as output, estimation, inference, prediction, or determination. As such, the input data can be used as an input to the trained machine learning module for providing corresponding results to kernel components and non-kernel components. For example, a trained machine learning module can generate results in response to requests. In some embodiments, a trained machine learning module can be executed by a portion of other software. For example, a trained machine learning module can be executed by a result daemon to be readily available to provide results upon request.

[0067] In some embodiments, the machine learning model comprises linear classifiers, logistic classifiers, Bayesian networks, random forest, neural networks, graph neural networks (GNN), message-passing graph neural networks (MPGNN), matrix factorization, hidden Markov model, support vector machine, K-means clustering, or K-nearest neighbor. In some embodiments, the machine learning model comprises a neural network. In some embodiments, the neural network is a convolutional neural network. In some embodiment, the convolutional neural network is a neural network, for example a graph neural network (GNN). In some embodiment, the convolutional neural network is a message-passing graph neural network (MPGNN).

[0068] In some embodiments, the machine learning model comprises embedding. In some embodiments, the training machine learning model is trained with protein data as an input. In some embodiments, the protein data is obtained from a database, or a sample. In some embodiments, the protein data is analyzed to identify loop-mediated protein-protein interactions. Thus, in some embodiments, the protein analyzed to identify loop-mediated protein-protein 17 IPTS / 128951704.1Attorney Docket No.: SES-014WO PATENT interactions is a training data. In some embodiments, the trained machine learning model is trained with the training data.

[0069] In some embodiments, the annotated epitopes, such as B-cell epitopes, T-cell epitopes, or others, are derived from the Immune Epitope Database and antibody-antigen complex structures from the Structural Antibody Database (as of May 2024, or any other available version). In some embodiments, structures are filtered to only retain complexes of antibodies bound to protein antigens rather than small molecules or other ligands. In some embodiments, and without limitation, structures are filtered to only retain complexes of antibodies bound to protein antigens rather than small molecules or other ligands, for example resulting in a total of 4420 structures containing 1897 unique protein antigens. Ultimately, the models disclose herein make predictions at the level of amino acid residues rather than proteins, and this dataset corresponds to over 4 million antigen amino acid residues and over 7 million antibody residues. In some mebodiments, given the structure of an antibody-antigen complex, the standard definition of the epitope is used as the set of residues in the antigen for which any atom lies within a distance of 4 Angstrom to any atom belonging to a residue in the antibody.

[0070] In some embodiments, any other database, or any combination of databases, may be used.

[0071] In various embodiments, the predictive model achieves a particular AUC performance metric. In various embodiments, the predictive model achieves an AUC of at least 0.60, at least 0.61, at least 0.62, at least 0.63, at least 0.64, at least 0.65, at least 0.66, at least 0.67, at least 0.68, at least 0.69, at least 0.70, at least 0.71, at least 0.72, at least 0.73, at least 0.74, at least 0.75, at least 0.76, at least 0.77, at least 0.78, at least 0.79, at least 0.80, at least 0.81, at least 0.82, at least 0.83, at least 0.84, at least 0.85, at least 0.86, at least 0.87, at least 0.88, at least 0.89, at least 0.90, at least 0.91, at least 0.92, at least 0.93, at least 0.94, at least 0.95, at least 0.96, at least 0.97, at least 0.98, or at least 0.99. In various embodiments, the predictive model achieves an AUC of at least 0.60. In various embodiments, the predictive model achieves an AUC of at least 0.61. In various embodiments, the predictive model achieves an AUC of at least 0.62. In various embodiments, the predictive model achieves an AUC of at least 0.63. In various embodiments, the predictive model achieves an AUC of at least 0.64. In various embodiments, the predictive model achieves an AUC of at least 0.65. In various embodiments, the predictive model achieves an AUC of at least 0.66. In various embodiments, the predictive model achieves 18 IPTS / 128951704.1Attorney Docket No.: SES-014WO PATENT an AUC of at least 0.67. In various embodiments, the predictive model achieves an AUC of at least 0.68. In various embodiments, the predictive model achieves an AUC of at least 0.69. In various embodiments, the predictive model achieves an AUC of at least 0.70. In various embodiments, the predictive model achieves an AUC of at least 0.71. In various embodiments, the predictive model achieves an AUC of at least 0.72. In various embodiments, the predictive model achieves an AUC of at least 0.73. In various embodiments, the predictive model achieves an AUC of at least 0.74. In various embodiments, the predictive model achieves an AUC of at least 0.75. In various embodiments, the predictive model achieves an AUC of at least 0.76. In various embodiments, the predictive model achieves an AUC of at least 0.77. In various embodiments, the predictive model achieves an AUC of at least 0.78. In various embodiments, the predictive model achieves an AUC of at least 0.79. In various embodiments, the predictive model achieves an AUC of at least 0.80. In various embodiments, the predictive model achieves an AUC of at least 0.81. In various embodiments, the predictive model achieves an AUC of at least 0.82. In various embodiments, the predictive model achieves an AUC of at least 0.83. In various embodiments, the predictive model achieves an AUC of at least 0.84. In various embodiments, the predictive model achieves an AUC of at least 0.85. In various embodiments, the predictive model achieves an AUC of at least 0.86. In various embodiments, the predictive model achieves an AUC of at least 0.87. In various embodiments, the predictive model achieves an AUC of at least 0.88. In various embodiments, the predictive model achieves an AUC of at least 0.89. In various embodiments, the predictive model achieves an AUC of at least 0.90. In various embodiments, the predictive model achieves an AUC of at least 0.91. In various embodiments, the predictive model achieves an AUC of at least 0.92. In various embodiments, the predictive model achieves an AUC of at least 0.93. In various embodiments, the predictive model achieves an AUC of at least 0.94. In various embodiments, the predictive model achieves an AUC of at least 0.95. In various embodiments, the predictive model achieves an AUC of at least 0.96. In various embodiments, the predictive model achieves an AUC of at least 0.97. In various embodiments, the predictive model achieves an AUC of at least 0.98. In various embodiments, the predictive module achieves an AUC of at least 0.99. Deploying the Predictive Model 19 IPTS / 128951704.1Attorney Docket No.: SES-014WO PATENT

[0072] Reference is made to FIG. 2A, which provides a flowchart with steps to predict an epitope of a protein of interest, specifically the antibody-agnostic prediction model, in accordance with an embodiment.

[0073] With reference to FIG. 2A, step 205 comprises obtaining or having obtained (e.g., generating from a data set) a graph representation of the protein sequences based on the protein of interest. In some embodiments, step 205 comprises generating the graph comprising a monomeric structure of the protein. In some embodiments, step 205 comprises generating a monomeric structure of the protein comprises generating a conformational variant of the monomeric structure of the protein, wherein each sidechain of the protein is in a minimized free energy conformation. In some embodiments, step 205 comprises characterizing one or more structural features of the protein by performing a structural analysis of individual amino acids of the protein. In some embodiments, step 205 comprises performing a pairwise amino acid to amino acid interaction analysis across at least one pair of amino acids in the protein to generate a feature graph.

[0074] In some embodiments, the graph representation of the protein comprises a plurality of nodes representing each amino acid of the protein, and a plurality of edges between each pair of nodes. In some embodiments, each edge represents a bond between each pair of nodes. In some embodiments, the bond is a peptide bond.

[0075] In some embodiments, the one or more structural features are node features and / or edge features. In some embodiments, the one or more structural features are node features. In some embodiments, the one or more structural features are edge features. As used herein, the nodes of the graph correspond to amino acids in the protein structure, while the edges correspond to structural linkages (e.g., bonds) between amino acids. The strength of the structural linkages between amino acids are represented by assigned weights of the edges. In some embodiments, using the protein structure, an edgeless version of the graph is generated where every amino acid in the protein structure is listed in the graph. In some embodiments, the edge feature comprises presence of at least one interaction between at least one pair of atoms. In some embodiments, the presence of the at least one interaction between the at least one pair of atoms comprises distance and angle between residue alpha-carbons, and / or an energy well. In some embodiments, the at least one interaction is a peptide bond, a backbone carbonyl interaction, a salt bridge, a pi-pi interaction, a hydrogen bond, or an ionic interaction. 20 IPTS / 128951704.1Attorney Docket No.: SES-014WO PATENT

[0076] In some embodiments, generating node features comprises assigning node vector values to each amino acid according to the atom type, distance from the amino acid’s alpha carbon, and curvature information. In some embodiments, the node vector values are aggregated to generate one or more structural features. In some embodiments, the node feature comprises a 3D coordinate of the amino acid. In some embodiments, the node feature comprises a secondary structure of the amin acid. In some embodiments, the node feature comprises standard properties of the amino acid. In some embodiments, the node feature comprises a 3D coordinate of the amino acid, a secondary structure of the amin acid, and standard properties of the amino acid.

[0077] In some embodiments, the feature graph is determined by performing a transform of one or more node features and / or edge features as provided herein. In some embodiments, the transform is a neural network, for example a graph neural network (GNN), such as those known in the art. In some embodiments, the graph neural network (GNN) is a message-passing graph neural network (MPGNN).

[0078] With reference to FIG. 2A, step 210 comprises providing the graph representation to a generalized machine learning algorithm model to generate a plurality of predicted probabilities, wherein the generalized machine learning algorithm model has been trained using a non-curated training data set, as provided herein, for example in FIG. 2B. Step 215 comprises classifying each of one or more amino acids of the protein as an epitope-member residue or not an epitope- member residue based on the plurality of predicted probabilities, as provided herein. Step 220 comprises generating an epitope prediction, such as a B-cell epitope, a T-cell epitope, or other epitope prediction comprising the classified epitope-member residues, as provided herein.

[0079] With reference to FIG. 2C, step 230 comprises providing the graph representation to a generalized machine learning algorithm model to generate a plurality of predicted probabilities, wherein the generalized machine learning algorithm model is trained using a training data set comprising at least 1,000 loop‑mediated protein‑protein interaction interfaces, as provided herein, for example in FIG. 2D. Step 235 comprises classifying each of one or more amino acids of the protein as an epitope-member residue or not an epitope-member residue based on the plurality of predicted probabilities, as provided herein. Step 240 comprises generating an epitope prediction, such as a B-cell epitope, a T-cell epitope, or other epitope prediction comprising the classified epitope-member residues, as provided herein. In some embodiments, the generalized machine learning algorithm model of step 230 trained on the trianing data is 21 IPTS / 128951704.1Attorney Docket No.: SES-014WO PATENT further refined for antibody-antigen interactions. In some embodiments, the generalized machine learning algorithm model of step 230 trained on the trianing data comprising the loop‑mediated protein‑protein interaction interfaces is further refined for antibody-antigen interactions. In some embodiments, the generalized machine learning algorithm model of step 230 trained on the trianing data comprising at least 1000 the loop‑mediated protein‑protein interaction interfaces is further refined for antibody-antigen interactions. In some embodiments, the generalized machine learning algorithm model of step 230 trained on the trianing data comprising at least 10000 the loop‑mediated protein‑protein interaction interfaces is further refined for antibody-antigen interactions.

[0080] With reference to FIG. 2B, data set of step 210 comprises a plurality of training epitope sequences of a plurality of training proteins 211, and labels for the plurality of training epitope sequences, wherein a label identifies whether a corresponding training epitope sequence was bound by a corresponding binder of a plurality of binders 212. In some embodiments, the plurality of training proteins comprises at least 500 different proteins. In some embodiments, the plurality of training proteins comprises at least 550 different proteins. In some embodiments, the plurality of training proteins comprises at least 600 different proteins. In some embodiments, the plurality of training proteins comprises at least 650 different proteins. In some embodiments, the plurality of training proteins comprises at least 700 different proteins. In some embodiments, the plurality of training proteins comprises at least 750 different proteins. In some embodiments, the plurality of training proteins comprises at least 800 different proteins. In some embodiments, the plurality of training proteins comprises at least 850 different proteins. In some embodiments, the plurality of training proteins comprises at least 900 different proteins. In some embodiments, the plurality of training proteins comprises at least 500 different binders. In some embodiments, the plurality of training proteins comprises at least 550 different binders. In some embodiments, the plurality of training proteins comprises at least 600 different binders. In some embodiments, the plurality of training proteins comprises at least 650 different binders. In some embodiments, the plurality of training proteins comprises at least 700 different binders. In some embodiments, the plurality of training proteins comprises at least 800 different proteins, wherein the plurality of binders comprises at least 700 different binders.

[0081] With reference to FIG. 2D, data set of step 230 comprises a plurality of training proteins 231, and labels for the plurality of training proteins, wherein a label identifies loop-mediated 22 IPTS / 128951704.1Attorney Docket No.: SES-014WO PATENT protein-protein interactions 232. In some embodiments, the plurality of training proteins comprises at least 500 different proteins. In some embodiments, the plurality of training proteins comprises at least 550 different proteins. In some embodiments, the plurality of training proteins comprises at least 600 different proteins. In some embodiments, the plurality of training proteins comprises at least 650 different proteins. In some embodiments, the plurality of training proteins comprises at least 700 different proteins. In some embodiments, the plurality of training proteins comprises at least 750 different proteins. In some embodiments, the plurality of training proteins comprises at least 800 different proteins. In some embodiments, the plurality of training proteins comprises at least 850 different proteins. In some embodiments, the plurality of training proteins comprises at least 900 different proteins. In some embodiments, the plurality of training proteins comprises at least 950 different proteins. In some embodiments, the plurality of training proteins comprises at least 1000 different proteins. In some embodiments, the plurality of training proteins comprises at least 1100 different proteins. In some embodiments, the plurality of training proteins comprises at least 1200 different proteins. In some embodiments, the plurality of training proteins comprises at least 1300 different proteins. In some embodiments, the plurality of training proteins comprises at least 1400 different proteins. In some embodiments, the plurality of training proteins comprises at least 1500 different proteins. In some embodiments, the plurality of training proteins comprises at least 1600 different proteins. In some embodiments, the plurality of training proteins comprises at least 1700 different proteins. In some embodiments, the plurality of training proteins comprises at least 1800 different proteins. In some embodiments, the plurality of training proteins comprises at least 1900 different proteins. In some embodiments, the plurality of training proteins comprises at least 2000 different proteins. In some embodiments, the plurality of training proteins comprises at least 2100 different proteins. In some embodiments, the plurality of training proteins comprises at least 2200 different proteins. In some embodiments, the plurality of training proteins comprises at least 2300 different proteins. In some embodiments, the plurality of training proteins comprises at least 2400 different proteins. In some embodiments, the plurality of training proteins comprises at least 2500 different proteins. In some embodiments, the plurality of training proteins comprises at least 2600 different proteins. In some embodiments, the plurality of training proteins comprises at least 2700 different proteins. In some embodiments, the plurality of training proteins comprises at least 2800 different proteins. In some 23 IPTS / 128951704.1Attorney Docket No.: SES-014WO PATENT embodiments, the plurality of training proteins comprises at least 2900 different proteins. In some embodiments, the plurality of training proteins comprises at least 3000 different proteins. In some embodiments, the plurality of training proteins comprises at least 3100 different proteins. In some embodiments, the plurality of training proteins comprises at least 3200 different proteins. In some embodiments, the plurality of training proteins comprises at least 3300 different proteins. In some embodiments, the plurality of training proteins comprises at least 3400 different proteins. In some embodiments, the plurality of training proteins comprises at least 3500 different proteins. In some embodiments, the plurality of training proteins comprises at least 3600 different proteins. In some embodiments, the plurality of training proteins comprises at least 3700 different proteins. In some embodiments, the plurality of training proteins comprises at least 3800 different proteins. In some embodiments, the plurality of training proteins comprises at least 3900 different proteins. In some embodiments, the plurality of training proteins comprises at least 4000 different proteins. In some embodiments, the plurality of training proteins comprises at least 4100 different proteins. In some embodiments, the plurality of training proteins comprises at least 4200 different proteins. In some embodiments, the plurality of training proteins comprises at least 4300 different proteins. In some embodiments, the plurality of training proteins comprises at least 4400 different proteins. In some embodiments, the plurality of training proteins comprises at least 4500 different proteins. In some embodiments, the plurality of training proteins comprises at least 4600 different proteins. In some embodiments, the plurality of training proteins comprises at least 4700 different proteins. In some embodiments, the plurality of training proteins comprises at least 4800 different proteins. In some embodiments, the plurality of training proteins comprises at least 4900 different proteins. In some embodiments, the plurality of training proteins comprises at least 5000 different proteins. In some embodiments, the plurality of training proteins comprises at least 5100 different proteins. In some embodiments, the plurality of training proteins comprises at least 5200 different proteins. In some embodiments, the plurality of training proteins comprises at least 5300 different proteins. In some embodiments, the plurality of training proteins comprises at least 5400 different proteins. In some embodiments, the plurality of training proteins comprises at least 5500 different proteins. In some embodiments, the plurality of training proteins comprises at least 5600 different proteins. In some embodiments, the plurality of training proteins comprises at least 5700 different 24 IPTS / 128951704.1Attorney Docket No.: SES-014WO PATENT proteins. In some embodiments, the plurality of training proteins comprises at least 5800 different proteins. In some embodiments, the plurality of training proteins comprises at least 5900 different proteins. In some embodiments, the plurality of training proteins comprises at least 6000 different proteins. In some embodiments, the plurality of training proteins comprises at least 6100 different proteins. In some embodiments, the plurality of training proteins comprises at least 6200 different proteins. In some embodiments, the plurality of training proteins comprises at least 6300 different proteins. In some embodiments, the plurality of training proteins comprises at least 6400 different proteins. In some embodiments, the plurality of training proteins comprises at least 6500 different proteins. In some embodiments, the plurality of training proteins comprises at least 6600 different proteins. In some embodiments, the plurality of training proteins comprises at least 6700 different proteins. In some embodiments, the plurality of training proteins comprises at least 6800 different proteins. In some embodiments, the plurality of training proteins comprises at least 6900 different proteins. In some embodiments, the plurality of training proteins comprises at least 7000 different proteins. In some embodiments, the plurality of training proteins comprises at least 7100 different proteins. In some embodiments, the plurality of training proteins comprises at least 7200 different proteins. In some embodiments, the plurality of training proteins comprises at least 7300 different proteins. In some embodiments, the plurality of training proteins comprises at least 7400 different proteins. In some embodiments, the plurality of training proteins comprises at least 7500 different proteins. In some embodiments, the plurality of training proteins comprises at least 7600 different proteins. In some embodiments, the plurality of training proteins comprises at least 7700 different proteins. In some embodiments, the plurality of training proteins comprises at least 7800 different proteins. In some embodiments, the plurality of training proteins comprises at least 7900 different proteins. In some embodiments, the plurality of training proteins comprises at least 8000 different proteins. In some embodiments, the plurality of training proteins comprises at least 8100 different proteins. In some embodiments, the plurality of training proteins comprises at least 8200 different proteins. In some embodiments, the plurality of training proteins comprises at least 8300 different proteins. In some embodiments, the plurality of training proteins comprises at least 8400 different proteins. In some embodiments, the plurality of training proteins comprises at least 8500 different proteins. In some embodiments, the plurality of training proteins comprises at 25 IPTS / 128951704.1Attorney Docket No.: SES-014WO PATENT least 8600 different proteins. In some embodiments, the plurality of training proteins comprises at least 8700 different proteins. In some embodiments, the plurality of training proteins comprises at least 8800 different proteins. In some embodiments, the plurality of training proteins comprises at least 8900 different proteins. In some embodiments, the plurality of training proteins comprises at least 9000 different proteins. In some embodiments, the plurality of training proteins comprises at least 9100 different proteins. In some embodiments, the plurality of training proteins comprises at least 9200 different proteins. In some embodiments, the plurality of training proteins comprises at least 9300 different proteins. In some embodiments, the plurality of training proteins comprises at least 9400 different proteins. In some embodiments, the plurality of training proteins comprises at least 9500 different proteins. In some embodiments, the plurality of training proteins comprises at least 9600 different proteins. In some embodiments, the plurality of training proteins comprises at least 9700 different proteins. In some embodiments, the plurality of training proteins comprises at least 9800 different proteins. In some embodiments, the plurality of training proteins comprises at least 9900 different proteins. In some embodiments, the plurality of training proteins comprises at least 10000 different proteins. In some embodiments, the plurality of training proteins comprises at least 10000 different proteins, or more.

[0082] Reference is made to FIG. 3A, which provides a flowchart with steps to predict an epitope of a protein of interest, specifically the antibody-specific prediction model, in accordance with an embodiment.

[0083] With reference to FIG. 3A, step 305 comprises obtaining or having obtained (e.g., generating from a data set) a graph representation of the protein and a graph representation of the binder based on the protein of interest and the antigen of interest. In some embodiments, step 305 comprises generating a protein graph and an antibody graph, each comprising a monomeric structure of the protein and the antibody. In some embodiments, step 305 comprises generating a conformational variant of each the monomeric structure of the protein and the monomeric structure of the antibody, wherein each sidechain of the protein and each sidechain of the antibody is in a minimized free energy conformation. In some embodiments, step 305 comprises characterizing one or more structural features of the protein and the antibody by independently performing a structural analysis of individual amino acids of the protein and the antibody. In some embodiments, step 305 comprises performing a pairwise amino acid to amino acid 26 IPTS / 128951704.1Attorney Docket No.: SES-014WO PATENT interaction analysis across at least one pair of amino acids in the protein and the antibody to generate an feature graph for each the protein and the antibody.

[0084] In some embodimetns, the protein graph and the antibody graph each comprises a plurality of nodes representing each amino acid of the protein and the antibody, and a plurality of edges between each pair of nodes. In some embodiments, each edge represents a bond between each pair of nodes. In some embodiments, the bond is a peptide bond.

[0085] Step 310 comprises transforming the obtained graph representation of the protein and graph representation of the binder to construct a bipartite graph, as provided herein, for example in FIG. 3B. Step 315 comprises providing the bipartite graph to a machine learning model to generate a plurality of predicted probabilities. In some embodimetns, the machine learning model of step 315 comprises unsupervised learning, supervised learning, semi-supervised learning, reinforcement learning, transfer learning, incremental learning, curriculum learning, and learning to learn. In some embodimetns, the machine learning model of step 315 comprises linear classifiers, logistic classifiers, Bayesian networks, random forest, neural networks, graph neural networks (GNN), message-passing graph neural networks (MPGNN), matrix factorization, hidden Markov model, support vector machine, K-means clustering, or K-nearest neighbor. In some embodimetns, the machine learning model of step 315 is trained using a training data set comprising at least 1,000 loop‑mediated protein‑protein interaction interfaces. In some embodiments, the training data set comprises at least 10000 loop‑mediated protein‑protein interaction interfaces. In some embodiments, the machine learning model trained using the training data is further refined for antibody-antigen interactions. Step 320 comprises classifying each amino acid of the protein as an epitope-member residue or not an epitope-member residue, and each amino acid of the binder as a paratope-member or not a paratope-member based on the plurality of predicted probabilities, as provided herein. Step 325 comprises identifying the paratope comprising one or more amino acid residues in the protein of interest associated with measures of non-structural binding relevance, as provided herein.

[0086] With reference to FIG. 3B, data set of step 310 comprises a plurality of protein embeddings 311, a plurality of binder embeddings 312, and a plurality of edges connecting protein embeddings and binder embeddings 312. In some embodiments, the plurality of edges comprise predicted weights determined through graph-based collaborative filtering. 27 IPTS / 128951704.1Attorney Docket No.: SES-014WO PATENT

[0087] In some embodiments, for the antibody-agnostic prediction model, an antigen ^ is takenas a collection of residues ^ = (K^, K^, … , K1), each with a corresponding binary labelcorresponding to whether it belongs to a B cell epitope or not. In some embodiments, for theantibody-specific prediction model, the examples are pairs (^, L), where ^ is an antigen and L isan antibody specifically targeting that antigen. In some embodiments, each A and B is represented as a collection of residues, where the label of a residue is 1 if and only if the residue belongs to the epitope or paratope for that specific antibody-antigen interaction. In some embodiments, a weighted cross-entropy loss function for the antibody-specific epitope predictor is used to account for the class imbalance. In some embodiments, to mitigate the class imbalance for the antibody-agnostic model and incur a lower penalty for false positive predictions, an in- batch contrastive loss is used as follows: 1U1U1U1W2 O 1ℒ = . . ℎS O S^ ^ T−. . ^ ℎ^ Twhere Q is atheembedding vector of positively resp negatively labeled residue ^, 3O is the number of positiveexamples (annotated epitope residues) and 3Sis the number of negative examples. In some embodiments, this loss deliberately separates the node embeddings of positive examples from those of negative examples to prevent overestimation of model accuracy due to class imbalance. Implementation of the Epitope Predictive Model

[0088] Without being bound by a particular theory, epitopes, such as a B-cell epitope, a T-cell epitope, or other epitopes, are specific regions or fragments of an antigen that are recognized and bound by immune cells, such as B-cell receptors (BCRs), T-cell receptors (TCRs), or antibodies. These epitopes can be composed of linear sequences of amino acids (linear epitopes) or three- dimensional conformations of the antigen's surface (conformational epitopes). When components of the adaptive immune system encounter an antigen, specific immune receptors—whether membrane‑bound antibodies (B‑cell receptors), soluble antibodies, or T‑cell receptors engaging peptide‑loaded major histocompatibility complexes (pMHC)—bind discrete molecular regions known as epitopes. Receptor engagement triggers cellular activation, clonal expansion, and 28 IPTS / 128951704.1Attorney Docket No.: SES-014WO PATENT differentiation into effector and memory populations that together drive immediate protection and long‑term immunity.

[0089] The goal of epitope prediction, irrespective of epitope class, is to pinpoint antigenic regions most likely to be recognized by immune receptors for use in vaccine design, immunomonitoring, therapeutic antibody or T‑cell engineering, and structure–function studies. Prediction frameworks must accommodate diverse epitope modalities: conformational or linear surfaces accessible to antibodies; processed linear peptides that bind specific HLA alleles for T‑cell recognition; and post‑translationally modified or neoantigenic determinants that arise in cancer or infection. Accurate mapping informs diagnostics, guides immunotherapy development, and accelerates strategies to prevent or treat infectious diseases, malignancies, and autoimmune disorders.

[0090] Disclosed herein are methods for predicting epitopes of a protein. Reference is now made to FIG.^1B, which depicts an exemplary block diagram for identifying epitopes from a protein sequence or structure using a unified epitope‑prediction model. As shown, FIG.^1B introduces a protein of interest^150, a Feature Graph module^160 that captures structural, biochemical, and processing attributes relevant to both antibody and T‑cell recognition, and an Epitope Predictive module^170 that outputs ranked candidate regions likely to serve as functional epitopes across immune contexts.

[0091] In some embodiments, when B-cells encounter an antigen, the BCR binds to the epitope, triggering the activation of the B-cell. This activation leads to the proliferation and differentiation of B-cells into plasma cells, which produce antibodies specific to the antigen, and memory B-cells, which provide long-term immunity.

[0092] The aim of B-cell epitope prediction is to aid in the identification of B-cell epitopes for practical purposes such as replacing the antigen for antibody production or conducting structure- function studies. Any region of the antigen that is exposed to solvents can be recognized by antibodies. Antibodies recognizing linear B-cell epitopes can recognize denatured antigens, while denaturing the antigen results in the loss of recognition for conformational B-cell epitopes. The majority of B-cell epitopes (approximately 90%) are conformational, and in fact, only a minority of native antigens contain linear B-cell epitopes.

[0093] B-cell epitopes play a role in the immune response and vaccine design. Understanding B-cell epitopes also aids in designing therapeutic antibodies for treating infectious diseases, 29 IPTS / 128951704.1Attorney Docket No.: SES-014WO PATENT cancers, and autoimmune disorders. Furthermore, epitope mapping is essential for diagnosing diseases and monitoring immune responses, as it helps in identifying which parts of an antigen are recognized by the immune system. This knowledge is vital for advancing immunotherapies and improving strategies for disease prevention and treatment.

[0094] Accordingly, disclosed herein are methods for predicting a B-cell epitope of a protein. Reference is now made to FIG. 1B, which depicts an exemplary block diagram for determining an epitope of a protein (such as a protein sequence, or a protein structure), specifically the antibody-agonistic model, in accordance with an embodiment. As one example, FIG. 1B introduces a protein of interest 150, a Feature Graph module 160, and an Epitope Predictive module 170.

[0095] In other embodiments, when T‑cells encounter antigen‑presenting cells displaying a peptide‑loaded major histocompatibility complex (pMHC), the T‑cell receptor (TCR) engages the peptide–MHC complex, triggering T‑cell activation. This activation initiates clonal expansion and differentiation into effector T‑cells—such as cytotoxic CD8⁺ T‑cells that lyse infected or malignant cells, and helper CD4⁺ T‑cells that coordinate broader immune responses—as well as memory T‑cells that confer long‑term cellular immunity.

[0096] The objective of T‑cell epitope prediction is to identify peptide sequences that can be processed, presented by specific HLA alleles, and recognized by TCRs for applications in vaccine design, immunomonitoring, and adoptive T‑cell therapies. Because TCR recognition is restricted to linear peptides bound within the MHC groove (typically 8–11 amino acids for class^I and 13–18 amino acids for class^II molecules), accurate prediction requires accounting for antigen processing, peptide binding affinity, and HLA polymorphism. Mapping T‑cell epitopes is pivotal for developing peptide and nucleic‑acid vaccines, engineering TCR‑ or CAR‑T therapeutics against cancers and infectious diseases, and diagnosing or modulating autoimmune conditions by pinpointing pathogenic peptides.

[0097] Accordingly, disclosed herein are methods for predicting a T‑cell epitope of a protein. Reference is now made to FIG.^1B, which illustrates an exemplary block diagram for identifying epitopes (e.g., from a protein sequence or structure) using a T‑cell agonistic model in accordance with one embodiment. As shown, FIG.^1B introduces a protein of interest^150, a Feature Graph module^160 configured to model antigen processing and HLA binding features, and a T‑Cell 30 IPTS / 128951704.1Attorney Docket No.: SES-014WO PATENT Epitope Predictive module^170 that outputs ranked candidate peptides likely to be recognized by TCRs.

[0098] In various embodiments, a protein of interest 150 can be a naturally occurring protein or a synthetic protein. In various embodiments, the protein of interest can be synthesized using any methods of protein synthesis including without limitation solid-phase peptide synthesis (SPPS), liquid-phase peptide synthesis (LPPS), recombinant DNA technology, cell-free protein synthesis (CFPS), native chemical ligation (NCL), expressed protein ligation (EPL), mirror image phage display, and any other methods of protein synthesis known to one skilled in the art.

[0099] The Feature Graph module 160 generally analyzes the sequence of a protein of interest 150 to determine the feature tensor of each amino acid of the protein sequence. For example, the Feature Graph model 160 may take the monomeric structure of the protein of interest 150 and use an algorithm to move the sidechains into a conformation with minimized free energy. The Feature Graph model 160 may then aggregate the atomic information for each amino acid residue’s sidechain. In some embodiment, for each atom a belonging to the sidechain of a residue ρ a feature vector h_a encoding the atom type, distance from the residue’s alpha carbon, and curvature (calculated as the discrete mean curvature of a sphere centered at the atom) may be constructed. In some embodiments, the remaining residue node features encode the 3D coordinates of the amino acid residue, secondary structure, and standard amino acid properties derived from reduced dimensional embeddings of physicochemical measurements of amino acids. In some embodiments, the edge features encode the distances and angles between residue alpha-carbons, as well as characterization of the residue interaction type as a peptide bond, backbone carbonyl interaction, salt bridge, pi-pi interaction, hydrogen bond, or ionic interaction. In some embodiments, the Feature Graph module 160 aggregates these feature vectors to arrive at a residue-level sidechain representation.

[0100] The epitope predictive module 170 in FIG. 1B may conduct steps for identifying an epitope. For example, the epitope predictive model 170 may derive feature embeddings for the amino acid residue nodes in a feature graph using message-passing graph neural networks (GNNs). In some embodiments, the graph neural network (GNN) is a graph attentional operator. In some embodiments, the epitope predictive model 170 may apply a single-layer perceptron with one-dimensional output and hyperbolic tangent activation to each fully updated node feature tensor to obtain the final model output in the form of node-level probabilities for the protein of 31 IPTS / 128951704.1Attorney Docket No.: SES-014WO PATENT interest. In some embodiments, the epitope predictive model 170 may take the protein of interest A as a collection of residues A=(r_1,r_2,…,r_N ), each with a corresponding binary label corresponding to whether it belongs to a B-cell epitope or not.

[0101] Reference is now made to FIG. 1C, which depicts an exemplary block diagram for determining an epitope of a protein (such as a protein sequence, or a protein structure), such as an antigen, and a paratope of a protein, such as an antibody, specifically the antibody-specific model, in accordance with an embodiment. As one example, FIG. 1C introduces a protein of interest 150, a feature graph module 160, a bipartite recommendation module 165, an epitope- paratope predictive module 175.

[0102] In various embodiments, a protein of interest 150 and 155 can be a naturally occurring protein or a synthetic protein, such as an antgen and an antibody. The Feature Graph module 160 generally analyzes the sequence of the protein of interest 150 and 155 to determine the feature tensor of each amino acid of the protein sequence, as explained herein.

[0103] The bipartite recommendation module 165 may may derive feature embeddings for the amino acid residue nodes in a feature graph using message-passing graph neural networks (GNNs). In some embodiments, the graph neural network (GNN) is a graph attentional operator. In some embodiments, the message-passing operations are applied ieratively and independently to the antibody and antigen feature graphs to produce embeddings of the antibody and antigen nodes. In some embodiments, the bipartite recommendation module 165 may train a second graph attention network (GAT) to predict the edge weights of a complete bipartite graph with disjoint vertex sets U and V.

[0104] The epitope-paratope predictive module 175 in FIG. 1B may conduct steps for identifying an epitope of an antigen and a paratope of an antibody. For example, the epitope- paratope predictive model 175 may apply a single-layer perceptron with one-dimensional output and hyperbolic tangent activation to each fully updated node feature tensor to obtain the final model output in the form of node-level probabilities for the antibody and antigen. In some embodiments, the epitope-paratope predictive model 175 may take the antigen A and the antibody B as pairs (A, B), where A is an antigen and B is an antibody specifically targeting that antigen. Antigen and antibody is each represented as a collection of residues, where the label of a residue is 1 if and only if the residue belongs to the epitope or paratope for that specific antibody-antigen interaction. 32 IPTS / 128951704.1Attorney Docket No.: SES-014WO PATENT Computer Implementation

[0105] The methods disclosed herein, such as the methods of generating a prediction of an epitope, such as a B-cell epitope, or a T-cell epitope, or other, are, in some embodiments, performed on one or more computers. For example, the building and deployment of a predictive model to analyze protein structural data, and database storage can be implemented in hardware or software, or a combination of both. In one embodiment, a machine-readable storage medium is provided, the medium comprising a data storage material encoded with machine readable data which, when using a machine programmed with instructions for using said data, is capable of displaying any of the datasets and execution and results of a predictive model. Such data can be used for a variety of purposes, such as patient monitoring, treatment considerations, and the like. Methods disclosed herein can be implemented in computer programs executing on programmable computers, comprising a processor, a data storage system (including volatile and non-volatile memory and / or storage elements), a graphics adapter, a pointing device, a network adapter, at least one input device, and at least one output device. Program code may be applied to input data to perform the functions described above and generate output information. The output information is applied to one or more output devices, in known fashion. The computer can be, for example, a personal computer, microcomputer, or workstation of conventional design.

[0106] Each program can be implemented in a high level procedural or object oriented programming language to communicate with a computer system. However, the programs can be implemented in assembly or machine language, if desired. In any case, the language can be a compiled or interpreted language. Each such computer program is preferably stored on a storage media or device (e.g., ROM or magnetic diskette) readable by a general or special purpose programmable computer, for configuring and operating the computer when the storage media or device is read by the computer to perform the procedures described herein. The system can also be considered to be implemented as a computer-readable storage medium, configured with a computer program, where the storage medium so configured causes a computer to operate in a specific and predefined manner to perform the functions described herein.

[0107] The signature patterns and databases thereof can be provided in a variety of media to facilitate their use. “Media” refers to a manufacture that contains the signature pattern 33 IPTS / 128951704.1Attorney Docket No.: SES-014WO PATENT information. The databases as described herein can be recorded on computer readable media, e.g. any medium that can be read and accessed directly by a computer. Such media include, but are not limited to: magnetic storage media, such as floppy discs, hard disc storage medium, and magnetic tape; optical storage media such as CD-ROM; electrical storage media such as RAM and ROM; and hybrids of these categories such as magnetic / optical storage media. One of skill in the art can readily appreciate how any of the presently known computer readable mediums can be used to create a manufacture comprising a recording of the present database information. "Recorded" refers to a process for storing information on computer readable medium, using any such methods as known in the art. Any convenient data storage structure can be chosen, based on the means used to access the stored information. A variety of data processor programs and formats can be used for storage, e.g. word processing text file, database format, etc.

[0108] FIG. 6 illustrates an example computer 500 for implementing the predictive models, methods, systems, and data described herein. The computer 600 includes at least one processor 602 coupled to a chipset 604. The chipset 604 includes a memory controller hub 620 and an input / output (I / O) controller hub 622. A memory 606 and a graphics adapter 612 are coupled to the memory controller hub 620, and a display 618 is coupled to the graphics adapter 612. A storage device 608, an input device 614, and network adapter 616 are coupled to the I / O controller hub 622. Other embodiments of the computer 600 have different architectures.

[0109] The storage device 608 is a non-transitory computer-readable storage medium such as a hard drive, compact disk read-only memory (CD-ROM), DVD, or a solid-state memory device. The memory 506 holds instructions and data used by the processor 602. The input device 614 is a touch-screen interface, a mouse, track ball, or other type of pointing device, a keyboard, or some combination thereof, and is used to input data into the computer 600. In some embodiments, the computer 600 may be configured to receive input (e.g., commands) from the input device 614 via gestures from the user. The graphics adapter 612 displays images and other information on the display 618. The network adapter 616 couples the computer 600 to one or more computer networks.

[0110] The computer 600 is adapted to execute computer program modules for providing functionality described herein. As used herein, the term “module” refers to computer program logic used to provide the specified functionality. Thus, a module can be implemented in 34 IPTS / 128951704.1Attorney Docket No.: SES-014WO PATENT hardware, firmware, and / or software. In one embodiment, program modules are stored on the storage device 608, loaded into the memory 606, and executed by the processor 602.

[0111] The types of computers 600 can vary depending upon the embodiment and the processing power required by the entity. For example, the can run in a single computer 600 or multiple computers 600 communicating with each other through a network such as in a server farm. The computers 600 can lack some of the components described above, such as graphics adapters 612, and displays 618.

[0112] Each program can be implemented in a high level procedural or object oriented programming language to communicate with a computer system. However, the programs can be implemented in assembly or machine language, if desired. In any case, the language can be a compiled or interpreted language. Each such computer program is preferably stored on a storage media or device (e.g., ROM or magnetic diskette) readable by a general or special purpose programmable computer, for configuring and operating the computer when the storage media or device is read by the computer to perform the procedures described herein. The system can also be considered to be implemented as a computer-readable storage medium, configured with a computer program, where the storage medium so configured causes a computer to operate in a specific and predefined manner to perform the functions described herein.

[0113] The signature patterns and databases thereof can be provided in a variety of media to facilitate their use. “Media” refers to a manufacture that contains the signature pattern information of the present embodiments. As applciable, the databases of the present embodiments can be recorded on computer readable media, e.g., any medium that can be read and accessed directly by a computer. Such media include, but are not limited to: magnetic storage media, such as floppy discs, hard disc storage medium, and magnetic tape; optical storage media such as CD-ROM; electrical storage media such as RAM and ROM; and hybrids of these categories such as magnetic / optical storage media. One of skill in the art can readily appreciate how any of the presently known computer readable mediums can be used to create a manufacture comprising a recording of the present database information. "Recorded" refers to a process for storing information on computer readable medium, using any such methods as known in the art. Any convenient data storage structure can be chosen, based on the means used to access the stored information. A variety of data processor programs and formats can be used for storage, e.g., word processing text file, database format, etc. 35 IPTS / 128951704.1Attorney Docket No.: SES-014WO PATENT

[0114] In some embodiments, provided herein is a system comprising a non-transitory computer-readable storage medium and a processor, wherein the non-transitory computer- readable storage medium comprises: a) a bipartite graph encoded on the non-transitory computer-readable storage medium; and b) instructions for generating a B-cell epitope prediction comprising a classified epitope-member residues from bipartite graph, wherein the bipartite graph is a transformation of a graph representation of a protein and a graph representation of a binder. In some embodiments, the encoded bipartite graph comprises a plurality of protein embeddings; a plurality of binder embeddings; and a plurality of edges connecting protein embeddings and binder embeddings, wherein the plurality of edges comprise predicted weights determined through graph-based collaborative filtering. In some embodiments, the instructions for generating a B-cell epitope prediction comprise instructions for: a) providing the bipartite graph to a machine learning model to generate a plurality of predicted probabilities; b) classifying each amino acid of the protein as an epitope-member residue or not an epitope- member residue, and each amino acid of the binder as a paratope-member or not a paratope- member based on the plurality of predicted probabilities; and c) generating the B-cell epitope prediction comprising the classified epitope-member residues. In some embodiments, the transformation of the graph representation of the protein and the graph representation of the binder comprises performing a transform of the bipartite graph. In somke embodiments, the transform is any transform provided herein, for example, a graph attention network (GAT). ENUMERATED EMBODIMENTS 1. A method for identifying a B-cell epitope of a protein, the method comprising: (a) obtaining or having obtained a graph representation of the protein; (b) providing the graph representation of the protein to a generalized machine learning algorithm model to generate a plurality of predicted probabilities, wherein the generalized machine learning algorithm model has been trained using a non-curated training data set comprising: a plurality of training epitope sequences of a plurality of training proteins; labels for the plurality of training epitope sequences, wherein a label identifies whether a corresponding training epitope sequence was bound by a corresponding binder of a plurality of binders, and 36 IPTS / 128951704.1Attorney Docket No.: SES-014WO PATENT wherein the plurality of training proteins comprises at least 800 different proteins, wherein the plurality of binders comprises at least 700 different binders, (c) classifying each of one or more amino acids of the protein as an epitope-member residue or not an epitope-member residue based on the plurality of predicted probabilities; and (d) generating a B-cell epitope prediction comprising the classified epitope-member residues. 2. The method of embodiment 1, wherein the step (a) comprises generating the graph representation of the protein comprising a monomeric structure of the protein. 3. The method of embodiment 2, wherein the step of generating a monomeric structure of the protein comprises generating a conformational variant of the monomeric structure of the protein, wherein each sidechain of the protein is in a minimized free energy conformation. 4. The method of embodiment 2, wherein the graph representation of the protein comprises a plurality of nodes representing each amino acid of the protein, and a plurality of edges between each pair of nodes. 5. The method of embodiment 2, wherein each edge represents a bond between each pair of nodes. 6. The method of embodiment 5, wherein the bond is a peptide bond. 7. The method of embodiment 1, wherein the step (a) comprises: (i) characterizing one or more structural features of the protein by performing a structural analysis of individual amino acids of the protein; and (ii) performing a pairwise amino acid to amino acid interaction analysis across at least one pair of amino acids in the protein to generate a feature graph. 8. The method of embodiment 7, wherein the one or more structural features are node features and / or edge features. 9. The method of any one of embodiment 8, wherein the method comprises generating a node feature by constructing a feature vector of each atom in the node, wherein the feature vector comprises atom type, distance from the amino acid’s alpha carbon, and curvature information. 37 IPTS / 128951704.1Attorney Docket No.: SES-014WO PATENT 10. The method of embodiment 9, wherein the feature vector of each atom in the node is aggregated to generate the one or more structural feature. 11. The method of any one of embodiment 7-10, wherein the node feature comprises a 3D coordinate of the amino acid, a secondary structure of the amin acid, and standard properties of the amino acid. 12. The method of embodiment 8, wherein the edge feature comprises presence of at least one interaction between at least one pair of atoms. 13. The method of embodiment 12, wherein the presence of the at least one interaction between the at least one pair of atoms comprises distance and angle between residue alpha-carbons, and / or an energy well. 14. The method of embodiment 13, wherein the at least one interaction is a peptide bond, a backbone carbonyl interaction, a salt bridge, a pi-pi interaction, a hydrogen bond, or an ionic interaction. 15. The method of embodiment 7, wherein the feature graph is determined by performing a transform of one or more node features and / or edge features as provided herein. 16. The method of embodiment 15, wherein the transform is a graph neural network (GNN). 17. The method of embodiment 16, wherein the graph neural network (GNN) is a message- passing graph neural network (MPGNN). 18. The method of embodiment 15, wherein the transform further comprises performing message passing via node and edge updates. 19. The method of embodiment 18, wherein the transform is a single-layer feedforward neural network. 21. The method of cany one of embodiments 14-19, wherein the feature graph is generated by iteratively analyzing each amino acid of the protein. 22. The method of embodiment 1, wherein the step (b) comprises generating a probability that the amino acid is the epitope-member. 23. The method of embodiment 1, wherein the step (d) comprises generating a set of amino acids that belong to the epitope. 38 IPTS / 128951704.1Attorney Docket No.: SES-014WO PATENT 24. The method of embodiment 1, wherein the machine learning algorithm model has been trained using at least 800 annotated B cell epitopes.25. The method of any one of embodiments 1-22, wherein the method is characterized by an AUROC of at least 0.705. 26. The method of any one of embodiments 1-22, wherein the method is characterized by an AUPRC of at least 0.256. 27. A method for identifying a B-cell epitope of a protein, and a paratope of a binder of the epitope, the method comprising: (a) obtaining or having obtained a graph representation of the protein and a graph representation of the binder; (b) transforming the obtained graph representation of the protein and the graph representation of the binder to construct a bipartite graph comprising: a plurality of protein embeddings; a plurality of binder embeddings; and a plurality of edges connecting protein embeddings and binder embeddings, wherein the plurality of edges comprise predicted weights determined through graph-based collaborative filtering; (c) providing the bipartite graph to a machine learning model to generate a plurality of predicted probabilities; (d) classifying each amino acid of the protein as an epitope-member residue or not an epitope-member residue, and each amino acid of the binder as a paratope-member or not a paratope-member based on the plurality of predicted probabilities; and (e) generating a B-cell epitope prediction comprising the classified epitope-member residues. 28. The method of embodiment 25, wherein the step (a) comprises generating a graph representation of the protein and a graph representation of the binder, each comprising a monomeric structure of the protein and the antibody. 29. The method of embodiment 26, wherein the step of generating a monomeric structure of the protein and the antibody comprises generating a conformational variant of each the monomeric 39 IPTS / 128951704.1Attorney Docket No.: SES-014WO PATENT structure of the protein and the monomeric structure of the antibody, wherein each sidechain of the protein and each sidechain of the antibody is in a minimized free energy conformation. 30. The method of embodiment 26, wherein the a graph representation of the protein and a graph representation of the binder each comprises a plurality of nodes representing each amino acid of the protein and the antibody, and a plurality of edges between each pair of nodes. 31. The method of embodiment 26, wherein each edge represents a bond between each pair of nodes. 32. The method of embodiment 29, wherein the bond is a peptide bond. 33. The method of embodiment 27, wherein the step (a) comprises: (i) characterizing one or more structural features of the protein and the antibody by independently performing a structural analysis of individual amino acids of the protein and the antibody; and (ii) performing a pairwise amino acid to amino acid interaction analysis across at least one pair of amino acids in the protein and the antibody to generate an feature graph for each the protein and the antibody. 34. The method of embodiment 33, wherein the one or more structural features are node features and / or edge features. 35. The method of any one of embodiment 33, wherein the method comprises generating a node feature by constructing a feature vector of each atom in the node, wherein the feature vector comprises atom type, distance from the amino acid’s alpha carbon, and curvature information. 36. The method of embodiment 35, wherein the feature vector of each atom in the node is aggregated to generate the one or more structural feature. 37. The method of any one of embodiment 33-36, wherein the node feature comprise a 3D coordinate of the amino acid, a secondary structure of the amin acid, and standard properties of the amino acid. 38. The method of embodiment 33, wherein the edge feature comprises presence of at least one interaction between at least one pair of atoms. 40 IPTS / 128951704.1Attorney Docket No.: SES-014WO PATENT 39. The method of embodiment 38, wherein the presence of the at least one interaction between the at least one pair of atoms comprises distance and angle between residue alpha-carbons, and / or an energy well. 40. The method of embodiment 39, wherein the at least one interaction is a peptide bond, a backbone carbonyl interaction, a salt bridge, a pi-pi interaction, a hydrogen bond, or an ionic interaction. 41. The method of embodiment 33, wherein the feature graph is determined by performing a transform of one or more node features and / or edge features as provided herein. 42. The method of embodiment 41, wherein the transform is a graph neural network (GNN). 43. The method of embodiment 42, wherein the graph neural network (GNN) is a message- passing graph neural network (MPGNN). 44. The method of embodiment 41, wherein the transform further comprises performing message passing via node and edge updates. 45. The method of embodiment 44, wherein the transform is a single-layer feedforward neural network. 46. The method of cany one of embodiments 40-45, wherein the feature graph is generated by iteratively analyzing each amino acid of the protein. 47. The method of embodiment 27, wherein the method further comprises, prior to step (c), predicting edge weights of the bipartite graph by performing a transform of the bipartite graph. 48. The method of embodiment 47, wherein the transform is a graph attention network (GAT). 49. The method of embodiment 27, wherein the step (c) comprises generating a probability that the amino acid of the protein is an epitope-member and that the amino acid of the antibody is a paratope-member. 50. The method of embodiment 27, wherein the step (e) comprises generating a set of amino acids that belong to the epitope and a set of amino acids that belong to the paratope. 51. The method of embodiment 50, wherein the epitope and the paratope interact with each other. 41 IPTS / 128951704.1Attorney Docket No.: SES-014WO PATENT 52. The method of embodiment 27, wherein the machine learning algorithm model has been trained using at least 4000 annotated B cell epitopes. 53. The method of any one of embodiments 27-52, wherein the method is characterized by an epitope AUROC of at least 0.761, at least 0.829, or at least 0.844. 54. The method of any one of embodiments 27-52, wherein the method is characterized by an epitope AUPRC of at least 0.272, or at least 0.412. 55. The method of any one of embodiments 27-52, wherein the method is characterized by a paratope AUROC of at least 0.923, at least 0.938, or at least 0.945. 56. The method of any one of embodiments 27-52, wherein the method is characterized by a paratope AUPRC of at least 0.510, at least 0.577, or at least 0.581. 57. A non-transitory computer-readable storage medium, the computer-readable storage medium comprising instructions that when executed by a processor, cause the processor to: (a) obtain or having obtained a graph representation of a protein; (b) provide the a graph representation of a protein to a generalized machine learning algorithm model to generate a plurality of predicted probabilities, wherein the generalized machine learning algorithm model has been trained using a non-curated training data set comprising: a plurality of training epitope sequences of a plurality of training proteins; labels for the plurality of training epitope sequences, wherein a label identifies whether a corresponding training epitope sequence was bound by a corresponding binder of a plurality of binders, and wherein the plurality of training proteins comprises at least 800 different proteins, wherein the plurality of binders comprises at least 700 different binders, (c) classify each of one or more amino acids of the protein as an epitope-member residue or not an epitope-member residue based on the plurality of predicted probabilities; and (d) generate a B-cell epitope prediction comprising the classified epitope-member residues. 42 IPTS / 128951704.1Attorney Docket No.: SES-014WO PATENT 58. The non-transitory computer-readable storage medium of embodiment 57, wherein the step (a) comprises generating the graph representation of a protein comprising a monomeric structure of the protein. 59. The non-transitory computer-readable storage medium of embodiment 58, wherein the step of generating a monomeric structure of the protein comprises generating a conformational variant of the monomeric structure of the protein, wherein each sidechain of the protein is in a minimized free energy conformation. 60. The non-transitory computer-readable storage medium of embodiment 58, wherein the graph representation of a protein comprises a plurality of nodes representing each amino acid of the protein, and a plurality of edges between each pair of nodes. 61. The non-transitory computer-readable storage medium of embodiment 58, wherein each edge represents a bond between each pair of nodes. 62. The non-transitory computer-readable storage medium of embodiment 61, wherein the bond is a peptide bond. 63. The non-transitory computer-readable storage medium of embodiment 57, wherein the step (a) comprises: (i) characterizing one or more structural features of the protein by performing a structural analysis of individual amino acids of the protein; and (ii) performing a pairwise amino acid to amino acid interaction analysis across at least one pair of amino acids in the protein to generate a feature graph. 64. The non-transitory computer-readable storage medium of embodiment 63, wherein the one or more structural features are node features and / or edge features. 65. The non-transitory computer-readable storage medium of any one of embodiment 64, wherein the non-transitory computer-readable storage medium comprises generating a node feature by constructing a feature vector of each atom in the node, wherein the feature vector comprises atom type, distance from the amino acid’s alpha carbon, and curvature information. 43 IPTS / 128951704.1Attorney Docket No.: SES-014WO PATENT 66. The non-transitory computer-readable storage medium of embodiment 64, wherein the feature vector of each atom in the node is aggregated to generate the one or more structural feature. 67. The non-transitory computer-readable storage medium of any one of embodiment 63-66, wherein the node feature comprise a 3D coordinate of the amino acid, a secondary structure of the amin acid, and standard properties of the amino acid. 68. The non-transitory computer-readable storage medium of embodiment 64, wherein the edge feature comprises presence of at least one interaction between at least one pair of atoms. 69. The non-transitory computer-readable storage medium of embodiment 68, wherein the presence of the at least one interaction between the at least one pair of atoms comprises distance and angle between residue alpha-carbons, and / or an energy well. 70. The non-transitory computer-readable storage medium of embodiment 69, wherein the at least one interaction is a peptide bond, a backbone carbonyl interaction, a salt bridge, a pi-pi interaction, a hydrogen bond, or an ionic interaction. 71. The non-transitory computer-readable storage medium of embodiment 70, wherein the feature graph is determined by performing a transform of one or more node features and / or edge features as provided herein. 72. The non-transitory computer-readable storage medium of embodiment 71, wherein the transform is a graph neural network (GNN). 73. The non-transitory computer-readable storage medium of embodiment 72, wherein the graph neural network (GNN) is a message-passing graph neural network (MPGNN). 74. The non-transitory computer-readable storage medium of embodiment 71, wherein the transform further comprises performing message passing via node and edge updates. 75. The non-transitory computer-readable storage medium of embodiment 74, wherein the transform is a single-layer feedforward neural network. 76. The non-transitory computer-readable storage medium of cany one of embodiments 70-75, wherein the feature graph is generated by iteratively analyzing each amino acid of the protein. 44 IPTS / 128951704.1Attorney Docket No.: SES-014WO PATENT 77. The non-transitory computer-readable storage medium of embodiment 57, wherein the step (b) comprises generating a probability that the amino acid is the epitope-member. 78. The non-transitory computer-readable storage medium of embodiment 57, wherein the step (d) comprises generating a set of amino acids that belong to the epitope. 79. The non-transitory computer-readable storage medium of embodiment 57, wherein the machine learning algorithm model has been trained using at least 800 annotated B cell epitopes.80. The non-transitory computer-readable storage medium of any one of embodiments 57-79, wherein performance of the B-cell epitope prediction is characterized by an AUROC of at least 0.705. 81. The non-transitory computer-readable storage medium of any one of embodiments 57-79, wherein performance of the B-cell epitope prediction is characterized by an AUPRC of at least 0.256. 82. A non-transitory computer-readable storage medium, the computer-readable storage medium comprising instructions that when executed by a processor, cause the processor to: (a) obtain or having obtained a graph representation of a protein and a graph representation of a binder; (b) transform the obtained the graph representation of the protein and the graph representation of the binder to construct a bipartite graph comprising: a plurality of protein embeddings; a plurality of binder embeddings; and a plurality of edges connecting protein embeddings and binder embeddings, wherein the plurality of edges comprise predicted weights determined through graph-based collaborative filtering; (c) provide the bipartite graph to a machine learning model to generate a plurality of predicted probabilities; (d) classify each amino acid of the protein as an epitope-member residue or not an epitope-member residue, and each amino acid of the binder as a paratope-member or not a paratope-member based on the plurality of predicted probabilities; and (e) generate a B-cell epitope prediction comprising the classified epitope-member residues. 45 IPTS / 128951704.1Attorney Docket No.: SES-014WO PATENT 83. The non-transitory computer-readable storage medium of embodiment 82, wherein the step (a) comprises generating the graph representation of the protein and the graph representation of the binder, each comprising a monomeric structure of the protein and the antibody. 84. The non-transitory computer-readable storage medium of embodiment 83, wherein the step of generating a monomeric structure of the protein and the antibody comprises generating a conformational variant of each the monomeric structure of the protein and the monomeric structure of the antibody, wherein each sidechain of the protein and each sidechain of the antibody is in a minimized free energy conformation. 85. The non-transitory computer-readable storage medium of embodiment 83, wherein the the graph representation of the protein and the graph representation of the binder each comprises a plurality of nodes representing each amino acid of the protein and the antibody, and a plurality of edges between each pair of nodes. 86. The non-transitory computer-readable storage medium of embodiment 83, wherein each edge represents a bond between each pair of nodes. 87. The non-transitory computer-readable storage medium of embodiment 86, wherein the bond is a peptide bond. 88. The non-transitory computer-readable storage medium of embodiment 82, wherein the step (a) comprises: (i) characterizing one or more structural features of the protein and the antibody by independently performing a structural analysis of individual amino acids of the protein and the antibody; and (ii) performing a pairwise amino acid to amino acid interaction analysis across at least one pair of amino acids in the protein and the antibody to generate an feature graph for each the protein and the antibody. 89. The non-transitory computer-readable storage medium of embodiment 88, wherein the one or more structural features are node features and / or edge features. 90. The non-transitory computer-readable storage medium of any one of embodiment 89, wherein the non-transitory computer-readable storage medium comprises generating a node 46 IPTS / 128951704.1Attorney Docket No.: SES-014WO PATENT feature by constructing a feature vector of each atom in the node, wherein the feature vector comprises atom type, distance from the amino acid’s alpha carbon, and curvature information. 91. The non-transitory computer-readable storage medium of embodiment 89, wherein the feature vector of each atom in the node is aggregated to generate the one or more structural feature. 92. The non-transitory computer-readable storage medium of any one of embodiment 88-91, wherein the node feature comprise a 3D coordinate of the amino acid, a secondary structure of the amin acid, and standard properties of the amino acid. 93. The non-transitory computer-readable storage medium of embodiment 89, wherein the edge feature comprises presence of at least one interaction between at least one pair of atoms. 94. The non-transitory computer-readable storage medium of embodiment 93, wherein the presence of the at least one interaction between the at least one pair of atoms comprises distance and angle between residue alpha-carbons, and / or an energy well. 95. The non-transitory computer-readable storage medium of embodiment 94, wherein the at least one interaction is a peptide bond, a backbone carbonyl interaction, a salt bridge, a pi-pi interaction, a hydrogen bond, or an ionic interaction. 96. The non-transitory computer-readable storage medium of embodiment 95, wherein the feature graph is determined by performing a transform of one or more node features and / or edge features as provided herein. 97. The non-transitory computer-readable storage medium of embodiment 96, wherein the transform is a graph neural network (GNN). 98. The non-transitory computer-readable storage medium of embodiment 97, wherein the graph neural network (GNN) is a message-passing graph neural network (MPGNN). 99. The non-transitory computer-readable storage medium of embodiment 96, wherein the transform further comprises performing message passing via node and edge updates. 100. The non-transitory computer-readable storage medium of embodiment 99, wherein the transform is a single-layer feedforward neural network. 47 IPTS / 128951704.1Attorney Docket No.: SES-014WO PATENT 101. The non-transitory computer-readable storage medium of cany one of embodiments 95-100, wherein the feature graph is generated by iteratively analyzing each amino acid of the protein. 102. The non-transitory computer-readable storage medium of embodiment 82, wherein the non- transitory computer-readable storage medium further comprises, prior to step (c), predicting edge weights of the bipartite graph by performing a transform of the bipartite graph. 103. The non-transitory computer-readable storage medium of embodiment 102, wherein the transform is a graph attention network (GAT). 104. The non-transitory computer-readable storage medium of embodiment 82, wherein the step (c) comprises generating a probability that the amino acid of the protein is an epitope-member and that the amino acid of the antibody is a paratope-member. 105. The non-transitory computer-readable storage medium of embodiment 82, wherein the step (e) comprises generating a set of amino acids that belong to the epitope and a set of amino acids that belong to the paratope. 106. The non-transitory computer-readable storage medium of embodiment 105, wherein the epitope and the paratope interact with each other. 107. The non-transitory computer-readable storage medium of embodiment 82, wherein the machine learning algorithm model has been trained using at least 4000 annotated B cell epitopes. 108. The non-transitory computer-readable storage medium of any one of embodiments 82-107, wherein performance of the B-cell epitope prediction is characterized by an epitope AUROC of at least 0.761, at least 0.829, or at least 0.844. 109. The non-transitory computer-readable storage medium of any one of embodiments 82-107, wherein performance of the B-cell epitope prediction is characterized by an epitope AUPRC of at least 0.272, or at least 0.412. 110. The non-transitory computer-readable storage medium of any one of embodiments 82-107, wherein performance of the B-cell epitope prediction is characterized by a paratope AUROC of at least 0.923, at least 0.938, or at least 0.945. 48 IPTS / 128951704.1Attorney Docket No.: SES-014WO PATENT 111. The non-transitory computer-readable storage medium of any one of embodiments 82-107, wherein performance of the B-cell epitope prediction is characterized by a paratope AUPRC of at least 0.510, at least 0.577, or at least 0.581. 112. A system comprising a non-transitory computer-readable storage medium and a processor, wherein the non-transitory computer-readable storage medium comprises: a) a bipartite graph encoded on the non-transitory computer-readable storage medium; and b) instructions for generating a B-cell epitope prediction comprising a classified epitope-member residues from bipartite graph, wherein the bipartite graph is a transformation of a graph representation of a protein and a graph representation of a binder. 113. The system of embodiment 112, wherein the encoded bipartite graph comprises: a plurality of protein embeddings; a plurality of binder embeddings; and a plurality of edges connecting protein embeddings and binder embeddings, wherein the plurality of edges comprise predicted weights determined through graph-based collaborative filtering. 114. The system of embodiments 112 or 113, wherein the instructions for generating a B-cell epitope prediction comprise instructions for a) providing the bipartite graph to a machine learning model to generate a plurality of predicted probabilities; b) classifying each amino acid of the protein as an epitope-member residue or not an epitope-member residue, and each amino acid of the binder as a paratope-member or not a paratope-member based on the plurality of predicted probabilities; and c) generating the B-cell epitope prediction comprising the classified epitope-member residues. 49 IPTS / 128951704.1Attorney Docket No.: SES-014WO PATENT 115. The system of any one of embodiments 112-114, wherein the transformation of the graph representation of the protein and the graph representation of the binder comprises performing a transform of the bipartite graph. 116. The method of embodiment 115, wherein the transform is a graph attention network (GAT). 117. A method for identifying a B-cell epitope of a protein, the method comprising: (a) obtaining or having obtained a graph representation of the protein; (b) providing the graph representation of the protein to a generalized machine learning algorithm model to generate a plurality of predicted probabilities, wherein the generalized machine learning algorithm model is trained using a training data set comprising at least 1,000 loop‑mediated protein‑protein interaction interfaces; (c) classifying each of one or more amino acids of the protein as an epitope-member residue or not an epitope-member residue based on the plurality of predicted probabilities; and (d) generating a B-cell epitope prediction comprising the classified epitope-member residues. 118. The method of embodiment 117, wherein the step (a) comprises generating the graph representation of the protein comprising a monomeric structure of the protein. 119. The method of embodiment 118, wherein the step of generating a monomeric structure of the protein comprises generating a conformational variant of the monomeric structure of the protein, wherein each sidechain of the protein is in a minimized free energy conformation. 120. The method of embodiment 118, wherein the graph representation of the protein comprises a plurality of nodes representing each amino acid of the protein, and a plurality of edges between each pair of nodes. 121. The method of embodiment 118, wherein each edge represents a bond between each pair of nodes. 122. The method of embodiment 121, wherein the bond is a peptide bond. 123. The method of embodiment 117, wherein the step (a) comprises: (i) characterizing one or more structural features of the protein by performing a structural analysis of individual amino acids of the protein; and 50 IPTS / 128951704.1Attorney Docket No.: SES-014WO PATENT (ii) performing a pairwise amino acid to amino acid interaction analysis across at least one pair of amino acids in the protein to generate a feature graph. 124. The method of embodiment 123, wherein the one or more structural features are node features and / or edge features. 125. The method of any one of embodiment 124, wherein the method comprises generating a node feature by constructing a feature vector of each atom in the node, wherein the feature vector comprises atom type, distance from the amino acid’s alpha carbon, and curvature information. 126. The method of embodiment 125, wherein the feature vector of each atom in the node is aggregated to generate the one or more structural feature. 127. The method of any one of embodiment 123-126, wherein the node feature comprises a 3D coordinate of the amino acid, a secondary structure of the amin acid, and standard properties of the amino acid. 128. The method of embodiment 124, wherein the edge feature comprises presence of at least one interaction between at least one pair of atoms. 129. The method of embodiment 128, wherein the presence of the at least one interaction between the at least one pair of atoms comprises distance and angle between residue alpha- carbons, and / or an energy well. 130. The method of embodiment 129, wherein the at least one interaction is a peptide bond, a backbone carbonyl interaction, a salt bridge, a pi-pi interaction, a hydrogen bond, or an ionic interaction. 131. The method of embodiment 123, wherein the feature graph is determined by performing a transform of one or more node features and / or edge features as provided herein. 132. The method of embodiment 131, wherein the transform is a graph neural network (GNN). 133. The method of embodiment 132, wherein the graph neural network (GNN) is a message- passing graph neural network (MPGNN). 134. The method of embodiment 131, wherein the transform further comprises performing message passing via node and edge updates. 51 IPTS / 128951704.1Attorney Docket No.: SES-014WO PATENT 135. The method of embodiment 134, wherein the transform is a single-layer feedforward neural network. 136. The method of cany one of embodiments 130-135, wherein the feature graph is generated by iteratively analyzing each amino acid of the protein. 137. The method of embodiment 117, wherein the generalized machine learning algorithm model comprises unsupervised learning, supervised learning, semi-supervised learning, reinforcement learning, transfer learning, incremental learning, curriculum learning, and learning to learn. 138. The method of embodiment 117, the generalized machine learning algorithm model comprises linear classifiers, logistic classifiers, Bayesian networks, random forest, neural networks, graph neural networks (GNN), message-passing graph neural networks (MPGNN), matrix factorization, hidden Markov model, support vector machine, K-means clustering, or K- nearest neighbor. 139. The method of any one of embodiments 117-138, wherein the training data set comprises at least 10000 loop‑mediated protein‑protein interaction interfaces. 140. The method of any one of embodiments 117-139, wherein the generalized machine learning algorithm model trained using the training data is further refined for antibody-antigen interactions. 141. The method of any one of embodiments 117-140, wherein the step (b) comprises generating a probability that the amino acid is the epitope-member. 142. The method of embodiment 117, wherein the step (d) comprises generating a set of amino acids that belong to the epitope. 143. The method of embodiment 117, wherein the machine learning algorithm model has been trained using at least 800 annotated B cell epitopes.144. The method of any one of embodiments 1-143, wherein the method is characterized by an AUROC of at least 0.705. 145. The method of any one of embodiments 1-143, wherein the method is characterized by an AUPRC of at least 0.256. 52 IPTS / 128951704.1Attorney Docket No.: SES-014WO PATENT 146. A method for identifying a B-cell epitope of a protein, and a paratope of a binder of the epitope, the method comprising: (a) obtaining or having obtained a graph representation of the protein and a graph representation of the binder; (b) transforming the obtained graph representation of the protein and the graph representation of the binder to construct a bipartite graph comprising: a plurality of protein embeddings; a plurality of binder embeddings; and a plurality of edges connecting protein embeddings and binder embeddings, wherein the plurality of edges comprise predicted weights determined through graph-based collaborative filtering; (c) providing the bipartite graph to a machine learning model to generate a plurality of predicted probabilities; (d) classifying each amino acid of the protein as an epitope-member residue or not an epitope-member residue, and each amino acid of the binder as a paratope-member or not a paratope-member based on the plurality of predicted probabilities; and (e) generating a B-cell epitope prediction comprising the classified epitope-member residues. 147. The method of embodiment 146, wherein the step (a) comprises generating a graph representation of the protein and a graph representation of the binder, each comprising a monomeric structure of the protein and the antibody. 148. The method of embodiment 147, wherein the step of generating a monomeric structure of the protein and the antibody comprises generating a conformational variant of each the monomeric structure of the protein and the monomeric structure of the antibody, wherein each sidechain of the protein and each sidechain of the antibody is in a minimized free energy conformation. 149. The method of embodiment 147, wherein the a graph representation of the protein and a graph representation of the binder each comprises a plurality of nodes representing each amino acid of the protein and the antibody, and a plurality of edges between each pair of nodes. 53 IPTS / 128951704.1Attorney Docket No.: SES-014WO PATENT 150. The method of embodiment 149, wherein each edge represents a bond between each pair of nodes. 151. The method of embodiment 150, wherein the bond is a peptide bond. 152. The method of embodiment 146, wherein the step (a) comprises: (i) characterizing one or more structural features of the protein and the antibody by independently performing a structural analysis of individual amino acids of the protein and the antibody; and (ii) performing a pairwise amino acid to amino acid interaction analysis across at least one pair of amino acids in the protein and the antibody to generate an feature graph for each the protein and the antibody. 153. The method of embodiment 152, wherein the one or more structural features are node features and / or edge features. 154. The method of any one of embodiment 152 or 153, wherein the method comprises generating a node feature by constructing a feature vector of each atom in the node, wherein the feature vector comprises atom type, distance from the amino acid’s alpha carbon, and curvature information. 155. The method of embodiment 154, wherein the feature vector of each atom in the node is aggregated to generate the one or more structural feature. 156. The method of any one of embodiment 152-155, wherein the node feature comprise a 3D coordinate of the amino acid, a secondary structure of the amin acid, and standard properties of the amino acid. 157. The method of embodiment 152, wherein the edge feature comprises presence of at least one interaction between at least one pair of atoms. 158. The method of embodiment 157, wherein the presence of the at least one interaction between the at least one pair of atoms comprises distance and angle between residue alpha- carbons, and / or an energy well. 159. The method of embodiment 158, wherein the at least one interaction is a peptide bond, a backbone carbonyl interaction, a salt bridge, a pi-pi interaction, a hydrogen bond, or an ionic interaction. 54 IPTS / 128951704.1Attorney Docket No.: SES-014WO PATENT 160. The method of embodiment 152, wherein the feature graph is determined by performing a transform of one or more node features and / or edge features as provided herein. 161. The method of embodiment 160, wherein the transform is a graph neural network (GNN). 162. The method of embodiment 161, wherein the graph neural network (GNN) is a message- passing graph neural network (MPGNN). 163. The method of embodiment 160, wherein the transform further comprises performing message passing via node and edge updates. 164. The method of embodiment 163, wherein the transform is a single-layer feedforward neural network. 165. The method of cany one of embodiments 152-164, wherein the feature graph is generated by iteratively analyzing each amino acid of the protein. 166. The method of embodiment 146, wherein the method further comprises, prior to step (c), predicting edge weights of the bipartite graph by performing a transform of the bipartite graph. 167. The method of embodiment 166, wherein the transform is a graph attention network (GAT). 168. The method of embodiment 146, wherein the machine learning model comprises unsupervised learning, supervised learning, semi-supervised learning, reinforcement learning, transfer learning, incremental learning, curriculum learning, and learning to learn. 169. The method of embodiment 146, the machine learning model comprises linear classifiers, logistic classifiers, Bayesian networks, random forest, neural networks, graph neural networks (GNN), message-passing graph neural networks (MPGNN), matrix factorization, hidden Markov model, support vector machine, K-means clustering, or K-nearest neighbor. 170. The method of any one of embodiments 146-169, wherein the machine learning model is trained using a training data set comprising at least 1,000 loop‑mediated protein‑protein interaction interfaces 171. The method of embodiment 170, wherein the training data set comprises at least 10000 loop‑mediated protein‑protein interaction interfaces. 172. The method of any one of embodiments 146-171, wherein the machine learning model trained using the training data is further refined for antibody-antigen interactions. 55 IPTS / 128951704.1Attorney Docket No.: SES-014WO PATENT 173. The method of embodiment 146, wherein the step (c) comprises generating a probability that the amino acid of the protein is an epitope-member and that the amino acid of the antibody is a paratope-member. 174. The method of embodiment 146, wherein the step (e) comprises generating a set of amino acids that belong to the epitope and a set of amino acids that belong to the paratope. 175. The method of embodiment 174, wherein the epitope and the paratope interact with each other. 176. The method of embodiment 146, wherein the machine learning algorithm model has been trained using at least 4000 annotated B cell epitopes. 177. The method of any one of embodiments 146-176, wherein the method is characterized by an epitope AUROC of at least 0.761, at least 0.829, or at least 0.844. 178. The method of any one of embodiments 146-176, wherein the method is characterized by an epitope AUPRC of at least 0.272, or at least 0.412. 179. The method of any one of embodiments 146-176, wherein the method is characterized by a paratope AUROC of at least 0.923, at least 0.938, or at least 0.945. 180. The method of any one of embodiments 146-176, wherein the method is characterized by a paratope AUPRC of at least 0.510, at least 0.577, or at least 0.581. 181. A non-transitory computer-readable storage medium, the computer-readable storage medium comprising instructions that when executed by a processor, cause the processor to: (a) obtain or having obtained a graph representation of a protein; (b) providing the graph representation of the protein to a generalized machine learning algorithm model to generate a plurality of predicted probabilities, wherein the generalized machine learning algorithm model is trained using a training data set comprising at least 1,000 loop‑mediated protein‑protein interaction interfaces; (c) classify each of one or more amino acids of the protein as an epitope-member residue or not an epitope-member residue based on the plurality of predicted probabilities; and 56 IPTS / 128951704.1Attorney Docket No.: SES-014WO PATENT (d) generate a B-cell epitope prediction comprising the classified epitope-member residues. 182. The non-transitory computer-readable storage medium of embodiment 181, wherein the step (a) comprises generating the graph representation of a protein comprising a monomeric structure of the protein. 183. The non-transitory computer-readable storage medium of embodiment 182, wherein the step of generating a monomeric structure of the protein comprises generating a conformational variant of the monomeric structure of the protein, wherein each sidechain of the protein is in a minimized free energy conformation. 184. The non-transitory computer-readable storage medium of embodiment 182, wherein the graph representation of a protein comprises a plurality of nodes representing each amino acid of the protein, and a plurality of edges between each pair of nodes. 185. The non-transitory computer-readable storage medium of embodiment 182, wherein each edge represents a bond between each pair of nodes. 186. The non-transitory computer-readable storage medium of embodiment 185, wherein the bond is a peptide bond. 187. The non-transitory computer-readable storage medium of embodiment 181, wherein the step (a) comprises: (i) characterizing one or more structural features of the protein by performing a structural analysis of individual amino acids of the protein; and (ii) performing a pairwise amino acid to amino acid interaction analysis across at least one pair of amino acids in the protein to generate a feature graph. 188. The non-transitory computer-readable storage medium of embodiment 187, wherein the one or more structural features are node features and / or edge features. 189. The non-transitory computer-readable storage medium of any one of embodiment 188, wherein the non-transitory computer-readable storage medium comprises generating a node feature by constructing a feature vector of each atom in the node, wherein the feature vector comprises atom type, distance from the amino acid’s alpha carbon, and curvature information. 57 IPTS / 128951704.1Attorney Docket No.: SES-014WO PATENT 190. The non-transitory computer-readable storage medium of embodiment 189, wherein the feature vector of each atom in the node is aggregated to generate the one or more structural feature. 191. The non-transitory computer-readable storage medium of any one of embodiment 187-190, wherein the node feature comprise a 3D coordinate of the amino acid, a secondary structure of the amin acid, and standard properties of the amino acid. 192. The non-transitory computer-readable storage medium of embodiment 189, wherein the edge feature comprises presence of at least one interaction between at least one pair of atoms. 193. The non-transitory computer-readable storage medium of embodiment 192, wherein the presence of the at least one interaction between the at least one pair of atoms comprises distance and angle between residue alpha-carbons, and / or an energy well. 194. The non-transitory computer-readable storage medium of embodiment 193, wherein the at least one interaction is a peptide bond, a backbone carbonyl interaction, a salt bridge, a pi-pi interaction, a hydrogen bond, or an ionic interaction. 195. The non-transitory computer-readable storage medium of embodiment 194, wherein the feature graph is determined by performing a transform of one or more node features and / or edge features as provided herein. 196. The non-transitory computer-readable storage medium of embodiment 195, wherein the transform is a graph neural network (GNN). 197. The non-transitory computer-readable storage medium of embodiment 196, wherein the graph neural network (GNN) is a message-passing graph neural network (MPGNN). 198. The non-transitory computer-readable storage medium of embodiment 195, wherein the transform further comprises performing message passing via node and edge updates. 199. The non-transitory computer-readable storage medium of embodiment 198, wherein the transform is a single-layer feedforward neural network. 200. The non-transitory computer-readable storage medium of cany one of embodiments 194- 199, wherein the feature graph is generated by iteratively analyzing each amino acid of the protein. 58 IPTS / 128951704.1Attorney Docket No.: SES-014WO PATENT 201. The non-transitory computer-readable storage medium of embodiment 181, wherein the generalized machine learning algorithm model comprises unsupervised learning, supervised learning, semi-supervised learning, reinforcement learning, transfer learning, incremental learning, curriculum learning, and learning to learn. 202. The non-transitory computer-readable storage medium of embodiment 181, the generalized machine learning algorithm model comprises linear classifiers, logistic classifiers, Bayesian networks, random forest, neural networks, graph neural networks (GNN), message-passing graph neural networks (MPGNN), matrix factorization, hidden Markov model, support vector machine, K-means clustering, or K-nearest neighbor. 203. The non-transitory computer-readable storage medium of any one of embodiments 181-202, wherein the training data set comprises at least 10000 loop‑mediated protein‑protein interaction interfaces. 204. The non-transitory computer-readable storage medium of any one of embodiments 181-203, wherein the generalized machine learning algorithm model trained using the training data is further refined for antibody-antigen interactions. 205. The non-transitory computer-readable storage medium of embodiment 181, wherein the step (b) comprises generating a probability that the amino acid is the epitope-member. 206. The non-transitory computer-readable storage medium of embodiment 181, wherein the step (d) comprises generating a set of amino acids that belong to the epitope. 207. The non-transitory computer-readable storage medium of embodiment 181, wherein the machine learning algorithm model has been trained using at least 800 annotated B cell epitopes. 208. The non-transitory computer-readable storage medium of any one of embodiments 181-207, wherein performance of the B-cell epitope prediction is characterized by an AUROC of at least 0.705. 209. The non-transitory computer-readable storage medium of any one of embodiments 181-207, wherein performance of the B-cell epitope prediction is characterized by an AUPRC of at least 0.256. 210. A non-transitory computer-readable storage medium, the computer-readable storage medium comprising instructions that when executed by a processor, cause the processor to: 59 IPTS / 128951704.1Attorney Docket No.: SES-014WO PATENT (a) obtain or having obtained a graph representation of a protein and a graph representation of a binder; (b) transform the obtained the graph representation of the protein and the graph representation of the binder to construct a bipartite graph comprising: a plurality of protein embeddings; a plurality of binder embeddings; and a plurality of edges connecting protein embeddings and binder embeddings, wherein the plurality of edges comprise predicted weights determined through graph-based collaborative filtering; (c) provide the bipartite graph to a machine learning model to generate a plurality of predicted probabilities; (d) classify each amino acid of the protein as an epitope-member residue or not an epitope-member residue, and each amino acid of the binder as a paratope-member or not a paratope-member based on the plurality of predicted probabilities; and (e) generate a B-cell epitope prediction comprising the classified epitope-member residues. 211. The non-transitory computer-readable storage medium of embodiment 210, wherein the step (a) comprises generating the graph representation of the protein and the graph representation of the binder, each comprising a monomeric structure of the protein and the antibody. 212. The non-transitory computer-readable storage medium of embodiment 211, wherein the step of generating a monomeric structure of the protein and the antibody comprises generating a conformational variant of each the monomeric structure of the protein and the monomeric structure of the antibody, wherein each sidechain of the protein and each sidechain of the antibody is in a minimized free energy conformation. 213. The non-transitory computer-readable storage medium of embodiment 211, wherein the the graph representation of the protein and the graph representation of the binder each comprises a plurality of nodes representing each amino acid of the protein and the antibody, and a plurality of edges between each pair of nodes. 214. The non-transitory computer-readable storage medium of embodiment 211, wherein each edge represents a bond between each pair of nodes. 60 IPTS / 128951704.1Attorney Docket No.: SES-014WO PATENT 215. The non-transitory computer-readable storage medium of embodiment 211, wherein the bond is a peptide bond. 216. The non-transitory computer-readable storage medium of embodiment 210, wherein the step (a) comprises: (i) characterizing one or more structural features of the protein and the antibody by independently performing a structural analysis of individual amino acids of the protein and the antibody; and (ii) performing a pairwise amino acid to amino acid interaction analysis across at least one pair of amino acids in the protein and the antibody to generate an feature graph for each the protein and the antibody. 217. The non-transitory computer-readable storage medium of embodiment 216, wherein the one or more structural features are node features and / or edge features. 218. The non-transitory computer-readable storage medium of any one of embodiment 217, wherein the non-transitory computer-readable storage medium comprises generating a node feature by constructing a feature vector of each atom in the node, wherein the feature vector comprises atom type, distance from the amino acid’s alpha carbon, and curvature information. 219. The non-transitory computer-readable storage medium of embodiment 217, wherein the feature vector of each atom in the node is aggregated to generate the one or more structural feature. 220. The non-transitory computer-readable storage medium of any one of embodiment 216-219, wherein the node feature comprise a 3D coordinate of the amino acid, a secondary structure of the amin acid, and standard properties of the amino acid. 221. The non-transitory computer-readable storage medium of embodiment 217, wherein the edge feature comprises presence of at least one interaction between at least one pair of atoms. 222. The non-transitory computer-readable storage medium of embodiment 221, wherein the presence of the at least one interaction between the at least one pair of atoms comprises distance and angle between residue alpha-carbons, and / or an energy well. 61 IPTS / 128951704.1Attorney Docket No.: SES-014WO PATENT 223. The non-transitory computer-readable storage medium of embodiment 222, wherein the at least one interaction is a peptide bond, a backbone carbonyl interaction, a salt bridge, a pi-pi interaction, a hydrogen bond, or an ionic interaction. 224. The non-transitory computer-readable storage medium of embodiment 223, wherein the feature graph is determined by performing a transform of one or more node features and / or edge features as provided herein. 225. The non-transitory computer-readable storage medium of embodiment 224, wherein the transform is a graph neural network (GNN). 226. The non-transitory computer-readable storage medium of embodiment 225, wherein the graph neural network (GNN) is a message-passing graph neural network (MPGNN). 227. The non-transitory computer-readable storage medium of embodiment 224, wherein the transform further comprises performing message passing via node and edge updates. 228. The non-transitory computer-readable storage medium of embodiment 227, wherein the transform is a single-layer feedforward neural network. 229. The non-transitory computer-readable storage medium of cany one of embodiments 223- 228, wherein the feature graph is generated by iteratively analyzing each amino acid of the protein. 230. The non-transitory computer-readable storage medium of embodiment 210, wherein the non-transitory computer-readable storage medium further comprises, prior to step (c), predicting edge weights of the bipartite graph by performing a transform of the bipartite graph. 231. The non-transitory computer-readable storage medium of embodiment 230, wherein the transform is a graph attention network (GAT). 232. The non-transitory computer-readable storage medium of embodiment 210, wherein the machine learning model comprises unsupervised learning, supervised learning, semi-supervised learning, reinforcement learning, transfer learning, incremental learning, curriculum learning, and learning to learn. 233. The non-transitory computer-readable storage medium of embodiment 210, the machine learning model comprises linear classifiers, logistic classifiers, Bayesian networks, random 62 IPTS / 128951704.1Attorney Docket No.: SES-014WO PATENT forest, neural networks, graph neural networks (GNN), message-passing graph neural networks (MPGNN), matrix factorization, hidden Markov model, support vector machine, K-means clustering, or K-nearest neighbor. 234. The non-transitory computer-readable storage medium of any one of embodiments 210-233, wherein the machine learning model is trained using a training data set comprising at least 1,000 loop‑mediated protein‑protein interaction interfaces 235. The non-transitory computer-readable storage medium of embodiment 234, wherein the training data set comprises at least 10000 loop‑mediated protein‑protein interaction interfaces. 236. The non-transitory computer-readable storage medium of any one of embodiments 210-235, wherein the machine learning model trained using the training data is further refined for antibody-antigen interactions. 237. The non-transitory computer-readable storage medium of embodiment 210, wherein the step (c) comprises generating a probability that the amino acid of the protein is an epitope- member and that the amino acid of the antibody is a paratope-member. 238. The non-transitory computer-readable storage medium of embodiment 210, wherein the step (e) comprises generating a set of amino acids that belong to the epitope and a set of amino acids that belong to the paratope. 239. The non-transitory computer-readable storage medium of embodiment 238, wherein the epitope and the paratope interact with each other. 240. The non-transitory computer-readable storage medium of embodiment 210, wherein the machine learning algorithm model has been trained using at least 4000 annotated B cell epitopes. 241. The non-transitory computer-readable storage medium of any one of embodiments 210-240, wherein performance of the B-cell epitope prediction is characterized by an epitope AUROC of at least 0.761, at least 0.829, or at least 0.844. 242. The non-transitory computer-readable storage medium of any one of embodiments 210-240, wherein performance of the B-cell epitope prediction is characterized by an epitope AUPRC of at least 0.272, or at least 0.412. 63 IPTS / 128951704.1Attorney Docket No.: SES-014WO PATENT 243. The non-transitory computer-readable storage medium of any one of embodiments 210-240, wherein performance of the B-cell epitope prediction is characterized by a paratope AUROC of at least 0.923, at least 0.938, or at least 0.945. 244. The non-transitory computer-readable storage medium of any one of embodiments 210-240, wherein performance of the B-cell epitope prediction is characterized by a paratope AUPRC of at least 0.510, at least 0.577, or at least 0.581. 245. A system comprising a non-transitory computer-readable storage medium and a processor, wherein the non-transitory computer-readable storage medium comprises: a) a bipartite graph encoded on the non-transitory computer-readable storage medium; and b) instructions for generating a B-cell epitope prediction comprising a classified epitope-member residues from bipartite graph, wherein the bipartite graph is a transformation of a graph representation of a protein and a graph representation of a binder. 246. The system of embodiment 245, wherein the encoded bipartite graph comprises: a plurality of protein embeddings; a plurality of binder embeddings; and a plurality of edges connecting protein embeddings and binder embeddings, wherein the plurality of edges comprise predicted weights determined through graph-based collaborative filtering. 247. The system of embodiments 245 or 246, wherein the instructions for generating a B-cell epitope prediction comprise instructions for a) providing the bipartite graph to a machine learning model to generate a plurality of predicted probabilities; b) classifying each amino acid of the protein as an epitope-member residue or not an epitope-member residue, and each amino acid of the binder as a paratope-member or not a paratope-member based on the plurality of predicted probabilities; and 64 IPTS / 128951704.1Attorney Docket No.: SES-014WO PATENT c) generating the B-cell epitope prediction comprising the classified epitope-member residues. 248. The system of any one of embodiments 245-247, wherein the transformation of the graph representation of the protein and the graph representation of the binder comprises performing a transform of the bipartite graph. 249. The system of embodiment 248, wherein the transform is a graph attention network (GAT). EXAMPLES

[0115] Below are examples of specific embodiments. The examples are offered for illustrative purposes only and are not intended to limit the scope. Efforts have been made to ensure accuracy with respect to numbers used, but some experimental error and deviation should be allowed for. Example 1: EpiGraph: Recommender-Style Graph Neural Networks for Highly Accurate Prediction of Conformational B-Cell Epitopes.

[0116] Current B-cell epitope predictors have highly varied output predictions and do not consistently agree with experimental data. For example, on a rigorous benchmarking task of nine leading epitope prediction tools, the ROC-AUC values for predicted epitopes were shown to be almost equivalent to random predictions. In particular, many methods did not have better Matthews correlation coefficients (MCC) with ground truth than generating random surface residues or random patches. Furthermore, only one of these methods (EpiPred, deprecated as of May 2022) predicts the epitope specific to a given antibody, rather than predicting all potential epitopes on a protein’s surface. This renders these methods challenging to apply in situations involving monoclonal antibodies, which are the largest class of clinically approved biologic drugs. Finally, while these methods train and test on existing epitope annotations derived from experimental data, they do not further validate their models by experimentally mapping previously unobserved epitopes, leaving open the question of their applicability to real-world prediction tasks. Predictors based on manually derived features

[0117] A variety of tools predict B cell epitopes based on linear sequences alone or via explicit geometric information derived from protein structure. Methods that only exploit linear sequence 65 IPTS / 128951704.1Attorney Docket No.: SES-014WO PATENT enjoy the advantage of significantly larger datasets of antibody-antigen interactions, but are highly limited in the sense that ~90% of known B cell epitopes are non-linear and exhibit complex 3D topologies. Methods incorporating structural information tend to score individual residues or small surface patches based their solvent accessibility and sidechain orientation, or compute antibody-antigen complementarity measures via computationally intensive docking procedures. While such methods have been available for almost twenty years, they have performed poorly both in silico and on de novo prediction tasks. Graph convolutions for protein interface prediction

[0118] Given the shortcomings of methods based on manually engineered protein surface features or computationally intensive docking studies, recent approaches have turned to deep learning as a means of integrating structural data directly into the training of epitope prediction models. The predictice models disclosed herein have sufficient data in terms of structures of antibody-antigen complexes (>4500 non-redundant structures, as opposed to only 150 in 2015) to plausibly train machine learning models capable of learning general principles of antibody- antigen recognition. Fout and coauthors consider the general problem of predicting the interfaces of protein-protein interactions. They represent the interaction partners as graphs (further details below) and perform message passing node updates on these graphs via a graph convolution operator 11^

[0119] In the as a in a graph representation of the protein’s structure, :\denotes the set of graph neighbors of node ^,[* , [1 , and [5 are learned weight^^^ is the graph adjacency matrix, ] is a biasvector, and Y is a non-linear activation function. They apply this message passing within a twin graph neural network architecture to classify residues as belonging to the interface or not (model architecture reproduced below in FIG. 7, panel C). However, their model was trained on a very small dataset (the Docking Benchmark Dataset, comprising 230 structures of protein-protein 66 IPTS / 128951704.1Attorney Docket No.: SES-014WO PATENT interactions). Similar architecture trained on a much larger dataset of ~4,500 antibody-antigen interactions exhibited poor performance (described in further detail below). A Recommender-style Graph Attention Network Graph-based collaborative filtering

[0120] The essential idea of the predictive models disclosed herein is to treat interactions between antibody and antigen amino acid residues analogously to user-item interactions in recommender systems based on collaborative filtering. In such systems, the task of predictinginteractions between ^ users ^ = {^^, ^^, … , ^^} and ^ items ^ = {^^, ^^, … , ^^} wasconsidered. Known user preferences are encoded by an interaction matrix ^ ∈ ^^×^ where^^^ = 1 if user ^^ interacts with item ^^ , and ^^^ = 0 otherwise. This interaction matrix wasinterpreted as the adjacency matrix of a bipartite graph ^ = {^ ∪ ^, ^}, where the edges ^ ={(^, ^) ∶ ^^^ = 1} correspond to user-item interactions. Graph-based collaborative filtering can bedescribed as the problem of inferring this graph topology for unknown user-item interactionsbased on observed data: ^^^ = (^, Θ), where is a learnable function and Θ represents themodel parameters.Node and edge features

[0121] The problem of predicting B-cell epitopes is approached via two paradigms: the antibody-agnostic view, which aims to identify all potential surface patches on a target protein that may be susceptible to binding by antibodies, and an antibody-specific view, which aims to identify the specific epitope targeted by an antibody that is known to bind the protein. In both cases, all involved proteins are represented via their structure, which may be experimentally determined or predicted via an algorithm such as AlphaFold2.

[0122] To avoid bias arising from conformational changes in the antibody and / or antigen upon complex formation, the antibody and antigen were separated into monomeric structures and use the Rosetta Relax protocol to move the sidechains into a conformation with minimized free energy. These relaxed structures are cast as graphs whose nodes correspond to amino acid residues, while edges correspond to peptide bonds or close contact between residues. To construct the node feature tensor, the atomic information for each residue’s sidechain was 67 IPTS / 128951704.1Attorney Docket No.: SES-014WO PATENT aggregated in the following way: for each atom " belonging to the sidechain of a residue ρ a feature vector ℎ%was constructed encoding the atom type, distance from the residue’s alpha carbon, and curvature (calculated as the discrete mean curvature of a sphere centered at the atom). A residue-level sidechain representation was then determined by aggregating these feature vectors over all atoms belonging to the residue’s sidechain &': )* 1ℎ( = φ , . ℎ% / where φ is a learned multi-layerresidue node features encode its 3D coordinates, secondary structure, and standard amino acid properties derived from reduced dimensional embeddings of physicochemical measurements of amino acids. The edge features encode the distances and angles between residue alpha-carbons, as well as characterization of the residue interaction type as a peptide bond, backbone carbonyl interaction, salt bridge, pi-pi interaction, hydrogen bond, or ionic interaction (FIG. 6, panel B). Transformer operations

[0123] The above representations of antigen and antibody structures were used to derive feature embeddings for the residue nodes using message-passing graph neural networks. The basic operation of the GNNs is the graph attentional operator first. In this setting, the node featurevectors are 0^, 0^, … , 01 ∈ ^2, where 3 is the number of nodes and 4 is the number of nodefeatures, and a weight matrix Θ ∈ ^5×2, where ^ is the number of edge features, which islearned by a multi-layer perceptron. A message passing was then performed via node and edge updates given by 0′^ = α^,^Θ80^ + . α^,^Θ;0^

[0124] Here, :(^)denotes the set of neighbors of node ^, and the attention weights are computed byAttorney Docket No.: SES-014WO PATENT

[0125] Here the attention mechanism is a single-layer feedforward neural network parametrized by the weight vector and <^,^are edge features.

[0126] The above operations are applied iteratively and independently to theantibody and (or antigen alone in the case of antibody-agnostic epitopeprediction) to produce embeddings ^ = {^ > >^, ^^, … , ^^} ⊂ ^ and ^ = {^^, ^^, … , ^^} ⊂ ^ ofthe antibody and antigen nodes, respectively. For antibody-specific epitope predictions, once these representations were obtained, a second graph attention network (GAT) was trained to predict the edge weights of a complete bipartite graph with disjoint vertex sets ^ and ^, akin to a user-item graph in typical recommender systems as described above. Following the final iteration of message passing over the antibody-antigen bipartite graph with predicted edge weights (or antigen graph alone for antibody-agnostic predictions), a single-layer perceptron with one-dimensional output and hyperbolic tangent activation was applied to each fully updated node feature tensor to obtain the final model output in the form of node-level probabilities for theantigen ρ@ @ @ D D D @^, ρ^, … , ρ^ ∈ A0,1B and antibody C^ , C^ , … , ρ^ ∈ A0,1B. Next, ρ^ = EHFG wasinterpreted as the probability that residue ^antigen belongs to theat which theantibody binds, and ρD^ = EHIJ as the probability that residue ^ in the antibody belongs to theparatope of this

[0127] The epitope prediction task (either antibody agnostic or antibody specific) is highly imbalanced, with only a small handful of annotated epitope residues out of hundreds to thousands of residues constituting a given antigen. It should be noted, however, that unannotated residues may indeed belong to an epitope for an antibody whose interaction with the antigen has yet to be characterized. As such, it’s important to caveat the antibody-agnostic model as primarily reliable for its positive predictive power. The same is not true of the antibody-specific model, since published structures of antibody-antigen interactions map the specific epitope targeted by the antibody, with all residues outside the interaction interface properly labeled as negatives.

[0128] An example from the training dataset for the antibody-agnostic predictor takes anantigen ^ as a collection of residues ^ = (K^, K^, … , K1), each with a corresponding binary labelcorresponding to whether it belongs to a B cell epitope or not. For the antibody-specificpredictor, the examples are pairs (^, L), where ^ is an antigen and L is an antibody specificallytargeting that antigen. Each is again represented as a collection of residues, where the label of a 69 IPTS / 128951704.1Attorney Docket No.: SES-014WO PATENT residue is 1 if and only if the residue belongs to the epitope or paratope for that specific antibody-antigen interaction. A weighted cross-entropy loss function was employed for the antibody-specific epitope predictor to account for the class imbalance described above. To mitigate the class imbalance for the antibody-agnostic model and incur a lower penalty for false positive predictions, an in-batch contrastive loss was employed as follows: 1U1U1U1W2. QRℎO S 1ℒ(^) = . O S^ , ℎ^ T − .. QRℎ^ , ℎ^ Twhere Q is a distance kernel, which takes to be cosine distance, resp denotes theembedding vector of positively resp negatively labeled 3O is the number of positiveexamples (annotated epitope residues) and 3Sis the number of negative examples. This loss deliberately separates the node embeddings of positive examples from those of negative examples to prevent overestimation of model accuracy due to class imbalance. Data and model training Data

[0129] The data are annotated B cell epitopes derived from the Immune Epitope Database and antibody-antigen complex structures from the Structural Antibody Database. Structures were filtered to only retain complexes of antibodies bound to protein antigens rather than small molecules or other ligands, resulting in a total of 4420 structures containing 1897 unique protein antigens. Ultimately, the models make predictions at the level of amino acid residues rather than proteins, and this dataset corresponds to over 4 million antigen amino acid residues and over 7 million antibody residues. Given the structure of an antibody-antigen complex, the standard definition of the epitope was used as the set of residues in the antigen for which any atom lies within a distance of 4 Angstrom to any atom belonging to a residue in the antibody (and vice- versa for the paratope, i.e. the set of residues in the antibody defined to interact with the antigen).

[0130] As several antigens are over-represented in the dataset (such as the Spike protein from SARS-CoV2) or exhibit high sequence and / or structure similarity to other antigens, the train-test data split strategy was carefully designed to minimize the possibility of information leakage during model validation. To that end, antigens and antibodies were clustered based on sequence 70 IPTS / 128951704.1Attorney Docket No.: SES-014WO PATENT and structure similarity, with the aim of partitioning the train and test sets into dissimilar sets of antigen-antibody pairs, or antigens alone for the antibody-agnostic prediction task. An antigen sequence similarity threshold of 90% and an antibody CDR3 sequence similarity threshold of 75% was determined based on the change in concavity in the number of antibody-antigen sequence clusters as a function of sequence similarity (FIG. 2, panel A). This resulted in a dataset comprising 4120 antibody-antigen complex structures for training and 300 such structures for validation, with no antigen resp antibody in the training dataset exhibiting more than 90% resp 75% sequence homology to any antigen resp antibody in the validation dataset. Training, validation, and testing

[0131] The predictive models were constructed using the PyTorch Geometric library for graph neural networks. A Bayesian hyperparameter optimization was applied to identify the ideal values for the training batch size, number of heads for the multi-headed attention operations, number of hidden dimensions, number of graph convolutional layers, and learning rate using the aforementioned validation set. This search was conducted automatically using the Sweeps functionality within the Weights & Biases platform. Training took roughly 3 days and 18 hours using an AdamW optimizer with a learning rate of 0.0005 using 4 Tesla V100-SXM2-16GB GPUs. Results In-silico benchmarking

[0132] The best-performing antibody-specific epitope prediction model achieves an AU-ROC for epitope residues of ~0.84, outperforming state-of-the-art structure-based B cell epitope prediction algorithms in terms of their self-reported metrics (Table 1). Since the epitope prediction task is highly imbalanced with a preponderance of non-epitope vs epitope residues, the AU-PRC for epitope residues is a more informative metric then AU-ROC, and the model achieves a value of ~0.412, higher than all other models considered. A number of example predictions for epitopes in the validation sets are shown in Figure 3A and 3B. In each case, the model accurately identifies both the epitope and paratope (region of the antibody binding to the antigen) for antigens that exhibit low sequence and structure homology to antigens in the training dataset. 71 IPTS / 128951704.1Attorney Docket No.: SES-014WO PATENT Comparison to non-recommender architectures

[0133] The recommender-style architecture was compared to alternative models similar to the twin neural network approach of Fout et al., based on graph convolution networks (GCNs) without a recommender-style bipartite graph or attention mechanisms (FIG. 7, panel C). These twin GCNs (with various optimized hyperparameter choices) exhibited poor generalization outside the training dataset and failed to yield a model capable of predicting epitopes in an antibody-specific fashion, with a best AU-PRC of 0.272 for epitope residues versus an AU-PRC of 0.412 for the best recommender-style graph attention network (FIG. 7, panel D, and Table 1). On inspection, the final layer node embeddings for example antigens were invariant under permutation of the partner antibody, while the antibody node embeddings varied weakly with the antigen. This suggests that the twin GCN model architecture over-weights the antigen node features, possibly due to the relative sparsity of antibody paratope residues relative to antigen epitope residues. This observation also emphasizes the need for an explicit attention mechanism such as that employed by EpiGraph to ensure that antibody-specific features are appropriately incorporated into the epitope prediction. Model Antibody Training Training Training AUPRC AUPRC AUPRC AUPRC specific data structures structures (epitope) (epitope) (paratope) (paratope)72 IPTS / 128951704.1Attorney Docket No.: SES-014WO PATENT EpiGraph Yes IEDB+SAb 4120 300 0.412 0.829 0.577 0.938 Ab-specific Dab GATValidation via wet lab experiments

[0134] A recent study has demonstrated that most existing tools for B cell epitope prediction perform no better than random guessing of protein surface patches. Since these models seemingly perform well in terms of in silico validation metrics, we sought to augment the in- silico measures of EpiGraph’s accuracy by comparing its epitope predictions to experimental results for de novo test cases. To that end, epitopes were experimentally mapped in both an antibody agnostic and antibody-specific fashion for a variety of antibodies whose epitopes have not previously been reported. Hydrogen-deuterium exchange

[0135] The low affinity immunoglobulin gamma Fc receptor FcɣRIIB plays a critical role in regulating antibody production by B-cells. An anti-FcɣRIIB benchmark antibody known as SM201 has been demonstrated to have a number of favorable biological properties that hypothetically derive from the specific epitope to which it binds on FcɣRIIB, but this epitope has not been publicly described. The antibody-specific epitope predictor was applied to map this epitope in silico, and then mapped the epitope experimentally via hydrogen-deuterium exchange mass spectrometry. The epitope predicted by the algorithm agreed well with the experimental results (FIG. 8, panel A). Crucially, the graph neural network assigned high probability to residues that differentiate FcɣRIIB from the highly homologous receptor FcɣRIIA, which SM201 is known not to bind.

[0136] The accuracy of this prediction was not the result of information leakage (for instance, the presence of a structure of an antibody bound to Fc_RIIb or a structurally homologous antigen at a homologous epitope). The Protein Data Bank (PDB) contains two publicly available structures of FcγRIIb in complex with an antibody. Both bind epitopes distant from that of SM201, indicating that EpiGraph’s accuracy is not the result of memorization. The PDB was 73 IPTS / 128951704.1Attorney Docket No.: SES-014WO PATENT searched for structurally homologous antigens using the global volume similarity measure based on BioZernike descriptors. Among the 44 structures containing proteins homologous to FcɣRIIB by this measure, only one was in complex with an antibody (antigen CD16a, PDB ID 7SEG), and the epitope on CD16a targeted by this antibody was distal from the site homologous to the SM201 epitope on FcɣRIIB, indicating that the accuracy of the graph neural network prediction was not influenced by information leakage from the training dataset. CryoEM

[0137] While HDX-MS effectively maps the epitope on the target antigen, it does not provide high-resolution information regarding the atomic details of the antibody-antigen interaction. To test the accuracy of the predictions at near-atomic resolution, cryogenic electron microscopy was performed to solve the structure of the T-cell activation surface protein marker PD-1 in complex with a novel antibody, whose epitope has not been previously reported. Again, the predictions of the model agree well with the experimental result, although several false positive residue predictions were observed (FIG. 8, panel B). Intriguingly, the model assigns the highest probability to a cluster of residues that exhibit the greatest extent of interaction with the antibody in the cryoEM structure, suggesting that the model has learned residue-level features relevant to specific antibody-antigen interactions rather than broadly selecting surface patches with general characteristics amenable to antibody binding. IVIG binding

[0138] As a final validation of the predictive models disclosed herein, the ability of the antibody-agnostic epitope prediction model was confirmed to identify the immunodominant regions of a protein – that is, the surface patches most likely to be bound by antibodies. As the example protein in this case, IdeS, a cysteine protease from the Gram positive human-infecting bacterium Streptococcus pyogenes that specifically cleaves all four classes of human immunoglobulin G at the Fc hinge region, was used. Because S. pyogenes is a human-infecting pathogen, most healthy donors have developed anti-IdeS IgG antibodies, which are refered to as pre-existing anti-drug antibodies, or ADAs for short.

[0139] Because of the polyclonal nature of the antibody response to IdeS, structural studies using isolated monoclonal antibodies are insufficient to experimentally map the 74 IPTS / 128951704.1Attorney Docket No.: SES-014WO PATENT immunodominant hotspots. Instead, mutagenesis studies were employed together with fluorescence-based measurements of polyclonal antibody binding to characterize the immunodominant surface regions on IdeS. Eight surface regions on the 3D crystal structure of the IdeS monomer (PDB 1Y08) were identified as likely immunodominant hotspots based on their solvent exposure, geometric structure, and biochemical composition, and systematically introduced mutations into these regions with the aim of interrupting antibody binding via alterations of the protein’s surface chemistry and conformation. The effects of these mutations on the binding of antibodies derived from the pooled serum of thousands of healthy donors (known as intravenous immunoglobulin or IVIG) were then measured, using an ELISA assay with optical density readout (OD450).

[0140] The resulting changes in binding of IVIG antibodies are illustrated in FIG. 9, panel A. The regions predicted by the model as the likely immunodominant hotspots indeed corresponded to areas where mutations significantly reduced IVIG antibody binding (regions A,D and G in FIG. 9, panel B). On the other hand, DiscoTope3, a popular antibody-agnostic epitope prediction tool, largely failed to accurately identify these regions. It instead prioritized regions near the enzymatic site of IdeS where it interacts with the Fc (regions E and F in FIG. 9, panel C), which have a lower average impact on antibody binding than the immunodominant regions predicted by the model.

[0141] This study developed and validated a set of recommender-style graph neural networks for the prediction of B-cell epitopes on general proteins. This recommender-style architecture has proven instrumental in obtaining antibody-specific epitope predictions, a significant advancement in the field of B cell epitope prediction. To date, methods to predict protein-protein interactions based on graph neural networks have employed twin neural network architectures with non-attentive graph convolutions. This study found that such architectures are unable to learn antibody-specific information, and recent benchmarking work has confirmed that such models do not produce reliable predictions.

[0142] The models outperform existing methods both in silico and via experimental validation, while also showing generalizability to a rigorously held-out test set. Importantly, the validation of the approach via wet lab experiments sets it apart from all other existing methods and provides robust evidence of the reliability of the predictions. In the first known experimental validation of an in silico B-cell epitope predictor, this study demonstrated EpiGraph’s accuracy in identifying 75 IPTS / 128951704.1Attorney Docket No.: SES-014WO PATENT the epitope on FcɣRIIB that is targeted by a benchmark antibody. Furthermore, the demonstration of EpiGraph’s ability to identify immunodominant hotspots on IdeS illustrates its applicability to the important task of reducing the immunogenicity of protein therapeutics.

[0143] EpiGraph could significantly improve the screening process for identifying antibodies with desired functionalities, thereby contributing to the development of more effective therapeutics. Moreover, it paves the way for de novo antibody design by supplying a reliable means to evaluate candidate antibody sequences in silico in terms of their likelihood to target the desired antigen and epitope.

[0144] This work has also highlighted the limitations of many existing epitope prediction tools, indicating a clear need for further research and methodological development in this area.

[0145] In conclusion, this study not only presents a novel and effective method for B cell epitope prediction but also underscores the importance of experimental validation in the development of computational epitope prediction tools. Example 2: The EpiGraph-Based Predictive Model.

[0146] The antibody-agnostic EpiGraph model takes an antigen structure as input, converts this structure to a graph representation, and predicts all potential immunodominant hotspots on the protein surface (FIG. 5, panel A). Antibody-specific EpiGraph model takes as input an antigen structure together with the structure of a binding antibody, featurizes these both as graphs, and predicts the specific epitope on the antigen surface to which the antibody binds as well as the antibody paratope, i.e. the antibody residues interacting with this epitope (FIG. 5, panel B). Example 3: CryoEM Structure of SM201-FcgRIIb Complex Confirms High Accuracy of the In Silico Epitope Prediction.

[0147] The prediction of the epitope for SM201 on FcγRIIβ obtained in Example 1 was previously validated via HDX, which is a low-resolution epitope-mapping method. A high- resolution cryoEM structure of SM201 bound to FcγRIIβ was determined using standard methods, as shown in FIG. 10. The cryoEM structure shows that the previously predicted epitope in Example 1 was accurate. Accordingly, the methods and systems provided herein are accurate at epitope prediction.

[0148] The entire disclosure of each of the patent and scientific documents referred to herein is incorporated by reference for all purposes. 76 IPTS / 128951704.1Attorney Docket No.: SES-014WO PATENT

[0149] The embodiments provided for herein may be embodied in other specific forms without departing from the spirit or essential characteristics thereof. The foregoing embodiments are therefore to be considered in all respects illustrative rather than limiting. Scope of the embodiments is thus indicated by the appended claims rather than by the foregoing description, and all changes that come within the meaning and range of equivalency of the claims are intended to be embraced therein. 77 IPTS / 128951704.1

Claims

Attorney Docket No.: SES-014WO PATENT CLAIMS 1. A method for identifying a B-cell epitope of a protein, the method comprising: (a) obtaining or having obtained a graph representation of the protein; (b) providing the graph representation of the protein to a generalized machine learning algorithm model to generate a plurality of predicted probabilities, wherein the generalized machine learning algorithm model is trained using a training data set comprising at least 1,000 loop‑mediated protein‑protein interaction interfaces; (c) classifying each of one or more amino acids of the protein as an epitope-member residue or not an epitope-member residue based on the plurality of predicted probabilities; and (d) generating a B-cell epitope prediction comprising the classified epitope-member residues.

2. The method of claim 1, wherein the step (a) comprises generating the graph representation of the protein comprising a monomeric structure of the protein.

3. The method of claim 2, wherein the step of generating a monomeric structure of the protein comprises generating a conformational variant of the monomeric structure of the protein, wherein each sidechain of the protein is in a minimized free energy conformation.

4. The method of claim 2, wherein the graph representation of the protein comprises a plurality of nodes representing each amino acid of the protein, and a plurality of edges between each pair of nodes.

5. The method of claim 2, wherein each edge represents a bond between each pair of nodes.

6. The method of claim 5, wherein the bond is a peptide bond.

7. The method of claim 1, wherein the step (a) comprises: (i) characterizing one or more structural features of the protein by performing a structural analysis of individual amino acids of the protein; and (ii) performing a pairwise amino acid to amino acid interaction analysis across at least one pair of amino acids in the protein to generate a feature graph. 78 IPTS / 128951704.1Attorney Docket No.: SES-014WO PATENT 8. The method of claim 7, wherein the one or more structural features are node features and / or edge features.

9. The method of any one of claim 8, wherein the method comprises generating a node feature by constructing a feature vector of each atom in the node, wherein the feature vector comprises atom type, distance from the amino acid’s alpha carbon, and curvature information.

10. The method of claim 9, wherein the feature vector of each atom in the node is aggregated to generate the one or more structural feature.

11. The method of any one of claim 7-10, wherein the node feature comprises a 3D coordinate of the amino acid, a secondary structure of the amin acid, and standard properties of the amino acid.

12. The method of claim 8, wherein the edge feature comprises presence of at least one interaction between at least one pair of atoms.

13. The method of claim 12, wherein the presence of the at least one interaction between the at least one pair of atoms comprises distance and angle between residue alpha-carbons, and / or an energy well.

14. The method of claim 13, wherein the at least one interaction is a peptide bond, a backbone carbonyl interaction, a salt bridge, a pi-pi interaction, a hydrogen bond, or an ionic interaction.

15. The method of claim 7, wherein the feature graph is determined by performing a transform of one or more node features and / or edge features as provided herein.

16. The method of claim 15, wherein the transform is a graph neural network (GNN).

17. The method of claim 16, wherein the graph neural network (GNN) is a message-passing graph neural network (MPGNN).

18. The method of claim 15, wherein the transform further comprises performing message passing via node and edge updates.

19. The method of claim 18, wherein the transform is a single-layer feedforward neural network.

20. The method of cany one of claims 14-19, wherein the feature graph is generated by iteratively analyzing each amino acid of the protein. 79 IPTS / 128951704.1Attorney Docket No.: SES-014WO PATENT 21. The method of claim 1, wherein the generalized machine learning algorithm model comprises unsupervised learning, supervised learning, semi-supervised learning, reinforcement learning, transfer learning, incremental learning, curriculum learning, and learning to learn.

22. The method of claim 1, the generalized machine learning algorithm model comprises linear classifiers, logistic classifiers, Bayesian networks, random forest, neural networks, graph neural networks (GNN), message-passing graph neural networks (MPGNN), matrix factorization, hidden Markov model, support vector machine, K-means clustering, or K-nearest neighbor.

23. The method of any one of claims 1-22, wherein the training data set comprises at least 10000 loop‑mediated protein‑protein interaction interfaces.

24. The method of any one of claims 1-23, wherein the generalized machine learning algorithm model trained using the training data is further refined for antibody-antigen interactions.

25. The method of any one of claims 1-24, wherein the step (b) comprises generating a probability that the amino acid is the epitope-member.

26. The method of claim 1, wherein the step (d) comprises generating a set of amino acids that belong to the epitope.

27. The method of claim 1, wherein the machine learning algorithm model has been trained using at least 800 annotated B cell epitopes.

28. The method of any one of claims 1-27, wherein the method is characterized by an AUROC of at least 0.

705.

29. The method of any one of claims 1-27, wherein the method is characterized by an AUPRC of at least 0.

256.

30. A method for identifying a B-cell epitope of a protein, and a paratope of a binder of the epitope, the method comprising: (a) obtaining or having obtained a graph representation of the protein and a graph representation of the binder; (b) transforming the obtained graph representation of the protein and the graph representation of the binder to construct a bipartite graph comprising: a plurality of protein embeddings; a plurality of binder embeddings; and 80 IPTS / 128951704.1Attorney Docket No.: SES-014WO PATENT a plurality of edges connecting protein embeddings and binder embeddings, wherein the plurality of edges comprise predicted weights determined through graph-based collaborative filtering; (c) providing the bipartite graph to a machine learning model to generate a plurality of predicted probabilities; (d) classifying each amino acid of the protein as an epitope-member residue or not an epitope-member residue, and each amino acid of the binder as a paratope-member or not a paratope-member based on the plurality of predicted probabilities; and (e) generating a B-cell epitope prediction comprising the classified epitope-member residues.

31. The method of claim 30, wherein the step (a) comprises generating a graph representation of the protein and a graph representation of the binder, each comprising a monomeric structure of the protein and the antibody.

32. The method of claim 31, wherein the step of generating a monomeric structure of the protein and the antibody comprises generating a conformational variant of each the monomeric structure of the protein and the monomeric structure of the antibody, wherein each sidechain of the protein and each sidechain of the antibody is in a minimized free energy conformation.

33. The method of claim 31, wherein the a graph representation of the protein and a graph representation of the binder each comprises a plurality of nodes representing each amino acid of the protein and the antibody, and a plurality of edges between each pair of nodes.

34. The method of claim 33, wherein each edge represents a bond between each pair of nodes.

35. The method of claim 34, wherein the bond is a peptide bond.

36. The method of claim 30, wherein the step (a) comprises: (i) characterizing one or more structural features of the protein and the antibody by independently performing a structural analysis of individual amino acids of the protein and the antibody; and (ii) performing a pairwise amino acid to amino acid interaction analysis across at least one pair of amino acids in the protein and the antibody to generate an feature graph for each the protein and the antibody. 81 IPTS / 128951704.1Attorney Docket No.: SES-014WO PATENT 37. The method of claim 36, wherein the one or more structural features are node features and / or edge features.

38. The method of any one of claim 36 or 37, wherein the method comprises generating a node feature by constructing a feature vector of each atom in the node, wherein the feature vector comprises atom type, distance from the amino acid’s alpha carbon, and curvature information.

39. The method of claim 38, wherein the feature vector of each atom in the node is aggregated to generate the one or more structural feature.

40. The method of any one of claim 36-39, wherein the node feature comprise a 3D coordinate of the amino acid, a secondary structure of the amin acid, and standard properties of the amino acid.

41. The method of claim 36, wherein the edge feature comprises presence of at least one interaction between at least one pair of atoms.

42. The method of claim 41, wherein the presence of the at least one interaction between the at least one pair of atoms comprises distance and angle between residue alpha-carbons, and / or an energy well.

43. The method of claim 42, wherein the at least one interaction is a peptide bond, a backbone carbonyl interaction, a salt bridge, a pi-pi interaction, a hydrogen bond, or an ionic interaction.

44. The method of claim 36, wherein the feature graph is determined by performing a transform of one or more node features and / or edge features as provided herein.

45. The method of claim 44, wherein the transform is a graph neural network (GNN).

46. The method of claim 45, wherein the graph neural network (GNN) is a message-passing graph neural network (MPGNN).

47. The method of claim 44, wherein the transform further comprises performing message passing via node and edge updates.

48. The method of claim 47, wherein the transform is a single-layer feedforward neural network.

49. The method of cany one of claims 36-48, wherein the feature graph is generated by iteratively analyzing each amino acid of the protein. 82 IPTS / 128951704.1Attorney Docket No.: SES-014WO PATENT 50. The method of claim 30, wherein the method further comprises, prior to step (c), predicting edge weights of the bipartite graph by performing a transform of the bipartite graph.

51. The method of claim 50, wherein the transform is a graph attention network (GAT).

52. The method of claim 30, wherein the machine learning model comprises unsupervised learning, supervised learning, semi-supervised learning, reinforcement learning, transfer learning, incremental learning, curriculum learning, and learning to learn.

53. The method of claim 30, the machine learning model comprises linear classifiers, logistic classifiers, Bayesian networks, random forest, neural networks, graph neural networks (GNN), message-passing graph neural networks (MPGNN), matrix factorization, hidden Markov model, support vector machine, K-means clustering, or K-nearest neighbor.

54. The method of any one of claims 30-53, wherein the machine learning model is trained using a training data set comprising at least 1,000 loop‑mediated protein‑protein interaction interfaces 55. The method of claim 54, wherein the training data set comprises at least 10000 loop‑mediated protein‑protein interaction interfaces.

56. The method of any one of claims 30-55, wherein the machine learning model trained using the training data is further refined for antibody-antigen interactions.

57. The method of claim 30, wherein the step (c) comprises generating a probability that the amino acid of the protein is an epitope-member and that the amino acid of the antibody is a paratope-member.

58. The method of claim 30, wherein the step (e) comprises generating a set of amino acids that belong to the epitope and a set of amino acids that belong to the paratope.

59. The method of claim 58, wherein the epitope and the paratope interact with each other.

60. The method of claim 30, wherein the machine learning algorithm model has been trained using at least 4000 annotated B cell epitopes.

61. The method of any one of claims 30-60, wherein the method is characterized by an epitope AUROC of at least 0.761, at least 0.829, or at least 0.

844.

62. The method of any one of claims 30-60, wherein the method is characterized by an epitope AUPRC of at least 0.272, or at least 0.

412. 83 IPTS / 128951704.1Attorney Docket No.: SES-014WO PATENT 63. The method of any one of claims 30-60, wherein the method is characterized by a paratope AUROC of at least 0.923, at least 0.938, or at least 0.

945.

64. The method of any one of claims 30-60, wherein the method is characterized by a paratope AUPRC of at least 0.510, at least 0.577, or at least 0.

581.

65. A non-transitory computer-readable storage medium, the computer-readable storage medium comprising instructions that when executed by a processor, cause the processor to: (a) obtain or having obtained a graph representation of a protein; (b) providing the graph representation of the protein to a generalized machine learning algorithm model to generate a plurality of predicted probabilities, wherein the generalized machine learning algorithm model is trained using a training data set comprising at least 1,000 loop‑mediated protein‑protein interaction interfaces; (c) classify each of one or more amino acids of the protein as an epitope-member residue or not an epitope-member residue based on the plurality of predicted probabilities; and (d) generate a B-cell epitope prediction comprising the classified epitope-member residues.

66. The non-transitory computer-readable storage medium of claim 65, wherein the step (a) comprises generating the graph representation of a protein comprising a monomeric structure of the protein.

67. The non-transitory computer-readable storage medium of claim 66, wherein the step of generating a monomeric structure of the protein comprises generating a conformational variant of the monomeric structure of the protein, wherein each sidechain of the protein is in a minimized free energy conformation.

68. The non-transitory computer-readable storage medium of claim 66, wherein the graph representation of a protein comprises a plurality of nodes representing each amino acid of the protein, and a plurality of edges between each pair of nodes.

69. The non-transitory computer-readable storage medium of claim 66, wherein each edge represents a bond between each pair of nodes. 84 IPTS / 128951704.1Attorney Docket No.: SES-014WO PATENT 70. The non-transitory computer-readable storage medium of claim 69, wherein the bond is a peptide bond.

71. The non-transitory computer-readable storage medium of claim 65, wherein the step (a) comprises: (i) characterizing one or more structural features of the protein by performing a structural analysis of individual amino acids of the protein; and (ii) performing a pairwise amino acid to amino acid interaction analysis across at least one pair of amino acids in the protein to generate a feature graph.

72. The non-transitory computer-readable storage medium of claim 71, wherein the one or more structural features are node features and / or edge features.

73. The non-transitory computer-readable storage medium of any one of claim 72, wherein the non-transitory computer-readable storage medium comprises generating a node feature by constructing a feature vector of each atom in the node, wherein the feature vector comprises atom type, distance from the amino acid’s alpha carbon, and curvature information.

74. The non-transitory computer-readable storage medium of claim 73, wherein the feature vector of each atom in the node is aggregated to generate the one or more structural feature.

75. The non-transitory computer-readable storage medium of any one of claim 71-74, wherein the node feature comprise a 3D coordinate of the amino acid, a secondary structure of the amin acid, and standard properties of the amino acid.

76. The non-transitory computer-readable storage medium of claim 73, wherein the edge feature comprises presence of at least one interaction between at least one pair of atoms.

77. The non-transitory computer-readable storage medium of claim 76, wherein the presence of the at least one interaction between the at least one pair of atoms comprises distance and angle between residue alpha-carbons, and / or an energy well.

78. The non-transitory computer-readable storage medium of claim 77, wherein the at least one interaction is a peptide bond, a backbone carbonyl interaction, a salt bridge, a pi-pi interaction, a hydrogen bond, or an ionic interaction. 85 IPTS / 128951704.1Attorney Docket No.: SES-014WO PATENT 79. The non-transitory computer-readable storage medium of claim 78, wherein the feature graph is determined by performing a transform of one or more node features and / or edge features as provided herein.

80. The non-transitory computer-readable storage medium of claim 79, wherein the transform is a graph neural network (GNN).

81. The non-transitory computer-readable storage medium of claim 80, wherein the graph neural network (GNN) is a message-passing graph neural network (MPGNN).

82. The non-transitory computer-readable storage medium of claim 79, wherein the transform further comprises performing message passing via node and edge updates.

83. The non-transitory computer-readable storage medium of claim 82, wherein the transform is a single-layer feedforward neural network.

84. The non-transitory computer-readable storage medium of cany one of claims 78-84, wherein the feature graph is generated by iteratively analyzing each amino acid of the protein.

85. The non-transitory computer-readable storage medium of claim 65, wherein the generalized machine learning algorithm model comprises unsupervised learning, supervised learning, semi- supervised learning, reinforcement learning, transfer learning, incremental learning, curriculum learning, and learning to learn.

86. The non-transitory computer-readable storage medium of claim 65, the generalized machine learning algorithm model comprises linear classifiers, logistic classifiers, Bayesian networks, random forest, neural networks, graph neural networks (GNN), message-passing graph neural networks (MPGNN), matrix factorization, hidden Markov model, support vector machine, K- means clustering, or K-nearest neighbor.

87. The non-transitory computer-readable storage medium of any one of claims 65-86, wherein the training data set comprises at least 10000 loop‑mediated protein‑protein interaction interfaces.

88. The non-transitory computer-readable storage medium of any one of claims 65-87, wherein the generalized machine learning algorithm model trained using the training data is further refined for antibody-antigen interactions. 86 IPTS / 128951704.1Attorney Docket No.: SES-014WO PATENT 89. The non-transitory computer-readable storage medium of claim 65, wherein the step (b) comprises generating a probability that the amino acid is the epitope-member.

90. The non-transitory computer-readable storage medium of claim 65, wherein the step (d) comprises generating a set of amino acids that belong to the epitope.

91. The non-transitory computer-readable storage medium of claim 65, wherein the machine learning algorithm model has been trained using at least 800 annotated B cell epitopes.

92. The non-transitory computer-readable storage medium of any one of claims 65-91, wherein performance of the B-cell epitope prediction is characterized by an AUROC of at least 0.

705.

93. The non-transitory computer-readable storage medium of any one of claims 65-91, wherein performance of the B-cell epitope prediction is characterized by an AUPRC of at least 0.

256.

94. A non-transitory computer-readable storage medium, the computer-readable storage medium comprising instructions that when executed by a processor, cause the processor to: (a) obtain or having obtained a graph representation of a protein and a graph representation of a binder; (b) transform the obtained the graph representation of the protein and the graph representation of the binder to construct a bipartite graph comprising: a plurality of protein embeddings; a plurality of binder embeddings; and a plurality of edges connecting protein embeddings and binder embeddings, wherein the plurality of edges comprise predicted weights determined through graph-based collaborative filtering; (c) provide the bipartite graph to a machine learning model to generate a plurality of predicted probabilities; (d) classify each amino acid of the protein as an epitope-member residue or not an epitope-member residue, and each amino acid of the binder as a paratope-member or not a paratope-member based on the plurality of predicted probabilities; and (e) generate a B-cell epitope prediction comprising the classified epitope-member residues. 87 IPTS / 128951704.1Attorney Docket No.: SES-014WO PATENT 95. The non-transitory computer-readable storage medium of claim 94, wherein the step (a) comprises generating the graph representation of the protein and the graph representation of the binder, each comprising a monomeric structure of the protein and the antibody.

96. The non-transitory computer-readable storage medium of claim 95, wherein the step of generating a monomeric structure of the protein and the antibody comprises generating a conformational variant of each the monomeric structure of the protein and the monomeric structure of the antibody, wherein each sidechain of the protein and each sidechain of the antibody is in a minimized free energy conformation.

97. The non-transitory computer-readable storage medium of claim 95, wherein the the graph representation of the protein and the graph representation of the binder each comprises a plurality of nodes representing each amino acid of the protein and the antibody, and a plurality of edges between each pair of nodes.

98. The non-transitory computer-readable storage medium of claim 95, wherein each edge represents a bond between each pair of nodes.

99. The non-transitory computer-readable storage medium of claim 98, wherein the bond is a peptide bond.

100. The non-transitory computer-readable storage medium of claim 94, wherein the step (a) comprises: (i) characterizing one or more structural features of the protein and the antibody by independently performing a structural analysis of individual amino acids of the protein and the antibody; and (ii) performing a pairwise amino acid to amino acid interaction analysis across at least one pair of amino acids in the protein and the antibody to generate an feature graph for each the protein and the antibody.

101. The non-transitory computer-readable storage medium of claim 100, wherein the one or more structural features are node features and / or edge features.

102. The non-transitory computer-readable storage medium of any one of claim 101, wherein the non-transitory computer-readable storage medium comprises generating a node feature by 88 IPTS / 128951704.1Attorney Docket No.: SES-014WO PATENT constructing a feature vector of each atom in the node, wherein the feature vector comprises atom type, distance from the amino acid’s alpha carbon, and curvature information.

103. The non-transitory computer-readable storage medium of claim 101, wherein the feature vector of each atom in the node is aggregated to generate the one or more structural feature.

104. The non-transitory computer-readable storage medium of any one of claim 100-103, wherein the node feature comprise a 3D coordinate of the amino acid, a secondary structure of the amin acid, and standard properties of the amino acid.

105. The non-transitory computer-readable storage medium of claim 101, wherein the edge feature comprises presence of at least one interaction between at least one pair of atoms.

106. The non-transitory computer-readable storage medium of claim 105, wherein the presence of the at least one interaction between the at least one pair of atoms comprises distance and angle between residue alpha-carbons, and / or an energy well.

107. The non-transitory computer-readable storage medium of claim 106, wherein the at least one interaction is a peptide bond, a backbone carbonyl interaction, a salt bridge, a pi-pi interaction, a hydrogen bond, or an ionic interaction.

108. The non-transitory computer-readable storage medium of claim 107, wherein the feature graph is determined by performing a transform of one or more node features and / or edge features as provided herein.

109. The non-transitory computer-readable storage medium of claim 108, wherein the transform is a graph neural network (GNN).

110. The non-transitory computer-readable storage medium of claim 109, wherein the graph neural network (GNN) is a message-passing graph neural network (MPGNN).

111. The non-transitory computer-readable storage medium of claim 108, wherein the transform further comprises performing message passing via node and edge updates.

112. The non-transitory computer-readable storage medium of claim 111, wherein the transform is a single-layer feedforward neural network.

113. The non-transitory computer-readable storage medium of cany one of claims 107-112, wherein the feature graph is generated by iteratively analyzing each amino acid of the protein. 89 IPTS / 128951704.1Attorney Docket No.: SES-014WO PATENT 114. The non-transitory computer-readable storage medium of claim 94, wherein the non- transitory computer-readable storage medium further comprises, prior to step (c), predicting edge weights of the bipartite graph by performing a transform of the bipartite graph.

115. The non-transitory computer-readable storage medium of claim 114, wherein the transform is a graph attention network (GAT).

116. The non-transitory computer-readable storage medium of claim 94, wherein the machine learning model comprises unsupervised learning, supervised learning, semi-supervised learning, reinforcement learning, transfer learning, incremental learning, curriculum learning, and learning to learn.

117. The non-transitory computer-readable storage medium of claim 94, the machine learning model comprises linear classifiers, logistic classifiers, Bayesian networks, random forest, neural networks, graph neural networks (GNN), message-passing graph neural networks (MPGNN), matrix factorization, hidden Markov model, support vector machine, K-means clustering, or K- nearest neighbor.

118. The non-transitory computer-readable storage medium of any one of claims 94-117, wherein the machine learning model is trained using a training data set comprising at least 1,000 loop‑mediated protein‑protein interaction interfaces 119. The non-transitory computer-readable storage medium of claim 118, wherein the training data set comprises at least 10000 loop‑mediated protein‑protein interaction interfaces.

120. The non-transitory computer-readable storage medium of any one of claims 94-119, wherein the machine learning model trained using the training data is further refined for antibody-antigen interactions.

121. The non-transitory computer-readable storage medium of claim 94, wherein the step (c) comprises generating a probability that the amino acid of the protein is an epitope-member and that the amino acid of the antibody is a paratope-member.

122. The non-transitory computer-readable storage medium of claim 94, wherein the step (e) comprises generating a set of amino acids that belong to the epitope and a set of amino acids that belong to the paratope. 90 IPTS / 128951704.1Attorney Docket No.: SES-014WO PATENT 123. The non-transitory computer-readable storage medium of claim 122, wherein the epitope and the paratope interact with each other.

124. The non-transitory computer-readable storage medium of claim 94, wherein the machine learning algorithm model has been trained using at least 4000 annotated B cell epitopes.

125. The non-transitory computer-readable storage medium of any one of claims 94-124, wherein performance of the B-cell epitope prediction is characterized by an epitope AUROC of at least 0.761, at least 0.829, or at least 0.

844.

126. The non-transitory computer-readable storage medium of any one of claims 94-124, wherein performance of the B-cell epitope prediction is characterized by an epitope AUPRC of at least 0.272, or at least 0.

412.

127. The non-transitory computer-readable storage medium of any one of claims 94-124, wherein performance of the B-cell epitope prediction is characterized by a paratope AUROC of at least 0.923, at least 0.938, or at least 0.

945.

128. The non-transitory computer-readable storage medium of any one of claims 94-124, wherein performance of the B-cell epitope prediction is characterized by a paratope AUPRC of at least 0.510, at least 0.577, or at least 0.

581.

129. A system comprising a non-transitory computer-readable storage medium and a processor, wherein the non-transitory computer-readable storage medium comprises: a) a bipartite graph encoded on the non-transitory computer-readable storage medium; and b) instructions for generating a B-cell epitope prediction comprising a classified epitope-member residues from bipartite graph, wherein the bipartite graph is a transformation of a graph representation of a protein and a graph representation of a binder.

130. The system of claim 129, wherein the encoded bipartite graph comprises: a plurality of protein embeddings; a plurality of binder embeddings; and 91 IPTS / 128951704.1Attorney Docket No.: SES-014WO PATENT a plurality of edges connecting protein embeddings and binder embeddings, wherein the plurality of edges comprise predicted weights determined through graph-based collaborative filtering.

131. The system of claims 129 or 130, wherein the instructions for generating a B-cell epitope prediction comprise instructions for a) providing the bipartite graph to a machine learning model to generate a plurality of predicted probabilities; b) classifying each amino acid of the protein as an epitope-member residue or not an epitope-member residue, and each amino acid of the binder as a paratope-member or not a paratope-member based on the plurality of predicted probabilities; and c) generating the B-cell epitope prediction comprising the classified epitope-member residues.

132. The system of any one of claims 129-131, wherein the transformation of the graph representation of the protein and the graph representation of the binder comprises performing a transform of the bipartite graph.

133. The system of claim 132, wherein the transform is a graph attention network (GAT). 92 IPTS / 128951704.1