Guided machine learning enabled generation of antigen receptors
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2026-02-09
- Publication Date
- 2026-08-13
Smart Images

Figure US2026014496_13082026_PF_FP_ABST
Abstract
Description
Attorney Ref.: 14786-074-228 (I03963-228074) / P40073-W0-1GUIDED MACHINE LEARNING ENABLED GENERATION OF ANTIGEN RECEPTORS CROSS REFERENCE TO RELATED APPLICATION
[0001] This application claims priority to U.S. Provisional Application No. 63 / 756,594, entitled “GUIDED MACHINE LEARNING ENABLED GENERATION OF ANTIGEN RECEPTORS” and filed on February 10, 2025, the disclosure of which is incorporated herein by reference in its entirety.TECHNICAL FIELD
[0002] The subject matter described herein relates generally to generative artificial intelligence and more specifically to guided machine learning enabled generation of antigen receptors.INTRODUCTION
[0003] A molecule is a group of two more atoms held together by chemical bonds. Molecules form the smallest identifiable unit into which a pure substance can be divided while still retaining the composition and chemical properties of that substance. Large molecules refer to those molecules that range between approximately 3000 Daltons and 150,000 Daltons in molecular weight. Large molecule therapeutics (also known as biopharmaceuticals, biotherapeutics, biologicals, or biologies) are often derivatives of natural human proteins, which modulate many essential cellular functions such as enzymatic reactions, transport of molecules, regulation and execution of a number of biological pathways, cell growth, proliferation, nutrient uptake, morphology, motility, intercellular communication, and / or the like. A single large molecule can include more than 1,300 amino acid residues linked by peptide bonds to form one or moreAttorney Ref.: 14786-074-228 (I03963-228074) / P40073-W0-1polypeptide. Due to their size and complexity, large molecule therapeutics are recombinantly produced by engineered cells instead of being chemically synthesized like the majority of small molecule drugs. Moreover, large molecule therapeutics are usually delivered through injection or infusion due to the ineffectiveness of oral administration. The development of a large molecule therapeutics may entail designing one or more sequences of amino acid residues capable of binding to a target (e.g., a protein, a nucleic acid, and / or the like) with sufficient specificity and absent undesirable traits such as immunogenicity, self-association, instability, and / or the like.SUMMARY
[0006] Systems, methods, and articles of manufacture, including computer program products, are provided for machine learning enabled generation of antigen receptors, such as B-cell receptors (BCRs) and T-cell receptors (TCRs), with context guidance. A viable protein therapeutic may be required to exhibit one or more properties of interest including, for example , affinity, avidity, anti-pathogen activity, and / or the like. In some cases, a protein design computation model may be trained to approximate a data distribution from a training dataset of known protein sequences, such as protein sequences exhibiting one or more properties of interest. However, applying the trained protein design computation model to sample indiscriminately from this data distribution may yield protein sequences lacking one or more properties of interest. Accordingly, in some cases, the generating of protein sequences by the protein design computation model may be conditioned on one or more contextual variables. In some cases, the one or more contextual variables may define a context associated with the one or more properties of interest including, for example, one or more target clonotypes (or clonal families), maturity, and / or the like. In some cases, the one or more contextual variables may include categorical variables, numerical variables, and / or the like. For instance, in some cases, protein sequences originatingAttorney Ref.: 14786-074-228 (I03963-228074) / P40073-W0-1from a target clonotype (or clonal family) or those produced by mature lymphocytes that have undergo affinity maturation may be more likely to exhibit one or more properties of interest. As such, in some cases, conditioning on the one or more contextual variables may increase the likelihood that the protein sequences generated by the protein design computation model can be developed into viable protein therapeutics.
[0007] In one aspect, there is provide a system for machine learning enabled generation of antigen receptors with context guidance. The system may include at least one data processor and at least one memory. The at least one memory may store instructions that result in operations when executed by the at least one data processor. The operations may include: identifying an antigen receptor repertoire having a plurality of clonotypes; identifying a target clonotype within the plurality of clonotypes in the antigen receptor repertoire; training a protein design computation model to approximate a data distribution of protein sequences of antigen receptors expressed by lymphocytes from the target clonotype; and applying the protein design computation model to generate, based at least on an input protein sequence, an output protein sequence corresponding to at least a portion of an antigen receptor expressed by lymphocytes from the target clonotype.
[0008] In another aspect, there is provide a computer-implemented method for machine learning enabled generation of antigen receptors with context guidance. The method may include: identifying an antigen receptor repertoire having a plurality of clonotypes; identifying a target clonotype within the plurality of clonotypes in the antigen receptor repertoire; training a protein design computation model to approximate a data distribution of protein sequences of antigen receptors expressed by lymphocytes from the target clonotype; and applying the protein design computation model to generate, based at least on an input protein sequence, an output proteinAttorney Ref.: 14786-074-228 (I03963-228074) / P40073-W0-1sequence corresponding to at least a portion of an antigen receptor expressed by lymphocytes from the target clonotype.
[0009] In another aspect, there is provided a computer program product for machine learning enabled generation of antigen receptors with context guidance. The computer program product may include a non-transitory computer readable medium storing instructions that result in operations when executed by at least one data processor. The operations may include: identifying an antigen receptor repertoire having a plurality of clonotypes; identifying a target clonotype within the plurality of clonotypes in the antigen receptor repertoire; training a protein design computation model to approximate a data distribution of protein sequences of antigen receptors expressed by lymphocytes from the target clonotype; and applying the protein design computation model to generate, based at least on an input protein sequence, an output protein sequence corresponding to at least a portion of an antigen receptor expressed by lymphocytes from the target clonotype.
[0010] In some variations, one or more features disclosed herein including the following features can optionally be included in any feasible combination.
[0011] In some variations, the protein design computation model includes a clonotype energy-based model (EBM) trained to approximate the data distribution of the protein sequences of antigen receptors expressed by lymphocytes from the target clonotype.
[0012] In some variations, the clonotype energy-based model (EBM) parameterizes an energy function that outputs a value indicative of whether an antigen receptor corresponding to a protein sequence generated from the input protein sequence is expressed by lymphocytes from the target clonotype.Attorney Ref.: 14786-074-228 (I03963-228074) / P40073-W0-1
[0013] In some variations, the energy function outputs a higher value for a protein sequence drawn from a higher density region of the data distribution populated by protein sequences of antigen receptors expressed by lymphocytes from the target clonotype.
[0014] In some variations, the energy function outputs a lower value for a protein sequence drawn from a lower density region of the data distribution populated by protein sequences of antigen receptors expressed by lymphocytes outside of the target clonotype.
[0015] In some variations, the output of the energy function comprises a value corresponding to a density at a location in the data distribution from which a protein sequence generated from the input protein sequence is drawn.
[0016] In some variations, the protein design computation model includes a score-based model trained to approximate the data distribution of protein sequences of antigen receptors expressed by lymphocytes from the target clonotype.
[0017] In some variations, the output of the score function comprises a value corresponding to a local density change at a location in the data distribution from which a protein sequence generated from the input protein sequence is drawn.
[0018] In some variations, the score-based model comprises a conditional score-based model trained with conditioning on the target clonotype and an unconditional score-based model trained without conditioning on the target clonotype.
[0019] In some variations, the score-based model comprises a linear combination of the conditional score-based model and the unconditional score-based model.
[0020] In some variations, a respective output of the conditional score-based model and the unconditional score-based model are weighted by a weight corresponding to a guidance level from the target clonotype.Attorney Ref.: 14786-074-228 (I03963-228074) / P40073-W0-1
[0021] In some variations, the output protein sequence is generated by modifying the input protein sequence over multiple iterations of gradient-based Markov Chain Monte Carlo (MCMC) sampling.
[0022] In some variations, the output protein sequence is generated by at least applying the protein design computation model to generate a first modified protein sequence having a first modification to the input protein sequence, applying the protein design computation model to generate a second modified protein sequence having a second modification to the input protein sequence, determining that an antigen receptor corresponding to the first modified protein sequence is more likely to be expressed by lymphocytes from the target clonotype than an antigen receptor corresponding to the second modified protein sequence, and applying the protein design computation model to further modify the first modified protein sequence instead of the second modified protein sequence.
[0023] In some variations, each of the first modification and the second modification include inserting, deleting, and / or changing a type of one or more amino acid residues in the input protein sequence.
[0024] In some variations, the output protein sequence is generated to correspond to the first modified protein sequence upon satisfying one or more criteria.
[0025] In some variations, the one or more criteria include at least one of (i) the input protein sequence having undergone a threshold quantity of gradient-based Markov Chain Monte Carlo (MCMC) sampling iterations to generate the first modified protein sequence, and (ii) the first modified protein sequence corresponding to at least the portion of the antigen receptor expressed by lymphocytes from the target clonotype.Attorney Ref.: 14786-074-228 (I03963-228074) / P40073-W0-1
[0026] In some variations, the data distribution comprises a noisy data distribution populated by noisy protein sequences of antigen receptors from the antigen receptor repertoire.
[0027] In some variations, the generating of the output protein sequence includes denoising the first modified protein sequence.
[0028] In some variations, the protein design computation model is trained to approximate a data distribution of protein sequences of antigen receptors expressed by mature lymphocytes that have undergone adequate affinity maturation.
[0029] In some variations, the protein design computation model includes a maturity energy-based model (EBM) trained to approximate the data distribution of protein sequences of antigen receptors expressed by mature lymphocytes.
[0030] In some variations, the maturity energy-based model (EBM) parameterizes an energy function that outputs a value indicative of whether an antigen receptor corresponding to a protein sequence generated from the input protein sequence is expressed by mature lymphocytes.
[0031] In some variations, the energy function outputs a higher value for a protein sequence drawn from a higher density region of the data distribution populated by protein sequences of antigen receptors expressed by mature lymphocytes.
[0032] In some variations, the energy function outputs a lower value for a protein sequence drawn from a lower density region of the data distribution populated by protein sequences of antigen receptors expressed by immature lymphocytes.
[0033] In some variations, the generating of the output protein sequence is guided by a composition of an energy function parameterized by the clonotype energy-based model (EBM) and the energy function parameterized by the maturity energy-based model (EBM) such that theAttorney Ref.: 14786-074-228 (I03963-228074) / P40073-W0-1output protein sequence corresponds to an antigen receptor expressed by a mature lymphocyte from the target clonotype.
[0034] In some variations, the composition comprises a product of experts or a weighted sum.
[0035] In some variations, a maturity of a lymphocyte is defined by one or more of (i) a quantity of somatic hypermutations (SHM) present in the lymphocyte, and (ii) a depth in a lymphocyte phylogenetic tree occupied by the lymphocyte.
[0036] In some variations, the target clonotype is defined by (i) a variable (V) gene segment, (ii) a joining (J) gene segment, and (ii) a third complementarity determining region (CDR3) length.
[0037] In some variations, the target clonotype comprises a clonotype that is enriched in the antigen receptor repertoire.
[0038] In some variations, the protein design computation model is trained to approximate the data distribution based on a plurality of positive samples.
[0039] In some variations, each positive sample includes a protein sequence of an antigen receptor paired with a correct clonotype.
[0040] In some variations, the protein design computation model is trained to approximate the data distribution based on a plurality of negative samples.
[0041] In some variations, each negative sample includes a protein sequence of an antigen receptor paired with an incorrect clonotype.
[0042] In some variations, the incorrect clonotype comprises a random clonotype or a different clonotype from a same clonal family as a correct clonotype.Attorney Ref.: 14786-074-228 (I03963-228074) / P40073-W0-1
[0043] In one aspect, there is provide a system for machine learning enabled generation of antigen receptors with context guidance. The system may include at least one data processor and at least one memory. The at least one memory may store instructions that result in operations when executed by the at least one data processor. The operations may include: identifying an antigen receptor repertoire having a plurality of different contexts; identifying a target context within the plurality of different contexts present in the antigen receptor repertoire; training a protein design computation model to approximate a data distribution of protein sequences of antigen receptors exhibiting the target context; and applying the protein design computation model to generate, based at least on an input protein sequence, an output protein sequence corresponding to at least a portion of an antigen receptor having the target context.
[0044] In another aspect, there is provide a computer-implemented method for machine learning enabled generation of antigen receptors with context guidance. The method may include: identifying an antigen receptor repertoire having a plurality of different contexts; identifying a target context within the plurality of different contexts present in the antigen receptor repertoire; training a protein design computation model to approximate a data distribution of protein sequences of antigen receptors exhibiting the target context; and applying the protein design computation model to generate, based at least on an input protein sequence, an output protein sequence corresponding to at least a portion of an antigen receptor having the target context.
[0045] In another aspect, there is provided a computer program product for machine learning enabled generation of antigen receptors with context guidance. The computer program product may include a non-transitory computer readable medium storing instructions that result in operations when executed by at least one data processor. The operations may include: identifying an antigen receptor repertoire having a plurality of different contexts; identifying a target contextAttorney Ref.: 14786-074-228 (I03963-228074) / P40073-W0-1within the plurality of different contexts present in the antigen receptor repertoire; training a protein design computation model to approximate a data distribution of protein sequences of antigen receptors exhibiting the target context; and applying the protein design computation model to generate, based at least on an input protein sequence, an output protein sequence corresponding to at least a portion of an antigen receptor having the target context.
[0046] In some variations, one or more features disclosed herein including the following features can optionally be included in any feasible combination.
[0047] In some variations, the target context is defined by one or more contextual variables.
[0048] In some variations, the one or more contextual variables include one or more of a categorical variable and a numerical variable.
[0049] In some variations, the target context comprises a target clonotype of lymphocytes expressing the antigen receptor.
[0050] In some variations, the target clonotype is defined by one or more of a variable (V) gene segment, a joint (J) gene segment, and a third complementarity determining region (CDR3) length.
[0051] In some variations, the target context comprise a maturity of lymphocytes expressing the antigen receptor.
[0052] In some variations, the maturity of a lymphocyte is defined by one or more of a quantity of somatic hypermutations (SHMs) present in an antigen receptor coding sequence of the lymphocyte relative to its ancestral lymphocyte, and a depth of the lymphocyte in a lymphocyte phylogenetic tree depicting the evolutionary relationships within the antigen receptor repertoire.Attorney Ref.: 14786-074-228 (I03963-228074) / P40073-W0-1
[0053] In some variations, the target context includes a target clonotype and a maturity of lymphocytes expressing the antigen receptor.
[0054] In some variations, the protein design computation model includes an energybased model (EBM) trained to approximate the data distribution of protein sequences of antigen receptors exhibiting the target context.
[0055] In some variations, the energy-based model (EBM) parameterizes an energy function that outputs a value indicative of whether an antigen receptor corresponding to a protein sequence generated from the input sequence exhibits the target context.
[0056] In some variations, the energy function outputs a higher value for a protein sequence drawn from a higher density region of the data distribution populated by protein sequences of antigen receptors exhibiting the target context
[0057] In some variations, the energy function outputs a lower value for a protein sequence drawn from a lower density region of the data distribution populated by protein sequences of antigen receptors without the target context
[0058] In some variations, the output of the energy function comprises a value corresponding to a density at a location in the data distribution from which a protein sequence generated from the input protein sequence is drawn.
[0059] In some variations, the protein design computation model includes a score-based function trained to approximate data the distribution of protein sequences of antigen receptors exhibiting the target context.
[0060] In some variations, the output of the score function comprises a value corresponding to a local density change at a location in the data distribution from which a protein sequence generated from the input protein sequence is drawn.Attorney Ref.: 14786-074-228 (I03963-228074) / P40073-W0-1
[0061] In some variations, the score-based model comprises a conditional score-based model trained with conditioning on the target context and an unconditional score-based model trained without conditioning on the target context.
[0062] In some variations, the score-based model comprises a linear combination of the conditional score-based model and the unconditional score-based model.
[0063] In some variations, a respective output of the conditional score-based model and the unconditional score-based model are weighted by a weight corresponding to a guidance level from the target context.
[0064] In some variations, the output protein sequence is generated by modifying the input protein sequence over multiple iterations of gradient-based Markov Chain Monte Carlo (MCMC) sampling.
[0065] In some variations, the output protein sequence is generated by at least applying the protein design computation model to generate a first modified protein sequence having a first modification to the input protein sequence, applying the protein design computation model to generate a second modified protein sequence having a second modification to the input protein sequence, determining that an antigen receptor corresponding to the first modified protein sequence is more likely to exhibit the target context than an antigen receptor corresponding to the second modified protein sequence, and applying the protein design computation model to further modify the first modified protein sequence instead of the second modified protein sequence.
[0066] In some variations, each of the first modification and the second modification include inserting, deleting, and / or changing a type of one or more amino acid residues in the input protein sequence.Attorney Ref.: 14786-074-228 (I03963-228074) / P40073-W0-1
[0067] In some variations, the output protein sequence is generated to correspond to the first modified protein sequence upon satisfying one or more criteria.
[0068] In some variations, the one or more criteria include at least one of (i) the input protein sequence having undergone a threshold quantity of gradient-based Markov Chain Monte Carlo (MCMC) sampling iterations to generate the first modified protein sequence, and (ii) the first modified protein sequence corresponding to at least the portion of the antigen receptor exhibiting the target context.
[0069] In some variations, the data distribution comprises a noisy data distribution populated by noisy protein sequences of antigen receptors from the antigen receptor repertoire.
[0070] In some variations, the generating of the output protein sequence includes denoising the first modified protein sequence.
[0071] In some variations, the protein design computation model comprises a composition of an energy function parameterized by a clonotype energy-based model (EBM) and an energy function parameterized by a maturity energy-based model (EBM)
[0072] In some variations, the clonotype energy-based model (EBM) is trained to approximate a data distribution of protein sequences of antigen receptors expressed by lymphocytes from a target clonotype comprising the target context.
[0073] In some variations, the maturity energy-based model (EBM) is trained to approximate a data distribution of protein sequences of antigen receptors expressed by mature lymphocytes further comprising the target context.
[0074] In some variations, the composition comprises a product of experts or a weighted sum.Attorney Ref.: 14786-074-228 (I03963-228074) / P40073-W0-1
[0075] In some variations, the output protein sequence is generated to correspond to an antigen receptor expressed by a mature lymphocyte from the target clonotype.
[0076] In some variations, the protein design computation model is trained to approximate the data distribution based on a plurality of positive samples.
[0077] In some variations, each positive sample includes a protein sequence of an antigen receptor paired with a correct context.
[0078] In some variations, the protein design computation model is trained to approximate the data distribution based on a plurality of negative samples.
[0079] In some variations, each negative sample includes a protein sequence of an antigen receptor paired with an incorrect context.
[0080] In some variations, the antigen receptor repertoire includes a plurality of B-cell receptors and / or T-cell receptors generated by one or more animals of a same or different species upon exposure to an antigen.
[0081] Implementations of the current subject matter can include, but are not limited to, methods consistent with the descriptions provided herein as well as articles that comprise a tangibly embodied machine-readable medium operable to cause one or more machines (e.g., computers, etc.) to result in operations implementing one or more of the described features. Similarly, computer systems are also described that may include one or more processors and one or more memories coupled to the one or more processors. A memory, which can include a non-transitory computer-readable or machine-readable storage medium, may include, encode, store, or the like one or more programs that cause one or more processors to perform one or more of the operations described herein. Computer implemented methods consistent with one or more implementations of the current subject matter can be implemented by one or more data processorsAttorney Ref.: 14786-074-228 (I03963-228074) / P40073-W0-1residing in a single computing system or multiple computing systems. Such multiple computing systems can be connected and can exchange data and / or commands or other instructions or the like via one or more connections, including, for example, to a connection over a network (e.g. the Internet, a wireless wide area network, a local area network, a wide area network, a wired network, or the like), via a direct connection between one or more of the multiple computing systems, etc.
[0082] The details of one or more variations of the subject matter described herein are set forth in the accompanying drawings and the description below. Other features and advantages of the subject matter described herein will be apparent from the description and drawings, and from the claims. While certain features of the currently disclosed subject matter are described for illustrative purposes in relation to the computational design of antigen receptors, such as B-cell receptors (BCRs), T-cell receptors (TCRs), and / or the like, it should be readily understood that such features are not intended to be limiting. The claims that follow this disclosure are intended to define the scope of the protected subject matter.DESCRIPTION OF DRAWINGS
[0083] The accompanying drawings, which are incorporated in and constitute a part of this specification, show certain aspects of the subject matter disclosed herein and, together with the description, help explain some of the principles associated with the disclosed implementations. In the drawings,
[0084] FIG. 1 depicts a system diagram illustrating an example of an antigen receptor design system, in accordance with some example embodiments;
[0085] FIG. 2A depicts a flowchart illustrating an example of a process for guided generation of antigen receptors, in accordance with some example embodiments;Attorney Ref.: 14786-074-228 (I03963-228074) / P40073-W0-1
[0086] FIG. 2B depicts a flowchart illustrating another example of a process for guided generation of antigen receptors, in accordance with some example embodiments;
[0087] FIG. 2C depicts a flowchart illustrating another example of a process for guided generation of antigen receptors, in accordance with some example embodiments;
[0088] FIG. 3 depicts a flowchart illustrating an example of a process for guided generation of antigen receptors with gradient-based Markov Chain Monte Carlo (MCMC) sampling, in accordance with some example embodiments; and
[0089] FIG. 4 (a) depicts a schematic diagram illustrating an example of one or more subclasses of lead protein molecules selected from the B-cell receptor (BCR) repertoire of an antigen-exposed animal’s immune system for use as context for conditioning the generation of novel protein sequences, in accordance with some example embodiments;
[0090] FIG. 4(b) depicts a schematic diagram illustrating an example of a variable (V) gene segment, a joint (J) gene segment, and a third complementarity determining region (CDR3) length defining a clonotype subclass for use as context for conditioning the generation of novel protein sequences, in accordance with some example embodiments;
[0091] FIG. 4(c) depicts a schematic diagram illustrating an energy based model (EBM) and a score-based denoiser implementing an example of a protein design computation model for context conditioned generation of novel protein sequences, in accordance with some example embodiments;
[0092] FIG. 5(a) depicts a graph illustrating the distribution of the embeddings of sample protein sequences in a training dataset and the distribution of the embeddings of novel protein sequences generated with context conditioning, in accordance with some example embodiments;Attorney Ref.: 14786-074-228 (I03963-228074) / P40073-W0-1
[0093] FIG. 5(b) depicts a graph illustrating the embeddings of novel protein sequences generated with context conditioning and differentiated based on target clonotypes, in accordance with some example embodiments;
[0094] FIG. 5(c) depicts a graph illustrating the embeddings of novel protein sequences generated with context conditioning and differentiated based on clonotypes, in accordance with some example embodiments;
[0095] FIG. 6(a) depicts a graph illustrating the correlation between the guidance level of context conditioning and accuracy of novel protein sequences generated with conditioning on individual and joint contextual variables, in accordance with some example embodiments;
[0096] FIG. 6(b) depicts a graph illustrating the correlation between the guidance level of context conditioning and the diversity of novel protein sequences generated with context conditioning, in accordance with some example embodiments;
[0097] FIG. 6(c) depicts a graph illustrating the correlation between the guidance level of context conditioning and the novelty of novel protein sequences generated with context conditioning, in accordance with some example embodiments;
[0098] FIG. 6(d) depicts a graph illustrating the correlation between the guidance level of context conditioning and the position-wise Kullback-Leibler (KL) divergence of novel protein sequences generated with context conditioning, in accordance with some example embodiments;
[0099] FIG. 7(a) depicts a graph illustrating the correlation between the guidance level of context conditioning and the accuracy of novel protein sequences generated with conditioning on variable (V) gene segment, in accordance with some example embodiments;Attorney Ref.: 14786-074-228 (I03963-228074) / P40073-W0-1
[0100] FIG. 7(b) depicts a graph illustrating the correlation between the guidance level of context conditioning and the accuracy of novel protein sequences generated with conditioning on variable (V) gene segment family, in accordance with some example embodiments;
[0101] FIG. 7(c) depicts a graph illustrating the correlation between the guidance level of context conditioning and the accuracy of novel protein sequences generated with conditioning on joint (J) gene segment, in accordance with some example embodiments;
[0102] FIG. 7(d) depicts a graph illustrating the correlation between the guidance level of context conditioning and the accuracy of novel protein sequences generated with conditioning on third complementarity determining region (CDR3) length, in accordance with some example embodiments;
[0103] FIG. 7(e) depicts a graph illustrating the correlation between the guidance level of context conditioning and the accuracy of novel protein sequences generated with conditioning on variable (V) gene segment, joint (J) gene segment, and third complementarity determining region (CDR3) length, in accordance with some example embodiments;
[0104] FIG. 7(f) depicts a graph illustrating the correlation between the guidance level of context conditioning and the uniqueness of novel protein sequences generated with context conditioning, in accordance with some example embodiments;
[0105] FIG. 7(g) depicts a graph illustrating the correlation between the guidance level of context conditioning and the position-wise Kullback-Leibler (KL) divergence of novel protein sequences generated with context conditioning, in accordance with some example embodiments;Attorney Ref.: 14786-074-228 (I03963-228074) / P40073-W0-1
[0106] FIG. 7(g) depicts a graph illustrating the correlation between the guidance level of context conditioning and the diversity of novel protein sequences generated with context conditioning, in accordance with some example embodiments;
[0107] FIG. 8 depicts Langevin Markov Chain Monte Carlo (MCMC) sampling trajectories illustrating novel protein sequences generated with context conditioning converging to a target variable (V) gene class, in accordance with some example embodiments;
[0108] FIG. 9(a) depicts a sequence logo plot illustrating the most common types of amino acid residue at each position of sample protein sequences in a training dataset, in accordance with some example embodiments;
[0109] FIG. 9(b) depicts a sequence logo plot illustrating the most common types of amino acid residue at each position of sample protein sequences from a subset of sampled protein sequences labeled as having a IGHV2-63 variable (V) gene segment, in accordance with some example embodiments;
[0110] FIG. 9(c) depicts a sequence logo plot illustrating the most common types of amino acid residue at each position of novel protein sequences generated with context conditioning on the IGHV2-63 variable (V) gene segment, in accordance with some example embodiments;
[0111] FIG. 10(a) depicts a graph illustrating the correlation between the guidance level of context conditioning on the quantity of somatic hypermutations and the accuracy of novel protein sequences generated with context conditioning as quantified by the quantity of somatic hypermutations relative to a target, in accordance with some example embodiments;
[0112] FIG. 10(b) depicts a graph illustrating the correlation between the guidance level of context conditioning on the quantity of somatic hypermutations and the accuracy of novelAttorney Ref.: 14786-074-228 (I03963-228074) / P40073-W0-1protein sequences generated with context conditioning as quantified by root-mean squared error (RMSE) relative to a target, in accordance with some example embodiments;
[0113] FIG. 11A depicts Langevin Markov Chain Monte Carlo (MCMC) sampling trajectories illustrating novel protein sequences generated with context conditioning IGHV6-5 variable (V) gene segment, in accordance with some example embodiments;
[0114] FIG. 11B depicts Langevin Markov Chain Monte Carlo (MCMC) sampling trajectories illustrating novel protein sequences generated with context conditioning IGHV2-63 variable (V) gene segment, in accordance with some example embodiments;
[0115] FIG. 11C depicts Langevin Markov Chain Monte Carlo (MCMC) sampling trajectories illustrating novel protein sequences generated with context conditioning IGHV5-29 variable (V) gene segment, in accordance with some example embodiments;
[0116] FIG. 12 depicts a block diagram illustrating an example of a computing system, in accordance with some example embodiments.
[0117] When practical, similar reference numbers denote similar structures, features, or elements.DETAILED DESCRIPTION
[0118] Designing a protein sequence exhibiting one or more properties of interest is a critical task in biomedicine and bioengineering. In the context of drug discovery, the protein sequence may be a protein therapeutic for treating, preventing, or curing diseases and medical conditions. Examples of protein therapeutics include antibodies, peptide hormones, growth factors, plasma proteins, enzymes, or hemolytic factors. The one or more properties of interest in the context of protein therapeutics may include binding affinity, binding specificity, functionality,Attorney Ref.: 14786-074-228 (I03963-228074) / P40073-W0-1and developability traits such as homogeneity, stability, solubility, and viscosity. The presence (and absence) of the one or more properties of interest may determine the fitness of the protein sequence as a viable protein therapeutic. As such, the process of designing a protein sequence exhibiting the one or more properties of interest typically includes two high-level phases: lead discovery (LD) followed by lead optimization (LO). The objective of lead discovery (LD) is to identify novel protein sequences exhibiting at least some level of fitness as a protein therapeutic. Those novel protein sequences then undergo lead optimization (LO), which aims to increase (or maximize) the fitness of the novel protein sequences identified through lead discovery (LD). The success of therapeutic protein design, particularly that of lead optimization (LO), tends to be heavily contingent on the quality of lead molecules provided by lead discovery (LD). However, conventional lead discovery (LD) techniques fail to consistently deliver high quality lead molecules capable of being developing into viable therapeutics.
[0119] To accelerate drug development and reduce reliance on expensive wet lab resources, drug discovery efforts, including lead discovery (LD) and lead optimization (LO), may leverage a variety of computational tools. However, conventional computation models, including generative models trained to generate novel protein molecules (e.g., lead molecules) for further optimization or experimental validation, may be prone to generating low quality protein molecules with poor chances of being developed into viable therapeutics. For example, a conventional computation model used for generating protein molecules (e.g., lead molecules) may be trained on a training dataset that includes sample protein sequences from the sequence neighborhoods of any known lead molecule and other related sequences. The resulting computation model may approximate a data distribution populated by protein sequences whose properties correspond to those found in the sample protein sequences. Although novel protein sequences may be generatedAttorney Ref.: 14786-074-228 (I03963-228074) / P40073-W0-1by sampling from this data distribution, unconditional sampling will often yield protein sequences with undesired properties at least because of the inability to avoid sampling from neighborhoods densely populated by protein sequences without any properties of interest . For instance, some protein sequences may be sampled from neighborhoods in the data distribution that originate from uninteresting modes in the training dataset, which in this context refer to sample protein sequences that appear frequently in the training dataset but are not necessarily suitable for further development. Some protein sequences sampled from the data distribution may be chimeras, which exhibit an undesired combination of properties that arise when different modes are incorrectly combined during the training of the computation model. Constrained sampling, which may include keeping some portions of the sampled protein sequences fixed, affords some control over the protein sequences generated by the computation model but enough to steer generation towards neighborhoods in the data distribution populated by high-fitness protein sequences.
[0120] Various embodiments of the present disclosure overcome the limitations of conventional computation models for generating protein sequences, such as those serving as lead molecules for lead optimization (LO), by conditioning the generation of protein sequences on context. For example, in some cases, the generation of protein sequences, such as antigen receptors (e.g., B-cell receptors (BCRs), T-cell receptors (TCRs), and / or the like) may be conditioned on one or more contextual variables defining a context. In some cases, the context may include protein subclasses such as those corresponding to clonal families, clonotypes, maturity, and / or the like. In some cases, protein sequences from some protein subclasses may be more likely to exhibit one or more properties of interest than those in other protein subclasses. Accordingly, unlike the unconditional sampling or constrained sampling associated with conventional computation models, various example embodiments of the protein design computation model described hereAttorney Ref.: 14786-074-228 (I03963-228074) / P40073-W0-1may be steered, through conditioning on one or more contextual variables, to generate protein sequences from certain protein subclasses, as defined by the one or more contextual variables. In some cases, conditioning on the one or more contextual variables may steer the protein design computation model to generate (or sample) protein sequences from neighborhoods in the data distribution populated by certain protein subclasses. For instance, in some cases, those protein subclasses may include protein sequences exhibiting one or more properties of interest and are therefore high-fitness protein sequences with a better chance of being developed into viable therapeutics. As such, in some cases, various embodiments of the protein design computation model described herein may be applied to increase the efficiency of lead discovery (LD) and improve the quality of lead molecules made available for subsequent lead optimization (LO).
[0121] An immune repertoire is the collection of antigen receptors, such as T-cell receptors (TCRs) and B-cell receptors (BCRs), generated by an immune system. Immune repertoires from animals that have been exposed to an antigen may be divided into clonotypes (or clonal families), each of which being a subclass of protein sequences derived from the same progenitor lymphocyte (e.g., B-cell, T-cell, and / or the like). In this context, the term “clonotype” or “clonal family” may refer to a group of lymphocytes (e.g., B cell, T cell, and / or the like) that evolved from the same ancestral lymphocyte (e.g., progenitor B cell, progenitor T cell, and / or the like). The deoxyribonucleic acid (DNA) of an ancestral lymphocyte may be diversified through mutations to form a lineage of related but not necessarily identical lymphocytes. For example, the deoxyribonucleic acid (DNA) of a lymphocyte includes the coding sequence for the antigen receptors (e.g., B-cell receptors (BCR) or T-cell receptors (TCRs)) expressed by the lymphocyte. The coding sequence of a lymphocyte includes a variable (V) gene segment, a diversity (D) segment, and a joining (J) gene segment, which code for the antigen-binding fragments (Fab) ofAttorney Ref.: 14786-074-228 (I03963-228074) / P40073-W0-1the antigen receptors expressed by the lymphocyte. The antigen-binding fragments (Fab) includes the variable region (Fv), which is responsible for recognizing specific antigens from pathogens (e.g., bacteria, viruses, parasites, and / or the like) or abnormal cells (e.g., cancers). Lymphocytes undergo affinity maturation to increase (or maximize) properties of interest (e.g., affinity, avidity, anti-pathogen activity, and / or the like) of the antigen receptors expressed by the lymphocytes through somatic hypermutations (SHM) of the antigen receptor coding sequences (e.g., V(D)J gene segments) in the lymphocytes. Antigen receptors expressed by “mature” lymphocyte, which has undergone adequate affinity maturation (e.g., a threshold quantity of cycles of affinity maturation), may therefore be more likely to exhibit desired properties than ones expressed by lymphocytes that has not undergone adequate affinity maturation. Variable-diversity-joining rearrangement (or V(D)J recombination) is one mechanism for achieving somatic hypermutations (SHM) in which the variable (V), joining (J), and in some cases, diversity (D) gene segments of a lymphocyte are rearranged in a nearly random fashion to generate a diverse pool of clonotypes, each of which being associated with a unique antigen receptor coding sequence (e.g., V(D)J gene segments). A single clonal family may include multiple clonotypes, each of which being a subset of lymphocytes with the same antigen receptor coding sequence.
[0122] In some example embodiments, the protein design computation model may be applied to generate protein sequences corresponding to antigen receptors, such as B-cell receptors (BCRs), T-cell receptors (TCRs), and / or the like. For example, in instances where the protein sequences are antigen receptors, the protein design computation model may generate protein sequences while conditioned on one or more contextual variables. In some cases, the one or more contextual variables may define context, one example of which is one or more target clonotypes (or clonal families). In some cases, the contextual variables defining the one or more targetAttorney Ref.: 14786-074-228 (I03963-228074) / P40073-W0-1clonotypes (or clonal families) may include the V-gene segment, the J-gene segment, and third complementarity determining region (CDR3) cluster of a protein sequencesuch that the conditioning on contextual variables steers the generation of protein sequences towards one or more target clonotypes or clonal families. In some cases, one or more protein sequences may be generated while conditioned on one or more contextual variables in order to perform hit expansion in which variants of a lead molecule exhibiting one or more properties of interest are computationally generated to exhibit similar properties of interest. As described in more detail below, various example embodiments of protein sequence generation described herein may be conditioned to steer generation towards those protein sequences exhibiting one or more properties of interest without guidance from one or more external property computation models. In some cases, this predictor-free guidance, which obviates reliance on gradients from external property computation models, may result in faster, more computationally efficient computational protein design while maintaining controllable mode exploration of the vast combinatorial space of possible protein sequences (e.g., approximately 20Npossible protein sequences for protein sequences containing N quantity of amino acid residues selected from twenty canonical amino acid residues).
[0123] In some example embodiments, one or more target clonotypes or clonal families may be identified based on data associated with one or more antigen receptor repertoires. For example, in some cases, the protein design computation model may be steered to generate protein sequences of antigen receptors from clonotypes (or clonal families) that are enriched in the antigen receptor repertoires of one or more animals and / or species of animals. In some cases, upon exposure to an antigen, such as a virus or a cell surface receptor, an animal’s immune system may generate antigen receptors capable of binding to the antigen. The term “antigen receptor repertoire” may refer to the range of antigen receptors expressed by the total lymphocyte (e.g., BAttorney Ref.: 14786-074-228 (I03963-228074) / P40073-W0-1cell or T cell) population of one or more animals or species of animals. It should be appreciate that the term “antigen receptor” as used herein may refer to a variety of different types of antigen receptor (or portions of antigen receptors) including, for example, B-cell receptors (BCRs), T-cell receptors (TCRs), full-length antibodies (e.g., immunoglobulin G (IgG), immunoglobulin M (IgM)), single-chain variable fragments (scFvs), antigen-binding fragments (Fabs), single-domain antibodies (e.g., variable heavy domain of heavy chain (VHH) or nanobodies), synthetic or semisynthetic binding proteins derived from or modeled after immune repertoires, and / or the like. A clonotype (or clonal family) that is enriched in one or more antigen receptor repertoires may be overrepresented (or present at a significantly higher than expected frequency) within the antigen receptor repertoires. This overrepresentation may indicate that the antigen receptors expressed by the lymphocytes that are members of the clonotype or clonal family are more likely to be binders for the antigen. Accordingly, in instances where the protein design computation model is applied to generate antigen receptors capable of binding to the antigen, the generation of protein sequences may be steered towards the clonotypes or clonal families that are enriched in antigen receptor repertoires at least because protein sequences sampled from the corresponding neighborhoods may be more likely to exhibit binding affinity towards the same antigen. Further optimization of these protein sequences, for example, through lead optimization (LO) is more likely to yield viable therapeutics. To date, conventional approaches to protein design, including computational methodologies leveraging artificial intelligence and machine learning, have failed to incorporate valuable data from antigen receptor repertoires, including for example, BCR repertoire data. Without insights gleaned from antigen receptor repertoires, which includes the clonotypes (or clonal families) of antigen receptors more likely to be binders of an antigen, conventional proteinAttorney Ref.: 14786-074-228 (I03963-228074) / P40073-W0-1design solutions do not support a controllable generative framework capable of yielding diverse protein sequences with the properties of interest associated with viable protein therapeutics.
[0124] In some example embodiments, the protein design computation model may generate protein sequences while conditioned on one or more contextual variables that define context, such as maturity. In some cases, conditioning on the one or more contextual variables defining maturity may steer the generation of protein sequences towards those produced by mature lymphocytes that have undergone adequate affinity maturation, such as a threshold quantity of cycles of affinity maturation. In this context, the one or more contextual variables defining the “maturity” of a lymphocyte may include the quantity of somatic hypermutations (SHMs) present in its antigen receptor coding sequence relative to its ancestral lymphocyte, the depth of the lymphocyte in a lymphocyte phylogenetic tree depicting the evolutionary relationships within a corresponding antigen receptor repertoire, and / or the like. As noted, a mature lymphocyte may be more likely to exhibit properties of interest, such as affinity, avidity, and anti-pathogen activity, at least because the lymphocyte have undergone adequate affinity maturation to increase (or maximize) these properties of interest. In some cases, the generation of protein sequences may be conditioned on contextual variables defining a certain clonotype (or clonal family) as well as maturity such that the resulting protein sequences belong to the clonotype (or clonal family) and are likely to have undergone adequate affinity maturation. As described in more detail below, the protein design computation model may include a composition of multiple energy-based models (EBMs), each of which being trained to approximate the data distribution of protein sequences with a different context. For example, in some cases, the composition of energy -based models (EBMs) may include one energy-based model (EBM) trained to approximate the data distributionAttorney Ref.: 14786-074-228 (I03963-228074) / P40073-W0-1of clonotypes (or clonal families) and a different energy-based model (EBM) trained to approximate the data distribution of maturity.
[0125] In some example embodiments, the protein design computation model may include an energy -based model (EBM) trained to approximate the data distribution of protein sequences from one or more target clonotypes (or clonal families), such as one or more clonotypes (or clonal families) that enriched (or overrepresented) in one or more antigen receptor repertoires. In some cases, the data distribution of target clonotypes (or clonal families) may include higher density regions populated by proteins sequences of antigen receptors belonging to the target clonotypes (or clonal families) and lower density regions populated by protein sequences of antigen receptors not belonging to the target clonotypes (or clonal families). In some cases, the parameters of the energy-based model (EBM) may be adjusted during the training of the energybased model (EBM) to approximate the aforementioned data distribution. Furthermore, the parameters of the energy-based model may parameterize a function (e.g., an energy function, score function, and / or the like) whose output or changes in whose output correspond to changes in the density of the data distribution. For example, in some cases, the function may assign a higher value (e.g., energy value, score, and / or the like) to the protein sequences of antigen receptors (e.g., B-cell receptors (BCRs) or T-cell receptors (TCRs)) belonging to the target clonotypes (or clonal families) and a lower value (e.g., energy value, score, and / or the like) to the protein sequences of those antigen receptors that do not. As described in more detail below, in some cases, the protein design computation model may be applied to generate protein sequences of antigen receptors while guided by the output of the function (e.g., gradient of the energy function, score of the score function, and / or the like) such that the protein design computation model samples, from the higher density regions of the data distribution, protein sequences of antigen receptors belonging to theAttorney Ref.: 14786-074-228 (I03963-228074) / P40073-W0-1target clonotypes (or clonal families) while avoiding the protein sequences of antigen receptors not belonging to the target clonotypes (or clonal families).
[0126] In some example embodiments, the protein design computation model may include an energy -based model (EBM) trained to approximate the data distribution of protein sequences of antigen receptors expressed by mature lymphocytes, which are lymphocytes that are more likely to have undergone adequate affinity maturation (e.g., a threshold quantity of cycles of affinity maturation). In some cases, the maturity of a lymphocyte may be quantified by the number of somatic hypermutations (SHM) present in its antigen receptor coding sequence relative to its ancestral lymphocyte. In some cases, the number of somatic hypermutations (SHM) may correspond to the edit distance between the antigen receptor coding sequence of the lymphocyte and that of the ancestral lymphocyte. Alternatively and / or additionally, the maturity of the lymphocyte may be quantified by the depth of the lymphocyte in a lymphocyte phylogenetic tree depicting the evolutionary relationships within a corresponding antigen receptor repertoire. In some cases, the data distribution may include higher density regions populated by the protein sequences of antigen receptors expressed by more mature lymphocytes, such as those lymphocytes exhibiting threshold number of somatic hypermutations (or edit distance) and / or occupying a threshold depth in the lymphocyte phylogenetic tree. In some cases, the data distribution may also include lower density regions populated by the protein sequences of antigen receptors expressed by less mature lymphocytes, such as those lymphocytes that fail to exhibit threshold number of somatic hypermutations (or edit distance) and / or occupy a threshold depth in the lymphocyte phylogenetic tree. In some cases, the parameters of the energy-based model (EBM) may be adjusted during the training of the energy-based model (EBM) to approximate the data distribution of the protein sequences of antigen receptors expressed by mature lymphocytes. Furthermore, inAttorney Ref.: 14786-074-228 (I03963-228074) / P40073-W0-1some cases, the parameters of the energy-based model (EBM) may parameterize a function (e.g., energy function, score function, and / or the like) whose output or changes in whose output correspond to changes in the density of the data distribution. For example, in some cases, this function may assign a higher value (e.g., energy value, score, and / or the like) to protein sequences of antigen receptors expressed by mature lymphocytes and a lower value (e.g., energy value, score, and / or the like) to the protein sequences of antigen receptors may immature lymphocytes. As described in more detail below, in some cases, the protein design computation model may be applied to generate protein sequences of antigen receptors while guided by the output of the function (e.g., gradient of the energy function, score of the score function, and / or the like) such that the protein design computation model samples, from the higher density regions of the data distribution, protein sequences of antigen receptors expressed by mature lymphocytes while avoiding those protein sequences of antigen receptors expressed by immature lymphocytes.
[0127] In some example embodiments, the protein design computation model may include a composition of the energy-based model (EBM) approximating the data distribution of the protein sequences of antigen receptors from target clonotypes (or clonal families) and the energy -based model (EBM) approximating the data distribution of the protein sequences of antigen receptors from mature lymphocytes (or lymphocytes more likely to have undergone adequate affinity maturation). In some cases, the generation of protein sequences may be guided by a total function (e.g., total energy function, total score function, and / or the like) summing the respective functions in order to steer the protein design computation model to sample, from the higher density regions of both data distributions, protein sequences of antigen receptors expressed by mature lymphocytes belonging to the target clonotypes (or clonal families). In some cases, the total function may be a product of experts or a weighted sum of the individual functions such that eachAttorney Ref.: 14786-074-228 (I03963-228074) / P40073-W0-1constituent function contributes a different (or equal) amount in terms of how much the generating of the protein sequences is influenced by the output of each function. As such, in some cases, it may be possible for the sampling to prioritize more (or less) the protein sequences generated by the protein design computation model being the protein sequences of antigen receptors from target clonotypes (or clonal families) over the protein sequences of antigen receptors expressed by mature lymphocytes.
[0128] In some example embodiments, the protein design computation model may modify an input protein sequence in order to generate an output protein sequence corresponding, for example, to an antigen receptor such as a B-cell receptor (BCR) or a T-cell receptor (TCR). In some cases, the modifying of the input protein sequence may be guided by the gradient of one or more functions (e.g., energy functions, score functions, and / or the like), such as the function of the energy -based model (EBM) approximating the data distribution of the protein sequences of antigen receptors from target clonotypes (or clonal families), the function of the energy-based model (EBM) approximating the data distribution of the protein sequences of antigen receptors expressed by mature lymphocytes, or a composition thereof. For example, in some cases, the protein design computation model may be applied to modify the input protein sequence over multiple iterations, with each iteration of modifications tantamount to drawing one or more samples (or modified protein sequences) from the data distributions. In some cases, the input protein sequence may be modified through gradient based Markov Chain Monte Carlo (MCMC) sampling (e.g., Langevin Markov Chain Monte Carlo (MCMC)), with each successive iteration further modifying a modified protein sequence sampled from a higher density region of the data distributions during a previous iteration. In some cases, guidance from the output of the one or more functions (e g., energy functions, score functions, and / or the like) may ensure successive samples (or modifiedAttorney Ref.: 14786-074-228 (I03963-228074) / P40073-W0-1protein sequences) are drawn from incrementally higher density regions of the data distributions, which are populated by protein sequences of antigen receptors from target clonotypes (or clonal families) and / or mature lymphocytes. In some cases, the one or more criteria may include the output protein sequence being drawn from a sufficiently high density region of the data distributions and is therefore assigned a threshold value (e.g., energy value, score, and / or the like) by the function. Alternatively and / or additionally, the oner or more criteria may include the protein design computation model having been applied to perform a threshold quantity of iterations of gradient-based Markov Chain Monte Carlo (MCMC) sampling (e.g., Langevin Markov Chain Monte Carlo (MCMC) sampling and / or the like).
[0129] In some example embodiments, the protein design computation model, including the clonotype energy-based model and the maturity energy-based model, may be trained to approximate the noisy data distributions of the protein sequences of antigen receptors expressed by mature lymphocytes from the target clonotype (or clonal family). For example, in some cases, instead of the protein sequences of antigen receptors from the antigen receptor repertoire of a single animal or multiple animals from the same or different species, the training dataset for training the protein design computation model may include noisy protein sequences. In some cases, a noisy protein sequence may be generated by adding noise, such as Gaussian noise, to the protein sequence of an antigen receptor from the antigen receptor repertoire of a single animal or multiple animals from the same or different species. Once trained, the protein design computation model may sample noisy protein sequences from the noisy data distributions by modifying an input protein sequence, for example, over multiple iterations of gradient-based Markov Chain Monte Carlo (MCMC) sampling (e.g., Langevin Markov Chain Monte Carlo (MCMC)). Sampling from the noisy data distribution may be more efficient at least because the noisy data distribution mayAttorney Ref.: 14786-074-228 (I03963-228074) / P40073-W0-1exhibit smoother transitions between higher and lower density regions. Contrastingly, when sampling from the true data distribution of clean (or noiseless) protein sequences of antigen receptors from the antigen receptor repertoire, the protein design computation model may be confined to areas within the immediate vicinity of the protein sequences of the antigen receptors in the antigen receptor repertoire due to the steep changes in densities across the true data distribution. The steep changes in densities give rise to the phenomenon of mode collapse where the protein design computation model is confined to sampling output protein sequences from regions within the immediate vicinity of the protein sequences of the antigen receptors in the antigen receptor repertoire. When trained to approximate the true data distribution of the clean (or noiseless) protein sequences of the antigen receptors from the antigen receptor repertoire, the protein design computation model may be less robust and capable of generating a smaller variety of output protein sequences than when the protein design computation model is trained to approximate the noisy data distribution.
[0130] FIG. 1 depicts a system diagram illustrating an example of a protein design system 100, in accordance with some example embodiments. Referring to FIG. 1, the protein design system 100 may include a protein design engine 110, a client device 120, and a data store 150. In the example shown in FIG. 1, the protein design engine 110, the client device 120, and the data store 130 may be communicatively coupled via a network 140. In some cases, the client device 130 may be a processor-based device including, for example, a workstation, a desktop computer, a laptop computer, a smartphone, a tablet computer, a wearable apparatus, and / or the like. In some cases, the network 140 may be a wired network and / or a wireless network including, for example, a local area network (LAN), a virtual local area network (VLAN), a wide area network (WAN), a public land mobile network (PLMN), the Internet, and / or the like. In someAttorney Ref.: 14786-074-228 (I03963-228074) / P40073-W0-1cases, the data store 120 may be a database including, for example, a relational database, a NoSQL database, a columnar database, an objected-oriented database, a key -value database, a hierarchical database, a document database, a graph database, and / or the like.
[0131] In some example embodiments, the protein design engine 110 may include a protein design computation model 115 trained to generate, based at least on an input protein sequence 142, an output protein sequence 144. In some cases, the input protein sequence 142 may be the protein sequence of at least a portion of antigen receptor, such as a B-cell receptor (BCR) or a T-cell receptor (TCR). In some cases, the antigen receptor corresponding to the input protein sequence 142 may exhibit one or more properties of interest, undesired properties, or no known properties at all. Alternatively, the input protein sequence 142 may be a noise sequence, which is randomly ordered sequence of amino acid residues without any known properties. In some cases, the protein design computation model 115 may be applied to generate the output protein sequence 144 by at least modifying the input protein sequence 142. For example, in some cases, the protein design computation model 115 may modify the input protein sequence 142 by inserting, deleting, and / or changing the type of one or more amino acid residues in the input protein sequence 142. In some cases, the protein design computation modelln 1 115 may generate the output protein sequence 142 by modifying the input protein sequence 142 over multiple iterations of gradientbased Markov Chain Monte Carlo (MCMC) sampling (e.g., Langevin Markov Chain Monte Carlo (MCMC)).
[0132] In some example embodiments, the protein design computation model 115 may be trained to generate the output protein sequence 144 to correspond to at least a portion of an antigen receptor (e.g., B-cell receptor (BCR) or T-cell receptor (TCR)) from a target clonotype (or clonal family), such as a clonotype (or clonal family) that is enriched (or overrepresented) in anAttorney Ref.: 14786-074-228 (I03963-228074) / P40073-W0-1antigen receptor repertoire 165 (e.g., a B-cell receptor (BCR) repertoire, a T-cell receptor (TCR) repertoire, and / or the like). Alternatively and / or additionally, the protein design computation model 115 may be trained to generate the output protein sequence 144 to correspond to at least a portion of an antigen receptor expressed by a mature lymphocyte (e.g., from the antigen receptor repertoire 165) that have undergone adequate affinity maturation (e.g., a threshold quantity of cycles of affinity maturation). It should be appreciated the antigen receptor repertoire 165 may include the antigen receptors expressed by the total lymphocyte (e.g., B cell or T cell) population of one or more animals or species of animals upon exposure to an antigen. Antigen receptors from a clonotype (or clonal family) that is enriched (or overrepresented) in the antigen receptor repertoire 165 as well as those expressed by mature lymphocytes may be more likely exhibit properties of interest, such as affinity, avidity, anti -pathogen activity, and / or the like. Accordingly, conditioning the generation of the output protein sequence 144 on clonotype (or clonal family) and / or lymphocyte maturity may increase the likelihood of the antigen receptor corresponding to the output protein sequence 144 generated by the protein design computation model 115 having one or more properties of interest (e.g., binding affinity, avidity, anti-pathogen activity, and / or the like).
[0133] Referring again to FIG. 1, in some example embodiments, the protein design computation model 115 may include a clonotype energy -based model (EBM) 150 trained to approximate the data distribution of the protein sequences of antigen receptors from the target clonotype (or clonal family) associated with one or more properties of interest, such as affinity, avidity, anti-pathogen activity, and / or the like. In some cases, the clonotype energy-based model 150 may be trained based on the antigen receptor repertoire 165. For example, in some cases, the antigen receptor repertoire 165 may include the total lymphocyte (e.g., B cell or T cell) populationAttorney Ref.: 14786-074-228 (I03963-228074) / P40073-W0-1of a single animal or multiple animals of the same or different species upon exposure to an antigen (e.g., virus or cell surface receptor). In some cases, the antigen receptor repertoire 165 may include lymphocytes from multiple clonotypes (or clonal families). A clonotype (or clonal family) that is enriched (or overrepresented) in the antigen receptor repertoire 165 may indicate that the antigen receptors expressed by the lymphocytes in the clonotype (or clonal family) are more likely to be binders for the antigen. As such, in some cases, a clonotype (or clonal family) that is enriched (or overrepresented) in the antigen receptor repertoire 156 may be identified as a target clonotype (or clonal family). In some cases, the training of the clonotype energy -based model 150 may include adjusting one or more parameters (e.g., weights, biases, and / or the like) of the clonotype energybased model 150 such that a function 151 (e.g., energy function, score function, and / or the like) parameterized by these parameters assigns a higher value (e.g., energy value, score, and / or the like) to a protein sequence corresponding to an antigen receptor from the target clonotype (or clonal family) and a lower value (e.g., energy value, score, and / or the like) to a protein sequence corresponding to an antigen receptor outside of the target clonotype (or clonal family).
[0134] As noted, in some cases, the output of the function 151 may indicate changes in the density of the data distribution of the protein sequences of antigen receptors expressed by lymphocytes from the target clonotype. For example, in some cases, the function 151 may be an energy function whose output indicates the density at a location from which the protein sequence of an antigen receptor is sampled. Alternatively and / or additionally, the function 151 may be a score function whose output indicates the local change in density at the location from which the protein sequence of an antigen receptor is sampled. In some cases, the function 151 may assign a higher value (e.g., energy value, score, and / or the like) to the protein sequences of antigen receptors (e.g., B-cell receptors (BCRs) or T-cell receptors (TCRs)) belonging to the target clonotype (orAttorney Ref.: 14786-074-228 (I03963-228074) / P40073-W0-1clonal family). Those protein sequences may be sampled from higher density regions of the data distribution. Contrastingly, a lower value (e.g., energy value, score, and / or the like) may be assigned to the protein sequences sampled from lower density regions of the data distribution, which are populated by the protein sequences of antigen receptors that do not belong to the target clonotype (or clonal family). Accordingly, in some cases, the modifying of the input protein sequence 142 may be guided by the output of the function 151 (e.g., gradient of the energy function, score of the score function, and / or the like) such that the output protein sequence 144 is sampled from a higher density region of the data distribution than the input protein sequence 142, meaning that the output protein sequence 144 is more likely to correspond to an antigen receptor from the target clonotype (or clonal family).
[0135] In some example embodiments, the protein design computation model 115 may include a maturity energy -based model (EBM) 152 trained to approximate the data distribution of the protein sequences of antigen receptors expressed by mature lymphocytes that have undergone adequate affinity maturation (e.g., a threshold quantity of cycles of affinity maturation). As noted, more mature lymphocytes may be more likely to generate antigen receptors exhibiting properties of interest (e.g., affinity, avidity, anti-pathogen activity, and / or the like) at least because the process of affinity maturation increases (or maximizes) these properties. In FIG. 1, the maturity energybased model 152 may parameterize a function 153 (e.g., energy function, score function, and / or the like) whose output corresponds to the changes in the density of the data distribution. In some cases, higher density regions of the data distribution may be populated by the protein sequences of antigen receptors expressed by mature lymphocytes while lower density regions of the data distribution may be populated by the protein sequences of antigen receptors expressed by immature lymphocytes. Moreover, in some cases, the function 153 may assign a higher value (e.g., energyAttorney Ref.: 14786-074-228 (I03963-228074) / P40073-W0-1value, score, and / or the like) to the protein sequences sampled from higher density regions of the data distribution, which are those of antigen receptors expressed by mature lymphocytes, whereas protein sequences sampled from lower density regions of the data distribution are assigned a lower value (e.g., energy value, score, and / or the like).
[0136] In some example embodiments, the maturity energy -based model 152 may be trained based on the antigen receptor repertoire 165, which includes antigen receptors expressed by lymphocytes from multiple clonotypes (or clonal families). In some cases, the “maturity” of an antigen receptor in the antigen receptor repertoire 165 may be quantified by the quantity of somatic hypermutations (SHM) present in the antigen receptor coding sequence of the lymphocyte expressing the antigen receptor and that of the ancestral lymphocyte, which corresponds to the edit distance between the two antigen receptor coding sequences. Alternatively and / or additionally, the maturity of the antigen receptor may be quantified by the depth of the lymphocyte expressing the antigen receptor in a lymphocyte phylogenetic tree depicting the evolutionary relationships within the antigen receptor repertoire 165. In some cases, the training of the maturity energybased model 152 may include adjusting one or more parameters (e.g., weights, biases, and / or the like) of the maturity energy -based model 152 such that the function 153 parameterized by these parameters assigns, to the protein sequence of an antigen receptor, a value (e.g., energy value, score, and / or the like) corresponding to the maturity of the antigen receptor as quantified, for example, by the edit distance, quantity of somatic hypermutations, and / or depth in the lymphocyte phylogenetic tree as described above. Alternatively, the parameters of the maturity energy-based model 152 may be adjusted during training such that the function 153 assigns a higher value (e.g., energy value, score, and / or the like) to the protein sequence of a antigen receptor expressed by a mature lymphocyte whose edit distance, number of somatic hypermutations, and / or depth in theAttorney Ref.: 14786-074-228 (I03963-228074) / P40073-W0-1lymphocyte phylogenetic tree satisfy one or more criteria. In such instances, the parameters of the maturity energy -based model 152 may be further adjusted during training such that the function 153 assigns a lower value (e.g., energy value, score, and / or the like) to the protein sequence of an antigen receptor expressed by an immature lymphocyte whose edit distance, number of somatic hypermutations, and / or depth in the lymphocyte phylogenetic tree fails to satisfy the one or more criteria.
[0137] In some example embodiments, the modifying of the input protein sequence 142 may be guided by the output of the function 153 (e.g., gradient of energy function, score of score function, and / or the like) such that the output protein sequence 144 is sampled from a higher density region of the data distribution than the input protein sequence 142. In some cases, guiding the modifying of the input protein sequence 142 based on the output of the function 153 may ensure that the output protein sequence 144 corresponds to an antigen receptor expressed by a mature lymphocyte that has undergone adequate affinity maturation and is therefore more likely to exhibit the one or more properties of interest (e.g., affinity, avidity, anti-pathogen activity, and / or the like).
[0138] In some example embodiments, the protein design computation model 115 may modify the input protein sequence 142 while guided by the gradient of a total function combining the function 151 (e.g., energy function, score function, and / or the like) of the clonotype energybased model 150 and the function 153 (e.g., energy function, score function, and / or the like) of the maturity energy -based model 152. Doing so may increase the likelihood of the output protein sequence 144 being sampled from higher density regions of the data distributions approximated by both the clonotype energy -based model 150 and the maturity energy -based model 152 such that the corresponding antigen receptor is both from the target clonotype (or clonal family) andAttorney Ref.: 14786-074-228 (I03963-228074) / P40073-W0-1expressed by a mature lymphocyte that has undergone adequate affinity maturation. In some cases, the total function may be a product of experts or a weighted sum of the function 151 and the function 153. As such, it is possible for the output (e.g., energy value, score, and / or the like) of the function 151 to contribute more (or less) than the output (e.g., energy value, score, and / or the like) of the function 153 when guiding the modifying of the input protein sequence 142. This difference in contribution may mean that the sampling from the data distributions may prioritize more (or less) the output protein sequence 144 corresponding to an antigen receptor from the target clonotype (or clonal family) over the output protein sequence 144 corresponding to an antigen receptor expressed by a mature lymphocyte .
[0139] FIG. 2A depicts a flowchart illustrating an example of a process 200 for guided generation of antigen receptors, in accordance with some example embodiments. Referring to FIGS. 1 and 2A, the process 200 may be performed by the protein design engine 110 to generate, for example, the output protein sequence 144 by at least modifying the input protein sequence 142. As described in more detail below, the modifying of the input protein sequence 142 may be conditioned on one or more contextual variables. In instances where the one or more contextual variables define a target clonotype (or clonal family) of the lymphocyte expressing an antigen receptor, such as a clonotype (or clonal family) that is enriched in an antigen receptor repertoire, the conditioning may steer the generation of the output protein sequence 144 towards the target clonotype (or clonal family) such that the output protein sequence 144 corresponds to an antigen receptor expressed by a lymphocyte from the target clonotype (or clonal family). The contextual variables defining the target clonotype (or clonal family) may include the variable (V) gene segment and the joint (J) gene segment of a lymphocyte and the third complementarity determining region (CDR3) cluster of the antigen receptor expressed by the lymphocyte. Alternatively and / orAttorney Ref.: 14786-074-228 (I03963-228074) / P40073-W0-1additionally, the one or more contextual variables may define the maturity of the antigen receptor, which may be quantified by edit distance, quantity of hypermutations, and / or depth in a lymphocyte phylogenetic tree associated with the antigen receptor coding sequence of the lymphocyte expressing the antigen receptor. In some cases, conditioning the generation of the output protein sequence 144 on the target clonotype and / or maturity may increase the likelihood of the antigen receptor corresponding to the output protein sequence 144 exhibiting one or more properties of interest, such as affinity, avidity, anti-pathogen activity, and / or the like.
[0140] At 202, an antigen receptor repertoire having a plurality of clonotypes is identified. In some example embodiments, the antigen receptor repertoire may include the range of antigen receptors expressed by the total lymphocyte (e.g., B cell or T cell) population of a single animal or multiple animals of the same or different species upon exposure to an antigen (e.g., virus or cell surface receptor). In some cases, the antigen receptor repertoire may include antigen receptors expressed by lymphocytes from multiple clonal families. In some cases, each clonal family including a group of lymphocytes that evolved from the same ancestral lymphocyte (e.g., progenitor B cell or progenitor T cell) through mutations (e.g., variable-diversity-joining V(D)J rearrangement or recombination) to have similar but not necessarily identical antigen receptor coding segments (e.g., variable (V) gene segment, diversity (D) segment, and joining (J) gene segment). As such, in some cases, a single clonal family may include multiple clonotypes, each of which being a subset of lymphocytes with the same antigen receptor coding sequence (e.g., V(D)J gene segments).
[0141] At 204, a target clonotype is identified within the plurality of clonotypes in the antigen receptor repertoire. In some example embodiments, the target clonotype may be a clonotype that is enriched (or overrepresented) in the antigen receptor repertoire. For example, inAttorney Ref.: 14786-074-228 (I03963-228074) / P40073-W0-1some cases, where the antigen receptor repertoire includes a disproportionately large quantity of antigen receptors expressed by the lymphocytes of a single clonotype, that clonotype may be identified as a target clonotype. In some cases, the clonotype being enriched (or overrepresented) in the antigen receptor repertoire may indicate that the antigen receptors expressed by the lymphocytes in the clonotype are more likely to be binders of the antigen. Accordingly, as described in more detail below, the generation of protein sequences corresponding to antigen receptors may be conditioned on the clonotype of the lymphocytes generating the antigen receptors such that the resulting protein sequences correspond to antigen receptors expressed by lymphocytes from the target clonotype.
[0142] At 206, a maturity is determined for each antigen receptor in the antigen receptor repertoire. In some example embodiments, the maturity of an antigen receptor may correspond to the likelihood of the lymphocyte expressing the antigen receptor having undergone adequate affinity maturation. In some cases, the maturity of the antigen receptor may be quantified by the quantity of somatic hypermutations (SHM) present in the antigen receptor coding segment (e.g., V(D)J gene segments) of the lymphocyte expressing the antigen receptor. In some cases, the quantity of somatic hypermutations (SHM) may correspond to the edit distance between the antigen receptor coding segment (e.g., V(D)J gene segments) of the lymphocyte expressing the antigen receptor and the antigen receptor coding segment (e.g., V(D)J gene segments) of the ancestral lymphocyte (e.g., progenitor B cell or progenitor T cell) from which the lymphocyte evolved. Alternatively and / or additionally, the maturing of the antigen receptor may be quantified by the depth in the lymphocyte expressing the antigen receptor in the lymphocyte phylogenetic tree depicting the evolutionary relationships within the antigen receptor repertoire. In some cases, lymphocytes undergo affinity maturation to increase the properties of interest (e.g., affinity,Attorney Ref.: 14786-074-228 (I03963-228074) / P40073-W0-1avidity, anti-pathogen activity, and / or the like) present in the antigen receptors expressed by lymphocytes. Antigen receptors expressed by mature lymphocytes that have undergone adequate affinity maturation are therefore more likely to exhibit properties of interest, such as affinity, avidity, anti-pathogen activity, and / or the like. Accordingly, as described in more detail below, the generation of protein sequences corresponding to antigen receptors may be conditioned on maturity (e.g., instead of or in addition to clonotype (or clonal family)) such that the resulting protein sequences correspond to antigen receptors expressed by mature lymphocytes that have undergone adequate affinity maturation (e.g., a threshold quantity of cycles of affinity maturation).
[0143] At 208, a protein design computation model is trained to trained to approximate a data distribution of protein sequences of antigen receptors expressed by lymphocytes from the target clonotype and a data distribution of protein sequences of antigen receptors expressed by mature lymphocytes. In some example embodiments, the protein design computation model may be trained, based on the antigen receptor repertoire, to approximate the data distribution of the protein sequences of antigen receptors expressed by mature lymphocytes from the target clonotype. In some cases, the protein design computation model may include multiple energybased models, each of which being trained to approximate a separate data distribution. For example, in some cases, the protein design computation model may include a clonotype energybased model (EBM) trained to approximate the data distribution of the protein sequences of antigen receptors expressed by lymphocytes from the target clonotype. In some cases, the training of the clonotype energy -based model may include adjusting one or more parameters (e.g., weights, biases, and / or the like) of the clonotype energy-based model such that a function (e.g., energy function, score function, and / or the like) parameterized by the parameters of the clonotype energybased model assigns a higher value (e.g., energy value, score, and / or the like) to the proteinAttorney Ref.: 14786-074-228 (I03963-228074) / P40073-W0-1sequences of antigen receptors expressed by lymphocytes from the target clonotype and a lower value (e.g., energy value, score, and / or the like) to the protein sequences of antigen receptors expressed by lymphocytes outside of the target clonotype. Furthermore, in some cases, the protein design computation model may include a maturity energy-based model (EBM) trained to approximate the data distribution the protein sequences of antigen receptors expressed by mature lymphocytes that have undergone adequate affinity maturation (e.g., a threshold quantity of cycles of affinity maturation). In some cases, the training of the maturity energy-based model (EBM) may include adjusting one or more parameters (e.g., weights, biases, and / or the like) of the maturity energy-based model such that a function (e.g., energy function, score function, and / or the like) parametrized by the parameters of the maturity energy-based model assigns a higher value (e.g., energy value, score, and / or the like) to the protein sequences of antigen receptors expressed by mature lymphocytes that have undergone adequate affinity maturation and a lower value (e.g., energy value, score, and / or the like) to the protein sequences of antigen receptors expressed by immature lymphocytes that have not undergone adequate affinity maturation.
[0144] In some example embodiments, the protein design computation model may be trained to approximate the noisy data distributions of the protein sequences of antigen receptors expressed by mature lymphocytes from the target clonotype. For example, in some cases, the clonotype energy-based model (EBM) may be trained to approximate the noisy data distribution of the protein sequences of antigen receptors expressed by lymphocytes from the target clonotype while the maturity energy-based model (EBM) may be trained to approximate the noisy data distribution of the protein sequences of antigen receptors expressed by mature lymphocytes that have undergone adequate affinity maturation (e.g., a threshold quantity of cycles of affinity maturation). In some cases, the protein design computation model may be trained on a trainingAttorney Ref.: 14786-074-228 (I03963-228074) / P40073-W0-1dataset that include noisy protein sequences, each of which being the protein sequence of an antigen receptor from the antigen receptor repertoire that has been adulterated with noise (e.g., Gaussian noise and / or the like). In some cases, the protein design computation model may be trained to approximate the noisy data distribution instead of the true data distribution of the clean (or noiseless) protein sequences of the antigen receptors from the antigen receptor repertoire to increase the robustness of the protein design computation model. For instance, the true data distribution of the clean (or noiseless) protein sequences of the antigen receptors from the antigen receptor repertoire may exhibit steep changes between the high and low density regions of the data distribution Sampling from the true data distribution may be therefore prone to mode collapse, a phenomenon in which the steep density changes confines the protein design computation model to sampling from within the immediate vicinity of the protein sequences of the antigen receptors in the antigen receptor repertoire. Contrastingly, the noisy data distribution may exhibit smoother transitions between the high and low density regions of the data distribution, which support a more efficient exploration of the data distribution and enable the protein design computation model to generate a greater variety of output protein sequences.
[0145] At 210, the protein design computation model is applied to generate, based at least on an input protein sequence, an output protein sequence corresponding to the protein sequence of at least the portion of the antigen receptor expressed by mature lymphocytes from the target clonotype. In some example embodiments, the trained protein design computation model may be applied to generate the output protein sequence by modifying the input protein sequence. In some cases, the input protein sequence may be a random or a noise protein sequence without any known properties of interest. Alternatively, in some cases, the input protein sequence may correspond to a lead molecule, such as an antibody identified through an animal immunizationAttorney Ref.: 14786-074-228 (I03963-228074) / P40073-W0-1campaign as having binding affinity towards a target antigen, in which case the output protein sequence may be generated as a part of a hit expansion effort to generate variants of the lead molecule that also exhibit similar properties of interest.
[0146] In some cases, the input protein sequence may be modified by inserting, deleting, and / or changing the type of one or more amino acid residues in the input protein sequences. Moreover, in some cases, the input protein sequence may be modified over multiple iterations of gradient-based Markov Chain Monte Carlo (MCMC) sampling (e.g., Langevin Markov Chain Monte Carlo (MCMC) sampling), with each iteration of gradient-based Markov Chain Monte Carlo (MCMC) sampling further modifying the input protein sequence to drawing one or more samples (or modified protein sequences) from incrementally higher density regions of the data distributions of protein sequences expressed by mature lymphocytes from the target clonotype. In some cases, the modifying of the input protein sequence may be guided by the function (e.g., energy function, score function, and / or the like) of the clonotype energy -based model, the function (e.g., energy function, score function, and / or the like) of the maturity energy-based model, or a composition of the thereof. For example, in some cases, the protein design computation model may be applied to modify the input protein sequence to generate a first modified protein sequence and a second modified protein sequence. In some cases, the function of the clonotype energybased model, the function of the maturity energy-based model, or a composition of the two functions may be applied to determine a first value (e.g., energy value, score, and / or the like) indicative of the likelihood of the first modified protein sequence being an antigen receptor expressed by a mature lymphocyte from the target clonotype. Furthermore, in some cases, the function of the clonotype energy-based model, the function of the maturity energy-based model, or a composition of the two functions may be applied to determine a second value (e.g., energyAttorney Ref.: 14786-074-228 (I03963-228074) / P40073-W0-1value, score, and / or the like) indicative of the likelihood of the second modified protein sequence being an antigen receptor expressed by a mature lymphocyte from the target clonotype. In some cases, the protein design computation model may be applied to further modify the first modified protein sequence instead of the second modified protein sequence if the first value and the second value indicate that the first modified protein sequence is more likely to be an antigen receptor expressed by a mature lymphocyte from the target clonotype than the second modified protein sequence. As described in more detail below, one of the modified protein sequences generated by the protein design computation model may be generated as the output protein sequence when one or more criteria are satisfied.
[0147] In some example embodiments, where the protein design computation model is trained to approximate the noisy data distributions of the protein sequences of antigen receptors expressed by mature lymphocytes from the target clonotype, the protein design computation model may be applied to draw samples (or modified protein sequences) from the noisy data distributions. In some cases, each sample (or modified protein sequence) drawn from the noisy data distribution may include noise (e.g., Gaussian noise and / or the like). Accordingly, in some cases, the generating of the output protein sequence may include denoising the corresponding sample (or modified protein sequence) drawn from the noisy data distribution. Doing so may be tantamount to projecting the output protein sequence from the noisy data distribution back to the true data distribution of the clean (or noiseless) protein sequences of antigen receptors from the antigen receptor repertoire. In some cases, the transition from the noisy data distribution (e.g., noisy continuous data distribution) back to the discrete sequence space may constitute a Neural Empirical Bayes projection using, for example, a least-squares estimator. In some cases, samples (or modified protein sequences) drawn during at least some previous iterations of gradient-basedAttorney Ref.: 14786-074-228 (I03963-228074) / P40073-W0-1Markov Chain Monte Carlo (MCMC), which correspond to the intermediate states between the input protein sequence and the output protein sequence, may also be denoised to enable an inspection of the trajectory of the gradient-based Markov Chain Monte Carlo (MCMC).
[0148] FIG. 2B depicts a flowchart illustrating another example of a process 250 for guided generation of antigen receptors, in accordance with some example embodiments. Referring to FIGS. 1 and 2B, the process 250 may be performed by the protein design engine 110 to generate, for example, the output protein sequence 144 by at least modifying the input protein sequence 142. As described in more detail below, the modifying of the input protein sequence 142 may be conditioned on one or more contextual variables. In instances where the one or more contextual variables define a target clonotype (or clonal family) of the lymphocyte expressing an antigen receptor, such as a clonotype (or clonal family) that is enriched in an antigen receptor repertoire, the conditioning may steer the generation of the output protein sequence 144 towards the target clonotype (or clonal family) such that the output protein sequence 144 corresponds to an antigen receptor expressed by a lymphocyte from the target clonotype (or clonal family). In some cases, the contextual variables defining the target clonotype (or clonal family) may include the variable (V) gene segment and the joint (J) gene segment of a lymphocyte and the third complementarity determining region (CDR3) cluster of the antigen receptor expressed by the lymphocyte.
[0149] At 252, an antigen receptor repertoire having a plurality of clonotypes is identified. In some example embodiments, the antigen receptor repertoire may include the range of antigen receptors expressed by the total lymphocyte (e.g., B cell or T cell) population of a single animal or multiple animals of the same or different species upon exposure to an antigen (e.g., virus or cell surface receptor). In some cases, the antigen receptor repertoire may include antigen receptors expressed by lymphocytes from multiple clonal families. As noted, in some cases, eachAttorney Ref.: 14786-074-228 (I03963-228074) / P40073-W0-1clonal family may include a group of lymphocytes that evolved from the same ancestral lymphocyte (e.g., progenitor B cell or progenitor T cell) through mutations (e.g., variablediversity-joining V(D)J rearrangement or recombination). As such, in some cases, the lymphocytes in each clonal family may have similar but not necessarily identical antigen receptor coding segments (e.g., variable (V) gene segment, diversity (D) segment, and joining (J) gene segment). Moreover, in some cases, a single clonal family may include multiple clonotypes, each of which being a subset of lymphocytes with the same antigen receptor coding sequence (e.g., V(D)J gene segments).
[0150] At 254, a target clonotype is identified within the plurality of clonotypes in the antigen receptor repertoire. In some example embodiments, the target clonotype may be a clonotype that is enriched (or overrepresented) in the antigen receptor repertoire. In some cases, an enriched clonotype in the antigen receptor repertoire may be a clonotype whose member lymphocytes express a disproportionately large quantity of antigen receptors in the antigen receptor repertoire. In some cases, the generation of protein sequences corresponding to antigen receptors may be conditioned on an enriched clonotype at least because the clonotype being enriched (or overrepresented) in the antigen receptor repertoire may be indicative of the antigen receptors expressed by the member lymphocytes being more likely to be binders of the antigen.
[0151] At 256, a protein design computation model is trained to approximate a data distribution of protein sequences of antigen receptors expressed by lymphocytes from the target clonotype. In some example embodiments, the protein design computation model may be trained, based on the antigen receptor repertoire, to approximate the data distribution of the protein sequences of antigen receptors expressed by lymphocytes from the target clonotype (e.g., an enriched clonotype and / or the like) in the antigen receptor repertoire. In some cases, the dataAttorney Ref.: 14786-074-228 (I03963-228074) / P40073-W0-1distribution may include high density regions populated by protein sequences of antigen receptors expressed by lymphocytes from the target clonotype and low density regions populated by protein sequences of antigen receptors expressed by lymphocytes outside of the target clonotype. Accordingly, in some cases, the protein design computation model may include a clonotype energy-based model (EBM) trained to approximate the data distribution. In some cases, one or more parameters (e.g., weights, biases, and / or the like) of the clonotype energy-based model may be adjusted during training such that a function (e.g., energy function, score function, and / or the like) parameterized by the parameters of the clonotype energy -based model assigns a higher value (e.g., energy value, score, and / or the like) to the protein sequences of antigen receptors expressed by lymphocytes from the target clonotype and a lower value (e.g., energy value, score, and / or the like) to the protein sequences of antigen receptors expressed by lymphocytes outside of the target clonotype. The protein design computation model may be used to approximate the data distribution of protein sequences of antigen receptors expressed by lymphocytes from the target clonotype by at least applying the function (e.g., energy function, score function, and / or the like) to determine the density of the region from which one or more protein sequences are drawn. The value (e.g., energy value, score, and / or the like) output by the function (e.g., energy function, score function, and / or the like) may indicate transitions between high and low density regions of the data distribution, thus identifying a protein sequence as corresponding to an antigen receptor expressed by a lymphocyte from the target clonotype or outside of the target clonotype.
[0152] At 258, the protein design computation model is applied to generate, based at least on an input protein sequence, an output protein sequence corresponding to at least a portion of an antigen receptor expressed by a lymphocyte from the target clonotype. In some example embodiments, the trained protein design computation model may be applied to generate the outputAttorney Ref.: 14786-074-228 (I03963-228074) / P40073-W0-1protein sequence by modifying the input protein sequence. In some cases, the input protein sequence may be a random or a noise protein sequence without any known properties of interest. Alternatively, in some cases, the input protein sequence may correspond to a lead molecule. For example, in some cases, the input protein sequence may be an antibody identified through an animal immunization campaign as having binding affinity towards a target antigen. In some cases, the input protein sequence may exhibit one or more properties of interest (e.g., binding affinity) attributable to the input protein sequence being expressed by a lymphocyte from the target clonotype. Accordingly, in some cases, the input protein sequence may be modified while conditioned on the target clonotype to generate the output protein sequence as a part of a hit expansion effort to generate variants of the lead molecule that also exhibit similar properties of interest.
[0153] In some cases, modifying the input protein sequence may be tantamount to applying the data distribution approximated by the clonotype energy-based model (EBM) of the trained protein design computation model to guide the modifying of the input protein sequence. In some cases, the modifying of the input protein sequence may include inserting, deleting, and / or changing the type of one or more amino acid residues in the input protein sequences. Moreover, in some cases, the input protein sequence may be modified over multiple iterations of gradientbased Markov Chain Monte Carlo (MCMC) sampling (e.g., Langevin Markov Chain Monte Carlo (MCMC) sampling), with each iteration of gradient-based Markov Chain Monte Carlo (MCMC) sampling applying the data distribution to further modify the input protein sequence. In some cases, modifying the input protein sequence may be tantamount to drawing one or more samples (or modified protein sequences) from the data distribution with guidance from the function (e.g., energy function, score function, and / or the like) parameterized by the clonotype energy-basedAttorney Ref.: 14786-074-228 (I03963-228074) / P40073-W0-1model (EBM). In some cases, with guidance from the values (e.g., energy values, scores, and / or the like) output by the function (e.g., energy function, score function, and / or the like), successive samples (or modified protein sequences) may be drawn from incrementally higher density regions of the data distributions of protein sequences expressed by mature lymphocytes from the target clonotype. For example, in some cases, the protein design computation model may be applied to modify the input protein sequence to generate a first modified protein sequence and a second modified protein sequence. In some cases, the function (e.g., energy function, score function, and / or the like) of the clonotype energy-based model may be applied to determine a first value (e.g., energy value, score, and / or the like) indicative of the likelihood of the first modified protein sequence being an antigen receptor expressed by a lymphocyte from the target clonotype. Furthermore, in some cases, the function (e.g., energy function, score function, and / or the like) of the clonotype energy -based model may be applied to determine a second value (e.g., energy value, score, and / or the like) indicative of the likelihood of the second modified protein sequence being an antigen receptor expressed by a lymphocyte from the target clonotype. In some cases, the protein design computation model may be applied to further modify the first modified protein sequence instead of the second modified protein sequence if the first value and the second value indicate that the first modified protein sequence is more likely to be an antigen receptor expressed by a lymphocyte from the target clonotype than the second modified protein sequence. In some cases, one of the modified protein sequences generated by the protein design computation model may be generated as the output protein sequence when one or more criteria are satisfied.
[0154] FIG. 2C depicts a flowchart illustrating another example of a process 270 for guided generation of antigen receptors, in accordance with some example embodiments. Referring to FIGS. 1 and 2C, the process 270 may be performed by the protein design engine 110 to generate,Attorney Ref.: 14786-074-228 (I03963-228074) / P40073-W0-1for example, the output protein sequence 144 by at least modifying the input protein sequence 142. As described in more detail below, the modifying of the input protein sequence 142 may be conditioned on one or more contextual variables defining a context. In some cases, the context may include a target clonotype (or clonal family) of the lymphocyte expressing an antigen receptor, such as a clonotype (or clonal family) that is enriched in an antigen receptor repertoire, in which case the one or more contextual variables may include the variable (V) gene segment, the joint (J) gene segment, and third complementarity determining region (CDR3) length. Alternatively and / or additionally, the context may include the maturity of the antigen receptor, in which case the one or more contextual variables may include edit distance, quantity of hypermutations, and / or depth in a lymphocyte phylogenetic tree associated with the antigen receptor coding sequence of the lymphocyte expressing the antigen receptor. In some cases, conditioning the generation of the output protein sequence 144 on the context defined by the one or more contextual variables may steer the protein design computation model 115 to generate (or sample) from neighborhoods in a data distribution populated by protein sequences having the context, thus increasing the likelihood of the antigen receptor corresponding to the output protein sequence 144 exhibiting one or more properties of interest associated with the context (e.g., affinity, avidity, anti-pathogen activity, and / or the like).
[0155] At 272, an antigen receptor repertoire having a plurality of different contexts is identified. In some example embodiments, the antigen receptor repertoire may include the range of antigen receptors expressed by the total lymphocyte (e.g., B cell, T cell, and / or the like) population of one or more animals exposed to an antigen (e.g., virus, cell surface receptor, and / or the like). In some cases, the antigen receptors in the antigen receptor repertoire may include antigen receptors having different contexts. One example of context includes the clonotype (orAttorney Ref.: 14786-074-228 (I03963-228074) / P40073-W0-1clonal family) of the lymphocytes expressing the antigen receptor. As noted, the antigen receptor repertoire may include antigen receptors expressed by lymphocytes from multiple clonal families, each of which including a group of lymphocytes evolved from the same ancestral lymphocyte (e.g., progenitor B cell or progenitor T cell) through mutations (e g., variable-diversity-joining V(D)J rearrangement or recombination). Another example of context includes the maturity of the lymphocyte expressing the antigen receptor. In some cases, the maturity of the lymphocyte may correspond to the extent of affinity maturation, a process including somatic hypermutation and clonal selection to incrementally increase the affinity of the antigen receptor.
[0156] At 274, a target context is identified within the plurality of different contexts present in the antigen receptor repertoire. In some example embodiments, the target context may include one or more target clonotypes (or clonal families). In some cases, each target clonotype may be a clonotype that is enriched (or overrepresented) in the antigen receptor repertoire. In some cases, an enriched clonotype in the antigen receptor repertoire may be a clonotype whose member lymphocytes express a disproportionately large quantity of antigen receptors in the antigen receptor repertoire. In some cases, antigen receptors expressed by lymphocytes from a target clonotype (or clonal family) may be more likely to exhibit one or more properties of interest, such as binding affinity towards a target antigen. In some cases, conditioning the generation of protein sequences corresponding to antigen receptors on a context including one or more target clonotypes (or clonal families) may result in the generation of antigen receptors expressed by lymphocytes from the one or more target clonotypes (or clonal families). Alternatively and / or additionally, the target context may include the maturity of the lymphocyte, which may be quantified edit distance, quantity of hypermutations, and / or depth in a lymphocyte phylogenetic tree associated with the antigen receptor coding sequence of the lymphocyte expressing the antigenAttorney Ref.: 14786-074-228 (I03963-228074) / P40073-W0-1receptor. In some cases, antigen receptors expressed by more mature lymphocytes, or lymphocytes that have undergo affinity maturation (or more cycles of affinity maturation), may be more likely to exhibit one or more properties of interest (e.g., binding affinity towards a target antigen). Accordingly, conditioning the generation of protein sequences corresponding to antigen receptors on a context including maturity may yield antigen receptors expressed by more mature lymphocytes.
[0157] At 276, a protein design computation model is trained to approximate a data distribution of protein sequences of antigen receptors exhibiting the target context. In some example embodiments, the protein design computation model may be trained, based on the antigen receptor repertoire, to approximate the data distribution of the protein sequences of antigen receptors exhibiting the target context (e.g., target clonotype (or clonal family), maturity, and / or the like). In some cases, the data distribution may include high density regions with a greater concentration of protein sequences of antigen receptors exhibiting the target context and low density regions with a lower concentration of protein sequences of antigen receptors without the target context. In some cases, the protein design computation model may include one or more energy-based models (EBMs). In some cases, the protein design computation model may include an ensemble of multiple energy -based models (EMBs), each of which being trained to approximate the data distribution of protein sequences exhibiting a different context. For example, in some cases, the protein design computation model may include a clonotype energy -based model trained to approximate the data distribution of protein sequences of antigen receptors expressed by lymphocytes from one or more target clonotypes (or clonal families). Furthermore, in some cases, the protein design computation model may include a maturity energy -based model (EBM) trainedAttorney Ref.: 14786-074-228 (I03963-228074) / P40073-W0-1to approximate the data distribution of protein sequences of antigen receptors expressed by mature lymphocytes (or lymphocytes that have undergone adequate affinity maturation).
[0158] In some example embodiments, the energy-based model (EBM) implementing the protein design computation model may be trained using contrastive divergence. In some cases, the training of the energy -based model (EBM) may further include data augmentation (e.g., contrastive divergence with data augmentation). In some cases, the training of the energy-based model (EBM) may include reducing (or minimizing) a contrastive divergence loss. In some cases, reducing (or minimizing) the contrastive divergence loss may encourage the model to assign lower energy to positive samples and higher energy to negative samples. For example, in some cases, a positive sample (ypOs’cPos)maY include a noised sequence ypos— x + J\f(O, <j2 / d) from the training dataset and its corresponding context c without noise added. In some cases, to create a smoother energy landscape, negative samples (yneg, cneg) may be augmented such that more negatives are sampled than positives. For instance, in some cases, three negatives may be sampled for each positive, each in a different way. In some cases, some negative samples may be generated when a random seed sequence undergoes Langevin Markov Chain Monte Carlo (MCMC) while guided by the score of an underfit energy-based model (EBM). In some cases, some negative samples may be generated by pairing the noisy positive sequencewith an incorrect context, such as an incorrect protein subclass. In some cases, the negative sample (ypOs'crand)may include a randomly drawn protein subclass crandas context while the negative sample (ypos,cfamiiy) may include a related but still incorrect protein subclass Cfamilyas context. Where the context is clonotype (or clonal family), the context c^amUymay denote a clonotype from the same gene family but a different gene.Attorney Ref.: 14786-074-228 (I03963-228074) / P40073-W0-1
[0159] In some example embodiments, instead of an energy-based model (EBM), it may also be possible for the protein design computation model to be at least partially implemented using a score-based model g^ (y, c). In some cases, the score-based model g<p (y, c) may be trained to directly approximate the score function characterizing the energy landscape of the data distribution of protein sequences exhibiting the target context. In some cases, the score-based model g<p(y, c) may be trained with a denoising objective with the context embedded and added to intermediate representations of the protein sequence to focus denoising on the target contexts (e g., target clonotype (or clonal family), maturity, and / or the like). In some cases, during the training of the score-based model g<p(y, c), the context in some training samples may be randomly replaced with a null 0 token with a certain probability (e.g., puncOnd = 0.2). Doing so may enable the score-based model g^y, c) to sample unconditionally when a null context 0 is encountered. In some cases, an unconditional score-based modelmay also permit the tuning of guidance level from context conditioning during sampling. That is, in some cases, sampling from the protein design computation model may include a linear combinationof the conditional score estimate g<p(.y,c and unconditional score estimate g<p(y, 0) with a weight m stipulating the level of guidance. In some cases, training the protein design computation model to approximate both the conditional score estimate gtp y.c) and unconditional score estimatemay enable subsequent generation of output sequences, for example, by sampling from the data distribution approximated by the trained protein design computation model, to be performed with predictor-free guidance towards output sequences exhibiting one or more properties of interest.
[0160] At 278, the protein design computation model may be applied to generate an output sequence corresponding to at least a portion of an antigen receptor having the target context.Attorney Ref.: 14786-074-228 (I03963-228074) / P40073-W0-1In some example embodiments, the trained protein design computation model may be applied to generate the output protein sequence by modifying an input protein sequence (e.g., a noise sequence of randomly ordered sequence of amino acid residues without any known properties). In some cases, applying the trained protein design computation model in this manner may be tantamount to sampling modified protein sequences from the data distribution approximated by the one or more energy -based models (EBMs) included in the protein design computation model. In some cases, the modifying of the input protein sequence may be conditioned on one or more contextual variables defining the target context. For example, in some cases, the one or more contextual variables may define one or more target clonotypes (or clonal families), maturity, and / or the like. Furthermore, in some cases, the one or more contextual variables may include categorical variables, numerical variables, and / or the like. For instance, in some cases, the one or more contextual variables may include categorical variables defining a variable (V) gene segment, a joint (J) gene segment, and / or the like. Alternatively and / or additionally, in some cases, the one or more contextual variables may include numerical variables defining a third complementarity determining region (CDR3) length, quantity of somatic hypermutations (SHMs), and / or the like.
[0161] In some example embodiments, the input protein sequence may be modified over multiple iterations of gradient-based Markov Chain Monte Carlo (MCMC) sampling (e.g., Langevin Markov Chain Monte Carlo (MCMC) sampling), with each iteration of gradient-based Markov Chain Monte Carlo (MCMC) sampling applying the data distribution to further modify the input protein sequence. In some cases, the modifying of the input protein sequence may be guided by a function (e.g., energy function, score function, and / or the like) parameterized by the one or more energy-based models (EBMs). In some cases, with guidance from the values (e.g., energy values, scores, and / or the like) output by the function (e.g., energy function, score function,Attorney Ref.: 14786-074-228 (I03963-228074) / P40073-W0-1and / or the like), successive samples (or modified protein sequences) may be drawn from incrementally higher density regions of the corresponding data distributions, which are more densely populated by protein sequences exhibiting the target context (e.g., target clonotype (or clonal family), maturity, and / or the like).
[0162] In some example embodiments, the energy-based model (EBM) implementing the protein design computation model may a conditional energy-based model (EBM) adapted to ingest a context c = (q, c2, ••• , cn). In some cases, each context c, may be embedded separately, followed by concatenation and projection to size dcontext. In some cases, categorical variables, such as variable (V) gene segment and joint (J) gene segment, may be one-hot encoded prior to being embedded, for example, by an embedding layer. In some cases, numerical variables, such as third complementarity determining region (CDR3) length and quantity of somatic hypermutations (SHMs), may be encoded using clustering encoding. In some cases, the full context embedding may be split into gain and bias terms to modify the embedding of the noisy input protein sequence.
[0163] In some example embodiments, Langevin dynamics with explicit inclusion of the target context may be applied to generate the output sequence during inference. For example, where the protein design computation model is implemented using one or more conditional energybased models (EBMs), the modified protein sequence that is generated at each step k may be given by<wherein 8 denotes step size, Vydenotes the gradient of the energy functionwith respect to noisy data y, and e denotes additional noise.Attorney Ref.: 14786-074-228 (I03963-228074) / P40073-W0-1
[0164] In some example embodiments, where the protein design computation model is implemented using one or more score-based models, the term Vy / e may be replaced with the learned approximation of the score function, denoted as s^ y, c, a>). In some cases, after T steps, clean discrete protein sequences may be recovered by returning from the noisy data distribution to a clean data distribution through denoising using a least square estimator as indicated below.<>It should be appreciated that with either the energy-based or the score-based implementation of the protein design computation model, the linear combinationof the conditional score estimate g<p(y, c) and unconditional score estimate g^ y, 0) may be used to denoise. In some cases, the context ctmay be specified explicitly during sampling to simplify inference.
[0165] FIG. 3 depicts a flowchart illustrating an example of a process 300 for guided generation of antigen receptors with gradient-based Markov Chain Monte Carlo (MCMC) sampling, in accordance with some example embodiments. Referring to FIGS. 1-3, the process 300 may be performed by the protein design engine 110 to generate, for example, the output protein sequence 144 by at least modifying the input protein sequence 142. In some cases, the process 300 may implement at least a portion of the operation 210 shown in FIG. 2A. As described in more detail below, the protein design engine 110 may apply the protein design computation model 115 to modify the input protein sequence 142 over multiple iterations of gradient-based Markov Chain Monte Carlo while guided by a function, such as an energy function, a score function, and / or the like. Moreover, the modifying of the input protein sequence 142 may be conditioned on one or more contextual variables defining a context such as, for example, a target clonotype (or clonal family) and / or maturity of the lymphocyte expressing an antigen receptor corresponding to the resulting output protein sequence 144. For example, in some cases, the modifying of the inputAttorney Ref.: 14786-074-228 (I03963-228074) / P40073-W0-1protein sequence 142 may be guided by the function (e g., energy function, score function, and / or the like) to sample from incrementally higher density regions of the data distributions populated by the protein sequences of antigen receptors expressed by mature lymphocytes from the target clonotype. In some cases, conditioning the generation of the output protein sequence 144 on context, such as clonotype and / or maturity, may increase the likelihood of the antigen receptor corresponding to the output protein sequence 144 exhibiting one or more properties of interest, such as affinity, avidity, anti-pathogen activity, and / or the like.
[0166] At 302, a trained protein design computation model is applied to generate a first modified protein sequence having a first modification to an input protein sequence. In some example embodiments, the protein design computation model may be trained to approximate the data distributions of the protein sequence of antigen receptors expressed by mature lymphocytes from a target clonotype, such as a clonotype that is enriched (or overrepresented) in the antigen receptor repertoire of a single animal or multiple animals of the same or different species. In some cases, the protein design computation model may be trained to approximate the noisy data distributions of noisy protein sequences generated by adding noise (e.g., Gaussian noise and / or the like) to the protein sequence of antigen receptors expressed by mature lymphocytes from the target clonotype. In some cases, once trained, the trained protein design computation model may be applied to modify the input protein sequence including by inserting, deleting, and / or changing the type of one or more amino acid residues in the input protein sequence. In some cases, the modifying of the input protein sequence may be tantamount to drawing a sample (or a modified protein sequence) from the data distributions (or noisy data distributions) of the protein sequence of antigen receptors expressed by mature lymphocytes from the target clonotype. As described in more detail below, the trained protein design computation model may be applied to modify theAttorney Ref.: 14786-074-228 (I03963-228074) / P40073-W0-1input protein sequence over multiple iterations gradient-based Markov Chain Monte Carlo (MCMC) sampling (e.g., Langevin Markov Chain Monte Carlo (MCMC) sampling). Accordingly, in some cases, additional samples (or modified protein sequences) may be drawn from the data distributions (or noisy data distribution) by applying the trained protein design computation model to further modify the input protein sequence or a modified protein sequence from a previous iteration of gradient-based Markov Chain Monte Carlo (MCMC) sampling.
[0167] At 304, the trained protein design computation model may be applied to generate a second modified protein sequence having a second modification to the input protein sequence. In some example embodiments, the trained protein design computation model may be applied to modify the input protein sequence and draw a different sample (or modified protein sequence) from the data distributions (or noisy data distributions) of the protein sequence of antigen receptors expressed by mature lymphocytes from the target clonotype (e.g., an enriched (or overrepresented) clonotype in an antigen receptor repertoire). In some cases, the trained protein design computation model may be applied to draw multiple samples (or modified protein sequences) during a single iteration of gradient-based Markov Chain Monte Carlo (MCMC) sampling (e.g., Langevin Markov Chain Monte Carlo (MCMC) sampling). In some cases, different modified protein sequences may be generated by applying different modifications to the same protein sequence from a previous iteration of gradient-based Markov Chain Monte Carlo (MCMC) sampling. Alternatively and / or additionally, different modified protein sequences may be generated by modifying different protein sequences from the previous iteration of gradient-based Markov Chain Monte Carlo (MCMC) sampling. Some or all of the samples (or modified protein sequences) from a current iteration of gradient-based Markov Chain Monte Carlo (MCMC) sampling (e g., Langevin MarkovAttorney Ref.: 14786-074-228 (I03963-228074) / P40073-W0-1Chain Monte Carlo (MCMC) sampling) may advance to a subsequent iteration for further modification by the trained protein design computation model.
[0168] At 306, that an antigen receptor corresponding to the first modified protein sequence is more likely to be expressed by a mature lymphocyte from a target clonotype than an antigen receptor corresponding to the second modified protein sequence is determined based at least on an output of one or more functions parameterized by the trained protein design computation model. In some example embodiments, the functions (e.g., energy functions, score functions, and / or the like) parameterized the trained protein design computation model may include a function parameterized by a clonotype energy-based model (EBM) and a function parameterized by a maturity energy-based model (EBM). These functions may output values (e.g., energy values, scores, and / or the like) indicative of the density of the data distribution from which each sample (or modified protein sequence) generated by the protein design computation model is drawn. As such, in some cases, the modifying of the input protein sequence may be guided by the functions including, for example, the values (e.g., scores) output by the functions, the gradients (e.g., changes in energy values) of the functions, and / or the like. For example, in some cases, the function parameterized by the clonotype energy-based model (EBM) may assign a higher value (e.g., energy value, score, and / or the like) to the protein sequences of antigen receptors (e.g., B-cell receptors (BCRs) or T-cell receptors (TCRs)) belonging to the target clonotypes and a lower value (e.g., energy value, score, and / or the like) to the protein sequences of those antigen receptors outside of the target clonotypes. Alternatively and / or additionally, the function parameterized by the maturity energy -based model (EBM) may a higher value (e.g., energy value, score, and / or the like) to protein sequences of antigen receptors expressed by mature lymphocytes and a lower value (e.g., energy value, score, and / or the like) to the protein sequences of antigen receptors mayAttorney Ref.: 14786-074-228 (I03963-228074) / P40073-W0-1immature lymphocytes. Tn some cases, the value output by a composition (e.g., product of experts, weighted sum, and / or the like) of the two functions for a protein sequence may indicate whether the corresponding antigen receptor is expressed by a mature lymphocyte from the target clonotype. For two samples (or modified protein sequences) drawn from the data distributions (or noisy data distributions) of the protein sequences of antigen receptors expressed by mature lymphocytes from the target clonotype during a single iteration of gradient-based Markov Chain Monte Carlo (MCMC) sampling, it may be possible that one sample (or modified protein sequence) is drawn from a higher density region of the data distributions than the other sample (or modified protein sequence). The one sample (or modified protein sequence) drawn from the higher density region of the data distributions may be more likely to correspond to an antigen receptor expressed by a mature lymphocyte from the target clonotype. As such, as described in more detail below, the sample (or modified protein sequence) that is drawn from the higher density region of the data distributions may advance to a subsequent iteration of gradient-based Markov Chain Monte Carlo (MCMC) sampling for further modifications by the trained molecule design computation model while the sample (or modified protein sequence) drawn from the lower density region of the data distributions may be preventing from advancing to the subsequent iteration.
[0169] At 308, the trained protein design computation model is applied to generate a third modified protein sequence by further modifying the first modified protein sequence instead of the second modified protein sequence. In some example embodiments, the protein design computation model may be applied to further modify the sample (or modified protein sequence) drawn from the higher density region of the data distributions (or noisy data distributions) instead of the sample (or modified protein sequence) from the lower density region of the data distributions (or noisy data distributions) at least because the sample (or modified protein sequence) from theAttorney Ref.: 14786-074-228 (I03963-228074) / P40073-W0-1higher density region of the data distributions (or noisy data distributions) may be more likely to correspond to an antigen receptor expressed by a mature lymphocyte from the target clonotype. For example, in some cases, the protein design computation model may be applied to further modify the sample (or modified protein sequence) by inserting, deleting, and / or changing the identity of one or more amino acid residues. In some cases, the protein design computation model may be applied over multiple iterations of gradient-based Markov Chain Monte Carlo (MCMC) sampling (e.g., Langevin Markov Chain Monte Carlo (MCMC) and / or the like) while guided by the output of the functions (e.g., energy functions, score functions, and / or the like) parameterized by the protein design computation model (e.g., the clonotype energy -based model, the maturity energy-based model, and / or the like). For instance, with each iteration of gradient-based Markov Chain Monte Carlo (MCMC) sampling, the protein design computation model may be applied to further modify samples (or modified protein sequences) drawn from incrementally higher density regions of the data distributions (or noisy data distributions) of the protein sequence of antigen receptors expressed by mature lymphocytes from the target clonotype (e.g., an enriched (or overrepresented) clonotype in an antigen receptor repertoire). Doing so may ensure that the resulting output protein sequence is sampled from a sufficiently high density region of the data distributions (or noisy data distributions). A sufficiently high density region of the data distributions (or noisy data distributions) are populated by the protein sequences of antigen receptors expressed by mature lymphocytes from the target clonotype, meaning that the output protein sequences may be more likely to exhibit the one or more properties of interest (e.g., affinity, avidity, anti-pathogen activity, and / or the like) associated with antigen receptors expressed by mature lymphocytes from the target clonotype and / or mature lymphocytes that have undergo adequate affinity maturation (e.g., a threshold quantity of cycles of affinity maturation).Attorney Ref.: 14786-074-228 (I03963-228074) / P40073-W0-1
[0113] At 310, upon satisfying one or more criteria, an output protein sequence corresponding to a modified protein sequence generated by the trained protein design computation model is generated. In some example embodiments, the protein design computation model may be applied to modify the input protein sequence over multiple iterations of gradient-based Markov Chain Monte Carlo (MCMC) sampling (e.g., Langevin Markov Chain Monte Carlo (MCMC) sampling and / or the like) until one or more criteria are satisfied. In some cases, the one or more criteria may include the output protein sequence being drawn from a sufficiently high density region of the data distributions (or noisy data distributions) of the protein sequences of antigen receptors expressed by mature lymphocytes from the target clonotype. Alternatively and / or additionally, the oner or more criteria may include the protein design computation model having performed a threshold quantity of iterations of gradient-based Markov Chain Monte Carlo (MCMC) sampling (e.g., Langevin Markov Chain Monte Carlo (MCMC) sampling and / or the like). In instances where the protein design computation model draws samples (or modified protein sequences) from the noisy data distributions of the protein sequences of antigen receptors expressed by mature lymphocytes from the target clonotype, the generating of the output protein sequence may further include denoising the corresponding modified protein sequence drawn from the noisy data distributions. In some cases, the denoising of the modified protein sequence may include removing the noise (e.g., Gaussian noise and / or the like) that is present in the modified protein sequence, thereby projecting the modified protein sequence from the noisy data distributions to the true data distributions of the protein sequences of antigen receptors expressed by mature lymphocytes from the target clonotype.
[0114] To further illustrate, consider an antigen receptor (e g., B-cell receptor (BCR) or T-cell receptor (TCR) sequence x and the contextual variables (e.g., variable (V) gene segment,Attorney Ref.: 14786-074-228 (I03963-228074) / P40073-W0-1joint (J) gene segment, and third complementarity determining region (CDR3) cluster) defining the clonotype of the antigen receptor x. In some cases, a protein design computation model (e.g., the protein design computation model 115 in FIG. 1) may be applied to generate output protein sequences that correspond to antigen receptors expressed from lymphocytes from the target clonotype that have undergone adequate affinity maturation (e.g., a threshold quantity of cycles of affinity maturation). In some cases, the maturity of a lymphocyte expressing an antigen receptor may be quantified by the quantity of somatic hypermutations (SHM) present in the antigen receptor coding sequence of the lymphocyte or the depth in the lymphocyte phylogenetic tree occupied by the lymphocyte.
[0115] In some example embodiments, separate functions (e.g., energy functions, score functions, and / or the like) may be defined for target clonotype and maturity. For example, one function (e.g., energy function, score function, and / or the like) may approximate the data distribution of the protein sequences of antigen receptors expressed by lymphocytes from the target clonotype while another function (e.g., energy function, score function, and / or the like) may approximate the data distribution of the protein sequences of antigen receptors expressed by lymphocytes that have undergone adequate affinity maturation (e.g., a threshold quantity of cycles of affinity maturation). Moreover, each function may be parameterized by the parameters (e.g., weights, biases, and / or the like) of a corresponding energy-based model (EBM) implemented, for example, using a neural network. This formulation is shown below^(y,c,m):ALx [K x [KJ X [K3] - Bwherein ALis a length L sequence of amino acid residues from the set of possible amino acid residues A (e.g., canonical amino acid residues), [Ki] denotes the set {1, ... , KJ of the possibleAttorney Ref.: 14786-074-228 (I03963-228074) / P40073-W0-1values of the first contextual variable [ / C2] denotes the set {1, ... , f2} °f the possible values of the second contextual variable [fc2], and [ / <3] denotes the set {1, ... , K3} of the possible values of the third contextual variable [fc3] .
[0116] In some cases, the target clonotype c may be defined by the three contextual variables K2. and / <3, representing the variable (V) gene segment, joining (J) gene segment, and third complementarity determining region (CDR3) cluster. In other words, the target clonotype c may be represented as c = [fc1(k2, k3],
[0117] As shown below, in some cases, the total energy function E (y, c, m) may be a product of experts or a weighted sum of the individual energy functions Ecand EA.E (y, c, m) = Ec(y, c) + EA(y, c, m)wherein A controls the contribution of each energy function and m is a hidden variable corresponding to the positions relevant to the maturity of the antigen receptor.
[0118] To generate one or more output protein sequences, each of which corresponding to at least a portion of an antigen receptor, the protein design computation model (e.g., the protein design computation model 115 shown in FIG. 1) may be applied to modify an input protein sequence over multiple iterations of gradient-based Markov Chain Monte Carlo (MCMC) sampling (e.g., Langevin Markov Chain Monte Carlo (MCMC) sampling and / or the like). In some cases, the modification of the input protein sequence at each iteration k of gradient-based Markov Chain Monte Carlo (MCMC) sampling may be defined as follows:wherein 6 denotes the step size anddenotes the gradient with respect to noisy protein sequence y. After T iterations of gradient-based Markov Chain Monte Carlo (MCMC) samplingAttorney Ref.: 14786-074-228 (I03963-228074) / P40073-W0-1or once the target clonotype is reached, y may be further refined with gradient descent (or discrete gradient flow) defined as follows:<
[0119] It should be appreciated that this second set of gradient-based Markov Chain Monte Carlo (MCMC) sampling iterations may represent smaller refines within the local neighborhood of the target clonotype. A smaller step size x] may be used and noise may be excluded to limit exploration while maximizing the maturity of the sampled sequences. Following this second set of gradient-based Markov Chain Monte Carlo (MCMC) sampling iterations, clean output protein sequences may be recovered, for example, using a least squares estimator as shown below, wherein c represents the target clonotype.
[0120] In some example embodiments, the energy-based models, such as the clonotype energy-based model (EBM) and maturity energy-based model (EBM), may be trained using a contrastive loss that encourages the models to assign lower values (e.g., energy values) to positive samples (e.g., protein sequences of antigen receptors expressed by mature lymphocytes from the target clonotype) and higher values (e.g., energy values) to negative samples (e.g., protein sequences of antigen receptors expressed by immature lymphocytes and / or lymphocytes outside of the target clonotype). The loss function £ may be defined as follows:wherein Dposand Dnegdenote the distributions of positive and negative samples, respectively. A positive sample (ypos, cpos) may include a noisy protein sequence yposfrom the training dataset and its corresponding clonotype c without added noise.Attorney Ref.: 14786-074-228 (I03963-228074) / P40073-W0-1
[0121] In some cases, for each positive sample, three (or another quantity of) negative samples may be sampled, each in a different way. For example, the first negative sample may come from an implementation of contrastive divergence in which a random seed undergoes gradient-based Markov Chain Monte Carlo (MCMC) sampling (e.g., Langevin Markov Chain Monte Carlo (MCMC) sampling) while guided by the score of an underfit energy-based model (EBM). Additional negative samples may be sampled by pairing a noisy protein sequence yposfrom a positive sample with an incorrect clonotype. One example of a negative sample includes the noisy protein sequence yposfrom the positive sample coupled with a randomly-drawn clonotype crand(e.g., (ypos, Gand))- Another example of a negative sample includes the noisy protein sequence yposfrom the positive sample coupled with related (e.g., from the same clonal family) but still incorrect clonotype cfamily(e.g., (ypos, cfamily)
[0122] FIG. 4 depicts a schematic diagram illustrating an overview of an example of context conditioning, in accordance with some example embodiments. In FIG. 4(a), B-cell receptor (BCR) repertoire data from the immune system of an animal exposed to an antigen is shown, with the protein subclasses of one or more lead molecules selected to serve as target contexts for conditioning the generation of novel protein sequences by a protein design computation model (e.g., the protein design computation model 115 of FIG. 1) trained on the repertoire. FIG. 4(b) shows, schematically, the contextual variables defining the target context for conditioning the generation of novel protein sequences by the protein design computation model. The contextual variables shown in FIG. 4(b) include the variable (V) gene segmentjoint (J) gene segment, and third complementarity determining region (CDR3) length of a target clonotype (or clonal family) serving as the target context. In FIG. 4(c), a protein design computation model implemented using a combination of an energy-based model (EBM) and a score-based denoiser isAttorney Ref.: 14786-074-228 (I03963-228074) / P40073-W0-1conditioned by adding the context embedding of the contextual variables shown in FIG. 4(b) to a noisy protein sequence y (generated by the addition of noise to the protein sequence x). In the example shown in FIG. 4(c), the protein design computation model generates noisy output sequences while guided by either the gradientyfe of the energy-based model (EBM) or the score function s^. As further shown in FIG. 4(c), the noisy output sequences may be denoised by a score-based denoiser to generate the corresponding clean output sequences x^. In some cases, the energy-based model (EBM) and the score-based denoiser may be trained separately at least because sampling and denoising may be decoupled, thus allowing the energy-based model (EBM) and the score-based denoiserto be trained separately yet used jointly to generate protein sequences. In some cases, the combination of the energy -based model (EBM) and the denoiser may yield fast, non-autoregressive sequence generation in which an entire protein sequence is generated in a single pass. Contrastingly, autoregressive methods, some of which are discussed in the experimental examples below, requires as many forward passes as the length of the protein sequence being generated. Accordingly, it should be appreciated that various example embodiments of the non-autoregressive approach described herein may be capable of generating protein sequences in less time and consumes less computational resources.
[0123] Experimental Examples
[0124] Various example embodiments of a context conditioned protein design computation model were evaluated using two B-cell receptor (BCR) repertoires datasets. The first B-cell receptor (BCR) repertoire dataset includes 1.5 million heavy-chain sequences from rats immunized with a transmembrane (TM) antigen. The transmembrane (TM) dataset includes 122 single-context modes (variable (V) gene segment) and 2,994 multi-context modes (variable (V) gene segment, joint (J) gene segment, and third complementarity determining region (CDR3)Attorney Ref.: 14786-074-228 (I03963-228074) / P40073-W0-1length). The second B-cell receptor (BCR) repertoire dataset includes 167,537 paired-chain sequences from patients receiving SARs-CoV-2 mRNA vaccines targeting the Spike (S) protein, annotated with heavy-chain and light-chain variable (V) genes. The Spike (S) protein dataset contains 2,797 multi-context modes (heavy chain and light chain variable (V) genes).
[0125] For each B-cell receptor (BCR) repertoire dataset, the protein design computation model was trained with relevant context included. For sampling, context conditions with a range of relative abundances in the training repertoire were selected. Given a target context, 100 samples are generated from each starting seed sequence where each seed sequence is a representative sequence from each variable (V) gene cluster in the training dataset. Doing so ensured no starting seed sequences are out-of-distribution to the trained model. Both the training sequences and final samples are labeled with consistent labels for context conditioning during training and sample labeling.
[0126] The performance of the protein design computation model was evaluated across three metrics. Fidelity is measured through both accuracy of individual and joint context conditions and position-wise Kullback-Leiber (KL) divergence. Position-wise Kullback-Leiber 11 2(KL) divergence is calculated as AA — Position — KL = ~ ?=i DKL(St II?)) , wherein S, and 7) are the amino acid distribution of the samples and the conditioned-matched training subset at sequence position i. Diversity is quantified as the ratio of unique to total samples while novelty is measured by edit distance to the nearest training sequence.
[0127] The performance of the protein design computation model was evaluated on single variable conditioning using a variable (V) gene segment as the target context. To do so, samples were generated by fixing a target variable (V) gene segment and undergoing 10 steps ofAttorney Ref.: 14786-074-228 (I03963-228074) / P40073-W0-1Langevin Markov Chain Monte Carlo (MCMC) prior to denoising the resulting noisy output sequence. Nine variable (V) gene segments from the transmembrane (TM) dataset were selected as contextual variables, representing various levels of abundance and proximity to other variable (V) genes. The generated protein sequences were embedded and visualized, as shown in FIG. 5(a). As further shown in FIGS. 5(b) and (c), the assigned labels of the samples closely matched the target variable (V) gene segment and always matched the target gene family (or the set of related variable (V) genes). Mismatches were likely due to the presence of training sequences from neighboring variable (V) genes within the same gene family and the blurry delineation between variable (V) gene definitions. Interestingly, mismatches still occurred when conditioning on variable (V) genes highly represented in the training dataset (e.g. IGHV2-12), suggesting that subclass definitions may matter more for conditioning than subclass representation in the training dataset.
[0128] The performance of the protein design computation model was also evaluated on multivariable conditioning across multiple categorical and numerical contextual variables simultaneously. For the paired-chain SARS-Cov2 B-cell receptor (BCR) repertoire dataset, the generation of protein sequences by the protein design computation model was conditioned on heavy-chain variable (V) gene segments across ten different context combinations found in the training repertoire. With the transmembrane (TM) dataset, the context included heavy chain variable (V) gene, heavy chain joint (J) gene, and third complementarity determining region (CDR3) length, the last of which is numerical variable with a range between 5 to 24. Fifteen target clonotypes defined by the three aforementioned contextual variables were selected, with three from each of the top 1% most represented clonotypes, the top 25%, the top 50%, and the bottom 25%Attorney Ref.: 14786-074-228 (I03963-228074) / P40073-W0-1as well as three that are not found in the training repertoire to assess the ability of the protein design computation model to generalize to unseen combinations of contextual variables.
[0129] FIGS. 6 and 7 depicts the results of the Spike (S) protein and transmembrane (TM) datasets, respectively. As the guidance level a) increases, the accuracy of samples being in the target context generally increases while position-wise Kullback-Leiber (KL) divergence decreases, implying a greater guidance level a> improves the fidelity of generated sequences and their distributions compared to a conditioned-matched training subset. Interestingly, the uniqueness and edit distance to the nearest training neighbor also increase with increased guidance level to, which indicates that the protein design computation model avoids mode collapse and recapitulating training sequences at high guidance levels. This phenomenon may be due to increasing guidance level to decreasing the contributed score of an unconditional model during sampling. The unconditional model was tasked with denoising samples from all clonotypes without context conditioning, a challenging task that may have required it to rely on a limited set of representative samples from each clonotype. This lowers the diversity of sampled sequences from the unconditional model as well as its contribution, such that accuracy and diversity increase with increased guidance level m.
[0130] The protein design computation model with context conditioning was also evaluated against two baselines: unconditional generation with one unconditional protein design computation model (UncPDCM) trained on protein sequences from a target context and one unconditional generative model (UncPDCM) trained on protein sequences from all context modes. Table 1 shows that the protein design computation model with context conditioning, denoted cPDCM in Table 1, outperforms the baselines by producing higher quality and more diverse samples, even when all training sequences are from the target subclass. The limited sampleAttorney Ref.: 14786-074-228 (I03963-228074) / P40073-W0-1uniqueness of training on all clonotypes further supports why increasing guidance level ) from context conditioning improves sample diversity.
[0131] Table 1: The context conditioned protein design computation model (cPDCM) outperforms unconditional baselines trained in a variety of scenarios on accuracy, uniqueness, and position-wise Kullback-Leiber divergence. The values shown are mean and standard deviations across 10 contexts from the SARS-CoV-2 B-cell receptor (BCR) repertoire dataset.
[0132] Table 2 shows the number of samples within 5 mutations of a known Spike (S) protein binder in CoVAbDab. When conditioned on context enriched in S binders, more samples are within 5 edits of CoVAbDab than conditioned on non-enriched context and thus more likely to be Spike (S) protein binders. This result suggests conditioning on the heavy- and light chain variable (V) genes alone as context can increase the efficiency of discovering novel binders, thus demonstrating the practicality of using various example embodiments of the context conditioned protein design computation model described herein to accelerate protein discovery.
[0133] Table 2: When using binding-enriched context, the protein design computation model generated more protein sequences (out of 36,000) near known SARS-CoV-2 binders in CoVAbDab compared to using a binding-depleted context. In Table 2, (HC only) denotes edit distance computed with heavy chain only and (HC+LC) denotes edit distance computed with both heavy- and light chains.Min edit Binding- Binding- Binding- Binding- rliGTQii enriched context depleted context enriched context depleted context (HC only)! (HC only)! (HC+LC) T (HC+LC)fAttomey Ref.: 14786-074-228 (I03963-228074) / P40073-W0-1CovAbDabbinders<=5 2444 238 147 0 <=3 255 2 6 0 <=1 2 0 0 0
[0134] Sampling trajectories of the context conditioned protein design computation model was analyzed by examining intermediate states during Langevin Markov Chain Monte Carlo (MCMC). For each intermediate state, a sample sequence is denoised to the clean sequence space and annotated with its variable (V) gene. For a given target variable (V) gene, 1000 sequences were sampled per starting seed sequence and the corresponding variable (V) gene annotations were tracked across ten Langevin Markov Chain Monte Carlo (MCMC) steps as well as the frequencies of different variable (V) genes at the final state. These sampling trajectories are shown in FIG. 8.
[0135] As shown in FIG. 11 A, most samples converged to the target variable (V) gene within 1-3 steps, with some exploration of nearby genes within the same gene family (e.g. IGHV2-X when conditioned on IGHV2-63), but not outside a family. Convergence is slower when related variable (V) genes exist in the training repertoire (FIG. 1 IB). In some cases, sequences converged to a closely related but incorrect variable (V) gene (FIG. 11C), likely due to imperfect separation between variable (V) gene definitions. Across the nine contexts assessed using the transmembrane (TM) dataset, 79.7% ( 33.9) of samples converged to the target context, six of which showed 99% convergence.
[0136] FIG. 9(a) depicts a sequence logo plot illustrating the most common types of amino acid residue at each position of sample protein sequences in a training dataset. In FIG. 9(b), a sequence logo plot illustrating the most common types of amino acid residue at each position of sample protein sequences from a subset of sampled protein sequences labeled as having a IGHV2-Attorney Ref.: 14786-074-228 (I03963-228074) / P40073-W0-163 variable (V) gene segment is shown. A sequence logo plot illustrating the most common types of amino acid residue at each position of novel protein sequences generated with context conditioning on the IGHV2-63 variable (V) gene segment is shown in FIG. 9(c).
[0137] FIG. 10(a) depicts a graph illustrating the correlation between the guidance level of context conditioning on the quantity of somatic hypermutations and the accuracy of novel protein sequences generated with context conditioning as quantified by the quantity of somatic hypermutations relative to a target. FIG. 10(b) depicts a graph illustrating the correlation between the guidance level of context conditioning on the quantity of somatic hypermutations and the accuracy of novel protein sequences generated with context conditioning as quantified by rootmean squared error (RMSE) relative to a target.
[0138] FIG. 12 depicts a block diagram illustrating an example of a computing system 1200, in accordance with some example embodiments. Referring to FIGS. 1-12, the computing system 1200 may be used to implement the molecule design engine 110, the analysis engine 120, the client device 130, and / or any components therein.
[0139] As shown in FIG. 12, the computing system 1200 can include a processor 1210, a memory 1220, a storage device 1230, and input / output devices 1240. The processor 1210, the memory 1220, the storage device 1230, and the input / output devices 1240 can be interconnected via a system bus 1250. The processor 1210 is capable of processing instructions for execution within the computing system 1200. Such executed instructions can implement one or more components of, for example, the molecule design engine 110, the analysis engine 120, the client device 130, and / or the like. In some example embodiments, the processor 1210 can be a singlethreaded processor. Alternately, the processor 1210 can be a multi -threaded processor. The processor 1210 is capable of processing instructions stored in the memory 1220 and / or on theAttorney Ref.: 14786-074-228 (I03963-228074) / P40073-W0-1storage device 1230 to display graphical information for a user interface provided via the input / output device 1240.
[0140] The memory 1220 is a computer readable medium such as volatile or non-volatile that stores information within the computing system 1200. The memory 1220 can store data structures representing configuration object databases, for example. The storage device 1230 is capable of providing persistent storage for the computing system 1200. The storage device 1230 can be a floppy disk device, a hard disk device, an optical disk device, or a tape device, or other suitable persistent storage means. The input / output device 1240 provides input / output operations for the computing system 1200. In some example embodiments, the input / output device 1240 includes a keyboard and / or pointing device. In various implementations, the input / output device 1240 includes a display unit for displaying graphical user interfaces.
[0141] According to some example embodiments, the input / output device 1240 can provide input / output operations for a network device. For example, the input / output device 1240 can include Ethernet ports or other networking ports to communicate with one or more wired and / or wireless networks (e.g., a local area network (LAN), a wide area network (WAN), the Internet).
[0142] In some example embodiments, the computing system 1200 can be used to execute various interactive computer software applications that can be used for organization, analysis and / or storage of data in various formats. Alternatively, the computing system 1200 can be used to execute any type of software applications. These applications can be used to perform various functionalities, e.g., planning functionalities (e.g., generating, managing, editing of spreadsheet documents, word processing documents, and / or any other objects, etc.), computing functionalities, communications functionalities, etc. The applications can include various add-inAttorney Ref.: 14786-074-228 (I03963-228074) / P40073-W0-1functionalities or can be standalone computing products and / or functionalities. Upon activation within the applications, the functionalities can be used to generate the user interface provided via the input / output device 1240. The user interface can be generated and presented to a user by the computing system 1200 (e.g., on a computer screen monitor, etc.).
[0143] One or more aspects or features of the subject matter described herein can be realized in digital electronic circuitry, integrated circuitry, specially designed ASICs, field programmable gate arrays (FPGAs) computer hardware, firmware, software, and / or combinations thereof. These various aspects or features can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device. The programmable system or computing system may include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other.
[0144] These computer programs, which can also be referred to as programs, software, software applications, applications, components, or code, include machine instructions for a programmable processor, and can be implemented in a high-level procedural and / or object-oriented programming language, and / or in assembly / machine language. As used herein, the term “machine-readable medium” refers to any computer program product, apparatus and / or device, such as for example magnetic discs, optical disks, memory, and Programmable Logic Devices (PLDs), used to provide machine instructions and / or data to a programmable processor, includingAttorney Ref.: 14786-074-228 (I03963-228074) / P40073-W0-1a machine-readable medium that receives machine instructions as a machine-readable signal. The term “machine-readable signal” refers to any signal used to provide machine instructions and / or data to a programmable processor. The machine-readable medium can store such machine instructions non-transitorily, such as for example as would a non-transient solid-state memory or a magnetic hard drive or any equivalent storage medium. The machine-readable medium can alternatively or additionally store such machine instructions in a transient manner, such as for example, as would a processor cache or other random access memory associated with one or more physical processor cores.
[0145] To provide for interaction with a user, one or more aspects or features of the subject matter described herein can be implemented on a computer having a display device, such as for example a cathode ray tube (CRT), a liquid crystal display (LCD), a light emitting diode (LED) monitor, or an organic light emitting diode (OLED) monitor for displaying information to the user and a keyboard and a pointing device, such as for example a mouse or a trackball, by which the user may provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well. For example, feedback provided to the user can be any form of sensory feedback, such as for example visual feedback, auditory feedback, or tactile feedback; and input from the user may be received in any form, including acoustic, speech, or tactile input. Other possible input devices include touch screens or other touch-sensitive devices such as single or multi-point resistive or capacitive track pads, voice recognition hardware and software, optical scanners, optical pointers, digital image capture devices and associated interpretation software, and the like.
[0146] In the descriptions above and in the claims, phrases such as “at least one of’ or “one or more of’ may occur followed by a conjunctive list of elements or features. The termAttorney Ref.: 14786-074-228 (I03963-228074) / P40073-W0-1“and / or” may also occur in a list of two or more elements or features. Unless otherwise implicitly or explicitly contradicted by the context in which it used, such a phrase is intended to mean any of the listed elements or features individually or any of the recited elements or features in combination with any of the other recited elements or features. For example, the phrases “at least one of A and B;” “one or more of A and B;” and “A and / or B” are each intended to mean “A alone, B alone, or A and B together.” A similar interpretation is also intended for lists including three or more items. For example, the phrases “at least one of A, B, and C;” “one or more of A, B, and C;” and “A, B, and / or C” are each intended to mean “A alone, B alone, C alone, A and B together, A and C together, B and C together, or A and B and C together.” Use of the term “based on,” above and in the claims is intended to mean, “based at least in part on,” such that an unrecited feature or element is also permissible.
[0147] The subject matter described herein can be embodied in systems, apparatus, methods, and / or articles depending on the desired configuration. The implementations set forth in the foregoing description do not represent all implementations consistent with the subject matter described herein. Instead, they are merely some examples consistent with aspects related to the described subject matter. Although a few variations have been described in detail above, other modifications or additions are possible. In particular, further features and / or variations can be provided in addition to those set forth herein. For example, the implementations described above can be directed to various combinations and subcombinations of the disclosed features and / or combinations and subcombinations of several further features disclosed above. In addition, the logic flows depicted in the accompanying figures and / or described herein do not necessarily require the particular order shown, or sequential order, to achieve desirable results. Other implementations may be within the scope of the following claims.
Claims
Attorney Ref.: 14786-074-228 (I03963-228074) / P40073-W0-1CLAIMSWhat is claimed is:
1. A computer-implemented method, comprising:identifying an antigen receptor repertoire having a plurality of clonotypes; identifying a target clonotype within the plurality of clonotypes in the antigen receptor repertoire;training a protein design computation model to approximate a data distribution of protein sequences of antigen receptors expressed by lymphocytes from the target clonotype; and applying the protein design computation model to generate, based at least on an input protein sequence, an output protein sequence corresponding to at least a portion of an antigen receptor expressed by lymphocytes from the target clonotype.
2. The method of claim 1, wherein the protein design computation model includes a clonotype energy-based model (EBM) trained to approximate the data distribution of the protein sequences of antigen receptors expressed by lymphocytes from the target clonotype.
3. The method of claim 2, wherein the clonotype energy -based model (EBM) parameterizes an energy function that outputs a value indicative of whether an antigen receptor corresponding to a protein sequence generated from the input protein sequence is expressed by lymphocytes from the target clonotype.
4. The method of claim 3, wherein the energy function outputs a higher value for a protein sequence drawn from a higher density region of the data distribution populated by protein sequences of antigen receptors expressed by lymphocytes from the target clonotype, and wherein the energy function outputs a lower value for a protein sequence drawn from a lower density region of the data distribution populated by protein sequences of antigen receptors expressed byAttorney Ref.: 14786-074-228 (I03963-228074) / P40073-W0-1lymphocytes outside of the target clonotype.
5. The method of any of claims 3 or 4, wherein the output of the energy function comprises a value corresponding to a density at a location in the data distribution from which a protein sequence generated from the input protein sequence is drawn.
6. The method of any of claims 1 to 5, wherein the protein design computation model includes a score-based model trained to approximate the data distribution of protein sequences of antigen receptors expressed by lymphocytes from the target clonotype.
7. The method of claim 6, wherein the output of the score function comprises a value corresponding to a local density change at a location in the data distribution from which a protein sequence generated from the input protein sequence is drawn.
8. The method of any of claims 6 or 7, wherein the score-based model comprises a conditional score-based model trained with conditioning on the target clonotype and an unconditional score-based model trained without conditioning on the target clonotype.
9. The method of claim 8, wherein the score-based model comprises a linear combination of the conditional score-based model and the unconditional score-based model, and wherein a respective output of the conditional score-based model and the unconditional scorebased model are weighted by a weight corresponding to a guidance level from the target clonotype.
10. The method of any of claims 1 to 9, wherein the output protein sequence is generated by modifying the input protein sequence over multiple iterations of gradient-based Markov Chain Monte Carlo (MCMC) sampling.
11. The method of any of claims 1 to 10, wherein the output protein sequence is generated by at leastAttorney Ref.: 14786-074-228 (I03963-228074) / P40073-W0-1applying the protein design computation model to generate a first modified protein sequence having a first modification to the input protein sequence,applying the protein design computation model to generate a second modified protein sequence having a second modification to the input protein sequence,determining that an antigen receptor corresponding to the first modified protein sequence is more likely to be expressed by lymphocytes from the target clonotype than an antigen receptor corresponding to the second modified protein sequence, andapplying the protein design computation model to further modify the first modified protein sequence instead of the second modified protein sequence.
12. The method of claim 11, wherein each of the first modification and the second modification include inserting, deleting, and / or changing a type of one or more amino acid residues in the input protein sequence.
13. The method of any of claims 11 or 12, further comprising:generating the output protein sequence to correspond to the first modified protein sequence upon satisfying one or more criteria.
14. The method of claim 13, wherein the one or more criteria include at least one of (i) the input protein sequence having undergone a threshold quantity of gradient-based Markov Chain Monte Carlo (MCMC) sampling iterations to generate the first modified protein sequence, and (ii) the first modified protein sequence corresponding to at least the portion of the antigen receptor expressed by lymphocytes from the target clonotype.
15. The method of any of claims 13 to 14, wherein the data distribution comprises a noisy data distribution populated by noisy protein sequences of antigen receptors from the antigen receptor repertoire.Attorney Ref.: 14786-074-228 (I03963-228074) / P40073-W0-116. The method of claim 15, wherein the generating of the output protein sequence includes denoising the first modified protein sequence.
17. The method of any of claims 1 to 16, further comprising:training the protein design computation model to approximate a data distribution of protein sequences of antigen receptors expressed by mature lymphocytes that have undergone adequate affinity maturation.
18. The method of claim 17, wherein the protein design computation model includes a maturity energy -based model (EBM) trained to approximate the data distribution of protein sequences of antigen receptors expressed by mature lymphocytes.
19. The method of claim 18, wherein the maturity energy -based model (EBM) parameterizes an energy function that outputs a value indicative of whether an antigen receptor corresponding to a protein sequence generated from the input protein sequence is expressed by mature lymphocytes.
20. The method of claim 19, wherein the energy function outputs a higher value for a protein sequence drawn from a higher density region of the data distribution populated by protein sequences of antigen receptors expressed by mature lymphocytes, and wherein the energy function outputs a lower value for a protein sequence drawn from a lower density region of the data distribution populated by protein sequences of antigen receptors expressed by immature lymphocytes.
21. The method of any of claims 18 to 20, wherein the generating of the output protein sequence is guided by a composition of an energy function parameterized by the clonotype energy -based model (EBM) and the energy function parameterized by the maturity energy-based model (EBM) such that the output protein sequence corresponds to an antigenAttorney Ref.: 14786-074-228 (I03963-228074) / P40073-W0-1receptor expressed by a mature lymphocyte from the target clonotype.
22. The method of claim 21, wherein the composition comprises a product of experts or a weighted sum.
23. The method of any of claims 17 to 22, wherein a maturity of a lymphocyte is defined by one or more of (i) a quantity of somatic hypermutations (SHM) present in the lymphocyte, and (ii) a depth in a lymphocyte phylogenetic tree occupied by the lymphocyte.
24. The method of any of claims 1 to 23, wherein the target clonotype is defined by (i) a variable (V) gene segment, (ii) a joining (J) gene segment, and (ii) a third complementarity determining region (CDR3) length.
25. The method of any of claims 1 to 24, wherein the target clonotype comprises a clonotype that is enriched in the antigen receptor repertoire.
26. The method of any of claims 1 to 25, wherein the protein design computation model is trained to approximate the data distribution based on a plurality of positive samples, and wherein each positive sample includes a protein sequence of an antigen receptor paired with a correct clonotype.
27. The method of any of claims 1 to 26, wherein the protein design computation model is trained to approximate the data distribution based on a plurality of negative samples, and wherein each negative sample includes a protein sequence of an antigen receptor paired with an incorrect clonotype.
28. The method of claim 27, wherein the incorrect clonotype comprises a random clonotype or a different clonotype from a same clonal family as a correct clonotype.
29. A system, comprising:at least one data processor; andAttorney Ref.: 14786-074-228 (I03963-228074) / P40073-W0-1at least one memory storing instructions, which when executed by the at least one data processor, result in operations comprising the method of any of claims 1 to 28.
30. A non-transitory computer readable medium storing instructions, which when executed by at least one data processor, result in operations comprising the method of any of claims 1 to 28.
31. A computer-implemented method, comprising:identifying an antigen receptor repertoire having a plurality of different contexts; identifying a target context within the plurality of different contexts present in the antigen receptor repertoire;training a protein design computation model to approximate a data distribution of protein sequences of antigen receptors exhibiting the target context; andapplying the protein design computation model to generate, based at least on an input protein sequence, an output protein sequence corresponding to at least a portion of an antigen receptor having the target context.
32. The method of claim 31, wherein the target context is defined by one or more contextual variables.
33. The method of claim 32, wherein the one or more contextual variables include one or more of a categorical variable and a numerical variable.
34. The method of any of claims 31 to 33, wherein the target context comprises a target clonotype of lymphocytes expressing the antigen receptor.
35. The method of claim 34, wherein the target clonotype is defined by one or more of a variable (V) gene segment, a joint (J) gene segment, and a third complementarity determining region (CDR3) length.Attorney Ref.: 14786-074-228 (I03963-228074) / P40073-W0-136. The method of any of claims 31 to 35, wherein the target context comprise a maturity of lymphocytes expressing the antigen receptor.
37. The method of claim 36, wherein the maturity of a lymphocyte is defined by one or more of a quantity of somatic hypermutations (SHMs) present in an antigen receptor coding sequence of the lymphocyte relative to its ancestral lymphocyte, and a depth of the lymphocyte in a lymphocyte phylogenetic tree depicting the evolutionary relationships within the antigen receptor repertoire.
38. The method of any of claims 31 to 37, wherein the target context includes a target clonotype and a maturity of lymphocytes expressing the antigen receptor.
39. The method of any of claims 31 to 38, wherein the protein design computation model includes an energy-based model (EBM) trained to approximate the data distribution of protein sequences of antigen receptors exhibiting the target context.
40. The method of claim 39, wherein the energy-based model (EBM) parameterizes an energy function that outputs a value indicative of whether an antigen receptor corresponding to a protein sequence generated from the input sequence exhibits the target context.
41. The method of claim 40, wherein the energy function outputs a higher value for a protein sequence drawn from a higher density region of the data distribution populated by protein sequences of antigen receptors exhibiting the target context, and wherein the energy function outputs a lower value for a protein sequence drawn from a lower density region of the data distribution populated by protein sequences of antigen receptors without the target context.
42. The method of any of claims 40 or 41, wherein the output of the energy function comprises a value corresponding to a density at a location in the data distribution from which a protein sequence generated from the input protein sequence is drawn.Attorney Ref.: 14786-074-228 (I03963-228074) / P40073-W0-143. The method of any of claims 31 to 42, wherein the protein design computation model includes a score-based function trained to approximate data the distribution of protein sequences of antigen receptors exhibiting the target context.
44. The method of claim 43, wherein the output of the score function comprises a value corresponding to a local density change at a location in the data distribution from which a protein sequence generated from the input protein sequence is drawn.
45. The method of any of claims 43 or 44, wherein the score-based model comprises a conditional score-based model trained with conditioning on the target context and an unconditional score-based model trained without conditioning on the target context.
46. The method of claim 45, wherein the score-based model comprises a linear combination of the conditional score-based model and the unconditional score-based model, and wherein a respective output of the conditional score-based model and the unconditional scorebased model are weighted by a weight corresponding to a guidance level from the target context.
47. The method of any of claims 31 to 46, wherein the output protein sequence is generated by modifying the input protein sequence over multiple iterations of gradient-based Markov Chain Monte Carlo (MCMC) sampling.
48. The method of any of claims 31 to 47, wherein the output protein sequence is generated by at leastapplying the protein design computation model to generate a first modified protein sequence having a first modification to the input protein sequence,applying the protein design computation model to generate a second modified protein sequence having a second modification to the input protein sequence,determining that an antigen receptor corresponding to the first modified protein sequenceAttorney Ref.: 14786-074-228 (I03963-228074) / P40073-W0-1is more likely to exhibit the target context than an antigen receptor corresponding to the second modified protein sequence, andapplying the protein design computation model to further modify the first modified protein sequence instead of the second modified protein sequence.
49. The method of claim 48, wherein each of the first modification and the second modification include inserting, deleting, and / or changing a type of one or more amino acid residues in the input protein sequence.
50. The method of any of claims 48 or 49, further comprising:generating the output protein sequence to correspond to the first modified protein sequence upon satisfying one or more criteria.
51. The method of claim 50, wherein the one or more criteria include at least one of (i) the input protein sequence having undergone a threshold quantity of gradient-based Markov Chain Monte Carlo (MCMC) sampling iterations to generate the first modified protein sequence, and (ii) the first modified protein sequence corresponding to at least the portion of the antigen receptor exhibiting the target context.
52. The method of any of claims 50 or 51, wherein the data distribution comprises a noisy data distribution populated by noisy protein sequences of antigen receptors from the antigen receptor repertoire.
53. The method of any of claims 50 to 52, wherein the generating of the output protein sequence includes denoising the first modified protein sequence.
54. The method of any of claims 31 to 53, wherein the protein design computation model comprises a composition of an energy function parameterized by a clonotype energybased model (EBM) and an energy function parameterized by a maturity energy -based modelAttorney Ref.: 14786-074-228 (I03963-228074) / P40073-W0-1(EBM).
55. The method of claim 54, wherein the clonotype energy-based model (EBM) is trained to approximate a data distribution of protein sequences of antigen receptors expressed by lymphocytes from a target clonotype comprising the target context, and wherein the maturity energy-based model (EBM) is trained to approximate a data distribution of protein sequences of antigen receptors expressed by mature lymphocytes further comprising the target context.
56. The method of any of claims 54 or 55, wherein the composition comprises a product of experts or a weighted sum.
57. The method of any of claims 54 to 56, wherein the output protein sequence is generated to correspond to an antigen receptor expressed by a mature lymphocyte from the target clonotype58. The method of any of claims 31 to 57, wherein the protein design computation model is trained to approximate the data distribution based on a plurality of positive samples, and wherein each positive sample includes a protein sequence of an antigen receptor paired with a correct context.
59. The method of any of claims 31 to 58, wherein the protein design computation model is trained to approximate the data distribution based on a plurality of negative samples, and wherein each negative sample includes a protein sequence of an antigen receptor paired with an incorrect context.
60. The method of any of claims 31 to 59, wherein the antigen receptor repertoire includes a plurality of B-cell receptors and / or T-cell receptors generated by one or more animals of a same or different species upon exposure to an antigen.
61. A system, comprising:Attorney Ref.: 14786-074-228 (I03963-228074) / P40073-W0-1at least one data processor; andat least one memory storing instructions, which when executed by the at least one data processor, result in operations comprising the method of any of claim 31 to 60.
62. A non-transitory computer readable medium storing instructions, which when executed by at least one data processor, result in operations comprising the method of any of claims 31 to 60.