Machine learning-enabled analysis of protein sequences and generation of protein sequences therewith

WO2026117531A1PCT designated stage Publication Date: 2026-06-04GENENTECH INC +2

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
GENENTECH INC
Filing Date
2025-11-25
Publication Date
2026-06-04

Smart Images

  • Figure US2025056972_04062026_PF_FP_ABST
    Figure US2025056972_04062026_PF_FP_ABST
Patent Text Reader

Abstract

A method may include determining a protein sequence comprising a plurality of amino acid residues. An embedding computation model may be applied to generate an embedding of the protein sequence. The embedding computation model may have been trained, through contrastive learning, to generate similar embeddings for protein sequences exhibiting a threshold similarity in one or more specific regions and dissimilar embeddings for protein sequences failing to exhibit the threshold similarity in the one or more specific regions. The embedding generated by the embedding computation model may prioritize similarities in the one or more specific regions (e.g., CDR3) over similarities in other regions (e.g., framework region). One or more related groups of protein sequences may be identified based on the embedding of the protein sequence. A visual representation may be generated based on the one or more groups of related protein sequences. Related systems and computer program products are also provided.
Need to check novelty before this filing date? Find Prior Art

Description

Attorney Ref.: 14786-058-228 (103963-228058) / P39076-WO-1MACHINE LEARNING-ENABLED ANALYSIS OF PROTEIN SEQUENCES AND GENERATION OF PROTEIN SEQUENCES THEREWITH CROSS REFERENCE TO RELATED APPLICATION

[0001] This application claims priority to U. S. Provisional Application No. 63 / 725,384, entitled “MACHINE LEARNING-ENABLED ANALYSIS OF PROTEIN SEQUENCES” and filed on November 26, 2024, the disclosure of which is incorporated herein by reference in its entirety.TECHNICAL FIELD

[0002] The subject matter described herein relates generally to the protein engineering and more specifically to techniques for protein analysis using high-dimensional embeddings of protein sequences generated by an embedding computation model and the generation of protein sequences based on such analysis.INTRODUCTION

[0002] Proteins are genetically encoded macromolecules with tremendous diversity in size and chemical composition. By regulating biological systems, proteins facilitate many essential cellular functions including, for example, enzymatic reactions, molecular transport, regulation and execution of a number of biological pathways, cell growth, proliferation, nutrient uptake, morphology, motility, intercellular communication, and / or the like. A protein structure may include one or more polypeptides, each of which including a sequence of amino acid residues linked together by peptide bonds (e.g., covalent peptide bonds). There are twenty canonical amino acid residues which, unlike non-canonical amino acid residues, are encoded directly by the genetic code. Each canonical amino acid residue includes the same backbone atoms (e.g., an amino groupNAI-5007186271V1 1Attorney Ref.: 14786-058-228 (103963-228058) / P39076-WO-1(NH2), an alpha carbon (C_a), and a carboxylic group (COOH)) coupled with a different combination of sidechain atoms (or R groups).

[0003] The primary structure of a protein molecule refers to the sequence of amino acid residues in each of the polypeptide chains forming the protein structure. The backbone atoms in adjacent amino acid residues that participate in the peptide bonds (e.g., covalent peptide bonds) therebetween form a repeating sequence of atoms known as the polypeptide backbone (or backbone) of the protein molecule. The secondary structure of the protein molecule refers to the local folded structures (e.g., a helixes, pleated sheet, and / or the like) that form within an individual polypeptide chain due to interactions between the backbone atoms (e.g., amino hydrogen atoms, carboxyl oxygen atoms, and / or the like). Further interactions (e.g., non-covalent bonds such as hydrogen bonding, ionic bonding, dipole-dipole interactions, and van der Waals forces) between the sidechains (or R-groups) of the amino acid residues in the protein molecule may cause folding within the individual polypeptide chains, thus forming the tertiary structure of the protein molecule. The tertiary structure of the protein molecule is also known as the conformation or the three-dimensional structure of the protein molecule. In protein molecules having multiple polypeptide chains, the protein molecule may also exhibit a quaternary structure, which is formed when the polypeptide chains are packed and held together by hydrogen bonds and van der Waals forces (e.g., between nonpolar sidechains).

[0004] The functions of a protein molecule may be contingent upon the sequence of amino acid residues in the polypeptide chains forming the protein molecule as well as the three-dimensional structure adopted by the polypeptide chains. For example, the primary structure of the protein molecule may determine the three-dimensional structure assumed by the protein molecule through the folding of the constituent polypeptide chains. In some cases, the bindingNAI-5007186271V1 2Attorney Ref.: 14786-058-228 (103963-228058) / P39076-WO-1affinity of the protein molecule towards a target molecule, such as a viral or tumor antigen, may depend on whether the polypeptide chains in the protein molecule are able to assume a three-dimensional structure that complements the three-dimensional structure of the target molecule and is sufficiently stable to allow a binding interaction between the two molecules. As such, one notable objective of computational protein design is to construct one or more protein sequences (e.g., antibodies and / or the like) that exhibit certain desirable properties. For instance, in the case of large molecule drug discovery (LMDD), computational protein design may seek to identify therapeutically viable protein sequences (e.g., antibodies and / or the like) with a variety of desirable properties such as expression, binding affinity towards a target molecule, binding specificity, stability, non-immunogenicity, human-ness, absence of self-association (or non-aggregation), lack of chemical liabilities (e.g., aspartate isomerization, oxidation, deamidation), and / or the like. SUMMARY

[0003] Systems, methods, and articles of manufacture, including computer program products, are provided for protein analysis in which high-dimensional embeddings of protein sequences generated by an embedding computation model are projected into a lower-dimensional space suitable for visualization (e.g., a two-dimensional space).

[0004] In one aspect, there is provide a system for machine learning enabled protein analysis and generation. The system may include at least one data processor and at least one memory. The at least one memory may store instructions that result in operations when executed by the at least one data processor. The operations may include: determining a selected protein sequence comprising a plurality of amino acid residues; determining two or more protein sequences that exhibit a threshold similarity in a specific region of each protein sequence; applying an embedding computation model to generate an embedding of the selected protein sequence,NAI-5007186271V1 3Attorney Ref.: 14786-058-228 (103963-228058) / P39076-WO-1where the embedding computation model has been trained to generate embeddings for the two or more determined protein sequences; identifying, based at least on the embedding of the selected protein sequence, one or more related groups of protein sequences; and generating, based at least on the one or more identified groups of related protein sequences, a visual representation.

[0005] In another aspect, there is provided a method for machine learning enabled protein analysis and generation. The method may include: determining a selected protein sequence comprising a plurality of amino acid residues; determining two or more protein sequences that exhibit a threshold similarity in a specific region of each protein sequence; applying an embedding computation model to generate an embedding of the selected protein sequence, where the embedding computation model has been trained to generate embeddings for the two or more determined protein sequences; identifying, based at least on the embedding of the selected protein sequence, one or more related groups of protein sequences; and generating, based at least on the one or more identified groups of related protein sequences, a visual representation.

[0006] In another aspect, there is provided a computer program product for machine learning enabled protein analysis and generation. The computer program product may include a non-transitory computer readable medium storing instructions that result in operations when executed by at least one data processor. The operations may include: determining a selected protein sequence comprising a plurality of amino acid residues; determining two or more protein sequences that exhibit a threshold similarity in a specific region of each protein sequence; applying an embedding computation model to generate an embedding of the selected protein sequence, where the embedding computation model has been trained to generate embeddings for the two or more determined protein sequences; identifying, based at least on the embedding of the selectedNAI-5007186271V1 4Attorney Ref.: 14786-058-228 (103963-228058) / P39076-WO-1protein sequence, one or more related groups of protein sequences; and generating, based at least on the one or more identified groups of related protein sequences, a visual representation.

[0007] In some variations, one or more features disclosed herein including the following features can optionally be included in any feasible combination.

[0008] In some variations, the embedding computation model has been trained such that an edit distance between the two or more similar embeddings satisfy one or more threshold criteria.

[0009] In some variations, two or more dissimilar protein sequences are determined. The dissimilar protein sequences comprising sequences that fail to exhibit the threshold similarity in the specific region of each protein sequence. The embedding computation model is trained to generate, for the two or more dissimilar protein sequences, two or more embeddings whose edit distance satisfy a different threshold criteria.

[0010] In some variations, two or more protein sequences that exhibit a threshold similarity across an entire length of each protein sequence are determined. The embedding computation model is trained to generate, for the two or more protein sequences with the threshold similarity across the entire length of each protein sequence, similar embeddings whose edit distance satisfy one or more threshold criteria.

[0011] In some variations, two or more protein sequences that fail to exhibit the threshold similarity across an entire length of each protein sequence are determined. The embedding computation model is trained to generate, for the two or more protein sequences without the threshold similarity across the entire length of each protein sequence, two or more dissimilar embeddings whose edit distance satisfy a different threshold criteria.NAI-5007186271V1 5Attorney Ref.: 14786-058-228 (103963-228058) / P39076-WO-1

[0012] In some variations, an additional embedding that generalizes, to a corresponding amino acid group, an identity of each amino acid residue in the specific region is generated for the selected protein sequence.

[0013] In some variations, a hybrid embedding comprising a combination of the embedding and the additional embedding is generated. The one or more related groups of protein sequences are identified based at least on the hybrid embedding.

[0014] In some variations, the combination comprises a union of the embedding and the additional embedding in which the embedding is weighted by one weight and the additional embedding is weighted by a different weight.

[0015] In some variations, the additional embedding is generated by at least assigning, to an amino acid residue group, each amino acid residue in the specific region of the selected protein sequence, and generating the additional embedding to include, for each amino acid residue in the specific region of the selected protein sequence, the amino acid residue group assigned to the amino acid residue.

[0016] In some variations, the amino acid residue group is one of a non-polar group, a polar group, an aromatic group, a positively charged (or basic) group, a negatively-charged (or acidic) group, or a special group.

[0017] In some variations, the one or more related groups of protein sequences are identified by at least determining, based at least on the embedding of the selected protein sequence, a cluster label assigning the selected protein sequence to a group of related protein sequences.

[0018] In some variations, the identifying of one or more related groups of protein sequences includes applying a clustering technique to a plurality of embeddings of a plurality of protein sequences including the embedding of the protein sequences.NAI-5007186271V1 6Attorney Ref.: 14786-058-228 (103963-228058) / P39076-WO-1

[0019] In some variations, the clustering technique comprises one of a hierarchal clustering technique, a centroid-based clustering technique, a density-based clustering technique, or a distribution-based clustering technique.

[0020] In some variations, the plurality of protein sequences comprise a repertoire of antibodies generated in response to exposure to a pathogen.

[0021] In some variations, the determining the selected protein sequence includes translating, into the plurality of amino acid residues, one or more corresponding coding sequences comprising raw sequencing data.

[0022] In some variations, the visual representation is generated to depict, for each related group of protein sequences, a corresponding cluster.

[0023] In some variations, the visual representation is generated by at least projecting the one or more groups of protein sequences into a two-dimensional or a three-dimensional space.

[0024] In some variations, the visual representation is generated to include a visual indicator representative of the selected protein sequence. The visual indicator is rendered in a color, shape, and / or size representative of one or more properties of the selected protein sequence.

[0025] In some variations, the threshold similarity is quantified by an edit distance.

[0026] In some variations, the specific region includes a complementarity determining region (CDR).

[0027] In some variations, the embedding computation model has been trained to prioritize similarities in the complementarity determining region (CDR) of the selected protein sequence over similarities in a framework region of the selected protein sequence when generating the embedding of the selected protein sequence.NAI-5007186271V1 7Attorney Ref.: 14786-058-228 (103963-228058) / P39076-WO-1

[0028] In some variations, the specific region includes a third complimentarity determining region (CDR3).

[0029] In some variations, the embedding computation model comprises an encoder of a language model having a transformer architecture.

[0030] In some variations, the embedding computation model has been trained to generate embeddings occupying a two-dimensional Euclidean space.

[0031] In some variations, the embedding computation model has been trained to generate embeddings occupying a hyperbolic space.

[0032] In some variations, the embedding computation model has been trained to generate the embedding of the selected protein sequence to lie below an embedding of a parent protein sequence of the selected protein sequence in a latent space occupied by a plurality of embeddings generated by the embedding computation model.

[0033] In some variations, the embedding computation model has been trained to generate the embedding of the selected protein sequence to lie above an embedding of a child protein sequence of the selected protein sequence in latent space.

[0034] In some variations, the embedding computation model has been trained with contrastive learning to generate the similar embeddings for the two or more similar protein sequences exhibiting the threshold similarity in the specific region and dissimilar embeddings for two or more dissimilar protein sequences without the threshold similarity in the specific region.

[0035] In some variations, the embedding computation model has been further finetuned to reduce an angle between embeddings of two or more related protein sequences exhibiting a parent-child hierarchical relationship and increase an angle between embeddings of two or more unrelated protein sequences absent the parent-child hierarchical relationship.NAI-5007186271V1 8Attorney Ref.: 14786-058-228 (103963-228058) / P39076-WO-1

[0036] In some variations, a generative context is determined based at least on an identified group of related protein sequences. A generative model is applied to generate, based at least on the generative context, one or more novel protein sequences.

[0037] In some variations, the generative model has been trained on a fill-in-the-middle (FIM) objective to generate the one or more novel protein sequences while leveraging a plurality of protein sequences forming the generative context.

[0038] In some variations, the generative model has been finetuned with direct preference optimization (DPO) to generate the one or more novel protein sequences by at least further evolving a subset of protein sequences occupying a region of interest in the generative context.

[0039] In some variations, the subset of protein sequences include children protein sequences exhibiting a greater distance to a germline and / or a greater edit distance to a naive repertoire than one or more parent protein sequences outside of the region of interest in the generative context.

[0040] In one aspect, there is provide a system for machine learning enabled protein analysis and generation. The system may include at least one data processor and at least one memory. The at least one memory may store instructions that result in operations when executed by the at least one data processor. The operations may include: identifying a pair of protein sequences as a positive pair based at least on the pair of protein sequences exhibiting a threshold similarity between a specific region of each protein sequence of the pair of protein sequences; identifying the pair of protein sequences as a negative pair based at least on the pair of protein sequences failing to exhibit the threshold similarity in the specific region of each protein sequence of the pair of protein sequences; training an embedding computation model by at least adjustingNAI-5007186271V1 9Attorney Ref.: 14786-058-228 (103963-228058) / P39076-WO-1one or more parameters of the embedding computation model to generate embeddings for each positive pair and embeddings for each negative pair; and applying the trained embedding computation model to generate an embedding for a protein sequence.

[0041] In another aspect, there is provided a method for machine learning enabled protein analysis and generation. The method may include: identifying a pair of protein sequences as a positive pair based at least on the pair of protein sequences exhibiting a threshold similarity between a specific region of each protein sequence of the pair of protein sequences; identifying the pair of protein sequences as a negative pair based at least on the pair of protein sequences failing to exhibit the threshold similarity in the specific region of each protein sequence of the pair of protein sequences; training an embedding computation model by at least adjusting one or more parameters of the embedding computation model to generate embeddings for each positive pair and embeddings for each negative pair; and applying the trained embedding computation model to generate an embedding for a protein sequence.

[0042] In another aspect, there is provided a computer program product for machine learning enabled protein analysis and generation. The computer program product may include a non-transitory computer readable medium storing instructions that result in operations when executed by at least one data processor. The operations may include: identifying a pair of protein sequences as a positive pair based at least on the pair of protein sequences exhibiting a threshold similarity between a specific region of each protein sequence of the pair of protein sequences; identifying the pair of protein sequences as a negative pair based at least on the pair of protein sequences failing to exhibit the threshold similarity in the specific region of each protein sequence of the pair of protein sequences; training an embedding computation model by at least adjusting one or more parameters of the embedding computation model to generate embeddings for eachNAI-5007186271V1 10Attorney Ref.: 14786-058-228 (103963-228058) / P39076-WO-1positive pair and embeddings for each negative pair; and applying the trained embedding computation model to generate an embedding for a protein sequence.

[0043] In some variations, one or more features disclosed herein including the following features can optionally be included in any feasible combination.

[0044] In some variations, the pair of protein sequence is identified as the positive pair further based at least on the pair of protein sequences exhibiting the threshold similarity across an entire length of each protein sequence.

[0045] In some variations, the pair of protein sequences are identified as the negative pair further based at least on the pair of protein sequences failing to exhibit the threshold similarity across the entire length of each protein sequence.

[0046] In some variations, the one or more parameters of the embedding computation model are adjusted to reduce an edit distance between embeddings of each positive pair.

[0047] In some variations, the one or more parameters of the embedding computation model are further adjusted to increase the edit distance between embeddings of each negative pair.

[0048] In some variations, the threshold similarity is quantified by an edit distance.

[0049] In some variations, the embedding computation model comprises an encoder of a language model having a transformer architecture.

[0050] In some variations, the specific region includes a complementarity determining region (CDR).

[0051] In some variations, the embedding computation model is trained to prioritize similarities in the complementarity determining regions (CDRs) of each pair of protein sequences over similarities in a framework regions each pair of protein sequences.NAI-5007186271V1 11Attorney Ref.: 14786-058-228 (103963-228058) / P39076-WO-1

[0052] In some variations, the specific region includes a third complementarity determining region (CDR3).

[0053] In some variations, a plurality of protein sequences are projected into a two-dimensional or a three-dimensional space. One or more protein sequences within a threshold distance of the protein sequence are identified, for a protein sequence of the plurality of protein sequences, as one or more neighboring protein sequences. The protein sequence and a neighboring protein sequence are identified as the positive pair based at least on the protein sequence and the neighboring protein sequence exhibiting the threshold degree of similarity in the specific region of each protein sequence.

[0054] In some variations, the protein sequence is excluded from being a part of any positive pairs and negative pairs based at least on the protein sequence failing to have a threshold quantity of neighboring protein sequences.

[0055] In some variations, the pair of protein sequences is identified as the positive pair further based at least on the pair of protein sequences including a protein sequence and a child protein sequence of the protein sequence. The embedding computation model is finetuned by further adjusting the one or more parameters of the embedding computation model to generate an embedding of the child protein sequence to lie beneath an embedding of the protein sequence in a latent space populated by a plurality of embeddings generated by the embedding computation model.

[0056] In some variations, the pair of protein sequences is identified as the positive pair further based at least on the pair of protein sequences including a protein sequence and a parent protein sequence from which the protein sequence descends. The embedding computation model is finetuned by further adjusting the one or more parameters of the embedding computation modelNAI-5007186271V1 12Attorney Ref.: 14786-058-228 (103963-228058) / P39076-WO-1to generate an embedding of the protein sequence to lie beneath an embedding of the parent protein sequence in a latent space populated by a plurality of embeddings generated by the embedding computation model.

[0057] In some variations, the pair of protein sequences is identified as the negative pair further based at least on the pair of protein sequences including two unrelated protein sequences. The embedding computation model is finetuned by further adjusting the one or more parameters of the embedding computation model to generate an embedding of one unrelated protein sequence to lie away from an embedding of the other unrelated protein sequence in a latent space populated by a plurality of embeddings generated by the embedding computation model.

[0058] In some variations, the one or more parameters of the embedding computation model are adjusted such that the embedding computation model generates similar embeddings for each positive pair and dissimilar embeddings for each negative pair.

[0059] In some variations, the embedding computation model is trained to generate embeddings occupying a two-dimensional Euclidean space.

[0060] In some variations, the embedding computation model is trained to generate embeddings occupying a hyperbolic space.

[0061] Implementations of the current subject matter can include, but are not limited to, methods consistent with the descriptions provided herein as well as articles that comprise a tangibly embodied machine-readable medium operable to cause one or more machines (e.g., computers, etc.) to result in operations implementing one or more of the described features. Similarly, computer systems are also described that may include one or more processors and one or more memories coupled to the one or more processors. A memory, which can include a non-transitory computer-readable or machine-readable storage medium, may include, encode, store, orNAI-5007186271V1 13Attorney Ref.: 14786-058-228 (103963-228058) / P39076-WO-1the like one or more programs that cause one or more processors to perform one or more of the operations described herein. Computer implemented methods consistent with one or more implementations of the current subject matter can be implemented by one or more data processors residing in a single computing system or multiple computing systems. Such multiple computing systems can be connected and can exchange data and / or commands or other instructions or the like via one or more connections, including, for example, to a connection over a network (e.g. the Internet, a wireless wide area network, a local area network, a wide area network, a wired network, or the like), via a direct connection between one or more of the multiple computing systems, etc.

[0062] The details of one or more variations of the subject matter described herein are set forth in the accompanying drawings and the description below. Other features and advantages of the subject matter described herein will be apparent from the description and drawings, and from the claims. While certain features of the currently disclosed subject matter are described for illustrative purposes in relation to therapeutic proteins (e.g., monoclonal antibodies (mABs), peptide hormones, growth factors, plasma proteins, enzymes, hemolytic factors, and / or the like), it should be readily understood that such features are not intended to be limiting. The claims that follow this disclosure are intended to define the scope of the protected subject matter.DESCRIPTION OF DRAWINGS

[0063] The accompanying drawings, which are incorporated in and constitute a part of this specification, show certain aspects of the subject matter disclosed herein and, together with the description, help explain some of the principles associated with the disclosed implementations. In the drawings,NAI-5007186271V1 14Attorney Ref.: 14786-058-228 (103963-228058) / P39076-WO-1

[0064] FIG. 1 depicts a system diagram illustrating an example of a protein analysis system, in accordance with some example embodiments;

[0065] FIG. 2A depicts a flowchart illustrating an example of a process for machine learning enabled protein analysis, in accordance with some example embodiments;

[0066] FIG. 2B depicts a flowchart illustrating an example of a process for training an embedding computation model to generate an embedding of a protein sequence that emphasize similarities in certain regions of the protein sequence, in accordance with some example embodiments;

[0067] FIG. 3A depicts a flowchart illustrating an example of a process for machine learning enabled analysis of protein sequences, in accordance with some example embodiments;

[0068] FIG. 3B depicts a flowchart illustrating another example of a process for machine learning enabled analysis of protein sequences, in accordance with some example embodiments;

[0069] FIG. 4 depicts a flowchart illustrating an example of a process for context-aware generation of novel protein sequences, in accordance with some example embodiments;

[0070] FIG. 5 depicts a schematic diagram illustrating an example of a workflow for machine learning enabled analysis of protein sequences, in accordance with some example embodiments;

[0071] FIG. 6 depicts a schematic diagram illustrating an example of a workflow for generating embeddings occupying a two-dimensional Euclidean space, in accordance with some example embodiments;NAI-5007186271V1 15Attorney Ref.: 14786-058-228 (103963-228058) / P39076-WO-1

[0072] FIG. 7A depicts a schematic diagram illustrating an example of a workflow for generating embeddings occupying a hyperbolic space, in accordance with some example embodiments;

[0073] FIG. 7B depicts a schematic diagram illustrating another example of a workflow for generating embeddings occupying a hyperbolic space, in accordance with some example embodiments;

[0074] FIG. 7C depicts two-dimensional visualizations of the embeddings of antibodies that illustrate the biological significance captured by the embeddings, in accordance with some example embodiments;

[0075] FIG. 8 depicts a graph illustrating an example of hierarchical loss, in accordance with some example embodiments;

[0076] FIG. 9A depicts the grouping and hierarchical structure that is present in the embeddings of protein sequences generated by an embedding computation model that has been finetuned (or post trained) with hierarchical loss, in accordance with some example embodiments;

[0077] FIG. 9B depicts the grouping and hierarchical structure that is present in the embeddings of protein sequences generated by an embedding computation model that has not been finetuned (or post trained) with hierarchical loss, in accordance with some example embodiments;

[0078] FIG. 10A depicts a schematic diagram illustrating the output of a generative model without finetuning for direct preference optimization (DPO), in accordance with some example embodiments;NAI-5007186271V1 16Attorney Ref.: 14786-058-228 (103963-228058) / P39076-WO-1

[0079] FIG. 10B depicts a schematic diagram illustrating the output of a generative model with finetuning for direct preference optimization (DPO), in accordance with some example embodiments;

[0080] FIG. 11 depicts a schematic diagram illustrating an example of a generalization scheme for generating an embedding of a protein sequence that generalizes one or more specific regions of the protein sequence, in accordance with some example embodiments;

[0081] FIG. 12A depicts a screenshot illustrating an example of a user interface, in accordance with some example embodiments;

[0082] FIG. 12B depicts a screenshot illustrating another example of a user interface, in accordance with some example embodiments;

[0083] FIG. 12C depicts a screenshot illustrating another example of a user interface, in accordance with some example embodiments;

[0084] FIG. 13 A depicts visual representations of groups of related protein sequences identified by imposing a different threshold for edit distance between two protein sequences, in accordance with some example embodiments;

[0085] FIG. 13B depicts visual representations of groups of related protein sequences identified by imposing a different average number of nearest neighbors for each protein sequence, in accordance with some example embodiments;

[0086] FIG. 13C depicts visual representations of groups of related protein sequences identified by assigning different weights to embeddings emphasizing the similarities in one or more specific regions of a protein sequence when generating a hybrid embedding, in accordance with some example embodiments;NAI-5007186271V1 17Attorney Ref.: 14786-058-228 (103963-228058) / P39076-WO-1

[0087] FIG. 14A depicts visual representations of the clusters of related protein sequences identified based on embeddings emphasizing similarities in one or more specific regions of each protein sequence, in accordance with some example embodiments;

[0088] FIG. 14B depicts visual representations of the clusters of related protein sequences identified based on embeddings emphasizing similarities in one or more specific regions of each protein sequence, in accordance with some example embodiments; and

[0089] FIG. 15 depicts a comparison of hyperbolic embeddings and Euclidean embeddings, in accordance with some example embodiments;

[0090] FIG. 16 depicts graphs illustrating a comparison of the performance of a generative model trained on a fill-in-the-middle (FIM) objective and further finetuned on hierarchical loss, in accordance with some example embodiments; and

[0091] FIG. 17 depicts a block diagram illustrating an example of a computing system, in accordance with some example embodiments.

[0092] When practical, similar reference numbers denote similar structures, features, or elements.DETAILED DESCRIPTION

[0093] Proteins are responsible for many essential cellular functions including, for example, enzymatic reactions, transport of molecules, regulation and execution of a number of biological pathways, cell growth, proliferation, nutrient uptake, morphology, motility, intercellular communication, and / or the like. An organism’s protein repertoire is encoded in its genome. For example, a coding sequence (or a protein-coding gene) is a deoxyribonucleic acid (DNA) or aNAI-5007186271V1 18Attorney Ref.: 14786-058-228 (103963-228058) / P39076-WO-1ribonucleic acid (RNA) sequence that specifies the sequence of amino acid residue in the polypeptide chain of a protein molecule. During protein synthesis, a coding sequence is transcribed into messenger ribonucleic acid (mRNA) that is then translated to form the polypeptide chain of the protein molecule. The diversity of proteins observed in nature is the result of evolutionary processes, including the duplication, divergence, and recombination of coding sequences. A protein molecule may include one or more domains, each of which being an evolutionary unit (e.g., of 100 to 250 amino acid residues) whose coding sequence can undergo duplication, divergence, and recombination. That novel protein molecules are formed by the duplication, divergence, and recombination of existing coding sequences means that at least some evolutionary relationships are present across different protein molecules. For instance, a domain family can contain protein molecules or, in some cases, portions of protein molecules, that descend from a common ancestor. Members of the same domain family may exhibit at least some similarities in the sequential order of amino acid residues and, by extension, similarities in function. Similarities (and dissimilarities) between protein molecules, such as those within the same domain family or, in some cases, across different domain families, may provide critical insights for engineering protein molecules with specific functions.

[0094] In the case of antibodies, for example, exposure to a pathogen may cause germline antibodies in the naive repertoire of antibodies to undergo affinity maturation, one mechanism of an adaptive immune system in which the coding sequences (or protein-coding genes) of the germline antibodies recombine to form novel antibodies with greater affinity towards the specific antigens present in the pathogen. The recombination of coding sequences (or protein-coding genes) gives rise to different clonotypes, each of which being a unique antigen receptor gene rearrangement nucleotide sequence that correspond to a population of identical cells. Antibodies NAI-5007186271V1 19Attorney Ref.: 14786-058-228 (103963-228058) / P39076-WO-1in the same clonotype (or subclonotype) may descend from a common ancestor (e.g., a fully rearranged, unmutated ancestor) and thus exhibit some measures of similarity therebetween. Thus, grouping a repertoire of antibodies into one or more clonotypes (and subclonotypes), for example, through multiple sequence alignment (MSA), is one example of conventional antibody repertoire analysis in which the diversity within an antibody repertoire is examined for insights into immune responses and for discovering new therapeutic agents. For example, clonotype grouping may identify two antibodies as originating from the same ancestor based on the two antibodies sharing the same variable (V) gene segment and / or joining (J) gene segment, and the sequence of amino acid residues forming the third complementarity region (CDR3) of each antibody exhibiting a threshold degree of similarity (e g., 70-100% similarity). However, drawbacks of conventional clonotype grouping include overlooking rare clonotypes, inaccurate clonotype identification due to sequencing errors, and improper gene assignment due to incomplete germline databases. Conventional antibody repertoire analysis also lack scalability and specificity. For instance, multiple sequence alignment (MSA) and clonotyping based analytics cannot scale to millions of antibody sequences, and are instead limited to small subsets of antibodies from each immunization campaign.

[0095] Generic machine learning based methodologies are also unsuitable for grouping protein molecules, such as antibodies, where different portions of a protein molecule have different impact on the properties of the protein molecule as a whole. Generic machine learning based methodologies, such as uniform manifold approximation and projection (UMAP) and t-distributed Stochastic Neighbor Embedding (t-SNE), group protein molecules based on similarities across entire protein molecules regardless of the significance of any specific regions therein. In doing so, generic machine learning based methodologies overlook key protein features, including the NAI-5007186271V1 20Attorney Ref.: 14786-058-228 (103963-228058) / P39076-WO-1significance of different regions of the protein molecule. For example, diversity in the complementarity determining regions of an antibody, particularly the third complementarity determining region (CDR3) of the heavy chain, may have a greater effect on the therapeutic properties of the antibody than diversity in the framework regions of the antibody. Thus, it may be more insightful for the grouping of antibodies to prioritize similarity in the complementarity determining regions of different antibodies than similarity in the respective framework regions. However, a generic machine learning based methodology would group two antibodies into the same cluster (or different clusters) regardless of whether the similarities between the two antibodies are present in the complementarity determining regions or the framework regions of each antibody. The region-agnostic groupings of antibodies resulting therefrom may be insufficiently granular for selecting candidate molecules therefrom.

[0096] Various example embodiments of the present disclosure addresses the shortcomings of conventional approaches to protein analytics, including antibody repertoire analysis, by at least identifying related protein sequences based on embeddings of protein sequences, each of which having been generated to emphasize similarities that exist within certain regions (e.g., one or more complementarity determining regions (CDR)) of the corresponding protein sequence. For example, in some example embodiments, a protein analysis engine may include an embedding model that generates, for a protein sequence, an embedding that captures the significance of the diversity (or conservation) in different regions of the protein sequences. Furthermore, in some cases, related protein sequences within a population of protein sequences, such as the repertoire of antibodies generated in response to pathogen exposure, may be identified based on the corresponding embeddings for visualization in the two-dimensional or three-dimensional space. In some cases, related protein sequences may be identified based on the NAI-5007186271V1 21Attorney Ref.: 14786-058-228 (103963-228058) / P39076-WO-1embedding of each protein sequence such that the groupings resulting therefrom prioritize similarities in certain regions. For instance, related antibodies identified based on the corresponding embeddings may exhibit a threshold degree of similarities in the third complementarity region (CDR3). However, two antibodies that exhibit similar framework regions may not necessarily be identified as related antibodies unless a threshold degree of similarities is also present in the third complementarity region (CDR3) of each antibody. This paradigm may be consistent with the observation that the sequence of amino acid residues forming the third complementarity region (CDR3), particularly the third complementarity determining region (CDR3) in the heavy chain, has a greater effect on the binding affinity and specificity of an antibody than the sequence of amino acid residues in the framework region.

[0097] In this context, a protein sequence may refer to the permutation of different amino acid residues. As such, in some cases, the embedding model may ingest a representation of a protein sequence that includes the identity (or type) and, in some cases, the relative position (or order), of each constituent amino acid residue. In some cases, the embedding model may include an encoder from a language model (or a large language model (LLM)), such as a language model (or large language model) with a transformer architecture, trained to generate the embedding of each protein sequence. In some cases, the embedding computation model may be trained, through contrastive learning (e.g., the reduction (or minimization) of a contrastive loss), to generate similar embeddings for similar protein sequences and dissimilar embeddings for dissimilar protein sequences. For example, in some cases, the training dataset for the embedding model may include one or more positive pairs of protein sequences, each of which being identified based on an edit distance (e g., Hamming distance and / or the like) between one or more specific regions (e.g., CDR3 region) of two constituent protein sequences satisfying one or more thresholds. NAI-5007186271V1 22Attorney Ref.: 14786-058-228 (103963-228058) / P39076-WO-1Alternatively and / or additionally, the training dataset for the embedding model may include one or more negative pairs of protein sequences, each of which being identified based on the edit distance (e.g., Hamming distance and / or the like) between the one or more specific regions (e.g., CDR3 region) of two constituent protein sequences failing to satisfy the one or more thresholds. In some cases, a specific region within a protein sequence may include a subset of amino acid residues in the protein sequence occupying one or more specific positions, serving one or more specific functionalities, determinative of one or more properties of the protein sequence, and / or the like. For instance, a specific region in the case of antibodies may be a region (e.g., variable region (Fv), a complementarity determining region (CDR), and / or the like) participating in a binding interaction with a target antigen and is therefore determinative of the binding affinity and specificity of antibodies. It should be appreciated that the same specific region in two or more different protein sequences may include different types of amino acid residues. Furthermore, in some cases, it may be possible for the same specific region in two or more different protein sequences to be at different positions in each protein sequence.

[0098] In some example embodiments, the embedding computation model may be trained to generate similar embeddings for each positive pair of protein sequences and dissimilar embeddings for each negative pair of protein sequences. In doing so, the embedding computation model may be trained to generate more similar embeddings for two protein sequences that exhibit similarities in the one or more specific regions than for two protein sequences whose similarities are limited to other regions. For example, the embedding computation model may generate more similar embeddings for two antibodies with similar third complementarity regions (CDR3) than for two antibodies with similar framework regions but dissimilar third complementarity regions (CDR3). In some cases, in addition to training the embedding computation model through NAI-5007186271V1 23Attorney Ref.: 14786-058-228 (103963-228058) / P39076-WO-1contrastive learning (e.g., the minimization (or reduction) of a contrastive loss), the embedding computation model may undergo post training to reduce (or minimize) a hierarchical loss (e.g., tree loss, angle-based entailment loss, and / or the like). As described in more detail below, imposing the hierarchical loss may train the embedding computation model to capture the hierarchical relationships that exist within an antibody repertoire, such as the hierarchical relationship between a parent protein sequence and a child protein sequence. In this context, a child protein sequence is a protein sequence that descended or evolved from a parent protein sequence. In the case of antibodies, for example, a child antibody may descend or evolve from a parent antibody as the result of affinity maturation In some cases, a parent protein sequence and a child protein sequence may be identified based on the extent of “entailment” therebetween. In the case of antibodies, the entailment between a parent antibody and a child antibody may be quantified by commonalities between the respective variable (V) and / or joining (J) gene segments.

[0099] In some example embodiments, the embedding computation model may be trained to generate embeddings that occupy a two-dimensional Euclidean space, which is a flat surface upon which the distance between embeddings reflect the extent of similarities in one or more specific regions of the underlying protein sequences. In some cases, one or more related groups of protein sequences may be identified by at least clustering the corresponding embeddings. It should be appreciated that grouping similar protein sequences using embeddings in two-dimensional Euclidean space do not adequately capture the hierarchical relationships that may exist between protein sequences, such as the hierarchical relationship within an antibody repertoire that arise due to affinity maturation. Hierarchical relationships in the resulting clusters of embeddings tend to be a mere byproduct of the embedding computation model being trained to generate similar embeddings for related protein sequences (e.g., parent and children antibodies) NAI-5007186271V1 24Attorney Ref.: 14786-058-228 (103963-228058) / P39076-WO-1with sufficient similarities in one or more specific regions (e.g., variable (V) gene segmentjoining (J) gene segment, and / or the like) rather than an actual recognition of the entailment between the related protein sequences. Accordingly, in some cases, instead of (or in addition to) being trained to generate embeddings occupying a two-dimensional Euclidean space, the embedding computation model may be trained to generate embeddings that occupy a hyperbolic space (e.g., Lorentz manifold and / or the like). In some cases, the curvature of the hyperbolic space may better preserve the hierarchical relationship between protein sequences, such as parent and child antibodies. As described in more detail below, the embedding computation model may be trained, based on a hierarchical loss (e.g., tree loss, angle-based entailment loss, and / or the like), to generate the embeddings of related antibodies (e g., parent and children antibodies) to occupy a geometric space (e.g., cone) within the hyperbolic space that reflects the hierarchical relationship therebetween. For instance, in some cases, the embedding computation model may be trained, in addition to reducing (or minimizing) the distance between the embeddings of related protein sequences, to also reduce (or minimize) the angle between such embeddings, such that the embedding of a child protein sequence occupy a geometric cone beneath the embedding of the parent protein sequence.

[0100] In some example embodiments, the embedding model may generate, for a protein sequence, a hybrid embedding that integrates at least two different embeddings of the protein sequence. For example, in some cases, the hybrid embedding may include one embedding generated by the embedding computation model, which has been trained to generate similar embeddings for protein sequences exhibiting similarities in one or more specific regions (e.g., CDR3) and dissimilar embeddings for protein sequences failing to exhibit similarities in one or more specific regions (e.g., CDR3). Furthermore, in some cases, the hybrid embedding may NAI-5007186271V1 25Attorney Ref.: 14786-058-228 (103963-228058) / P39076-WO-1include an additional embedding that generalizes the one or more specific regions. For an antibody, the additional embedding may generalize a particular complementarity determining region (CDR), e.g., the third complementarity determining region (CDR3), by assigning each constituent amino acid residue to an amino acid group including, for example, a non-polar group, a polar group, an aromatic group, a positive-charged (basic) group, a negatively-charged (acidic) group, a special group, and / or the like. In some cases, the hybrid embedding may be a combination (eg., union) of the embedding and the additional embedding. In some cases, the clusters of protein sequences that are identified based on the hybrid embedding may exhibit more refined boundaries than the clusters of protein sequences identified based on the embedding (or the additional embedding) alone.

[0101] In some example embodiments, the embeddings generated by the embedding computation model may be applied to toward context-based generation of novel protein sequences. For example, in some cases, a generative model may be trained using one or more sets of protein sequences, such as the groups of related protein sequences identified based on the embeddings generated by the embedding computation model. In some cases, the generative model (e.g., a language model) may be pretrained on a fill-in-the-middle (FIM) objective in which the generative model is trained to determine the identities of a contiguous span of amino acid residues in a set of protein sequences, such as a group of related protein sequences identified based on the embeddings generated by the embedding computation model. For instance, in some cases, the individual protein sequences in the set of protein sequences may be concatenated for ingestion by the generative model. In some cases, the identities of the individual amino acid residues in the continuous span of amino acid residues may be masked and, in some cases, an unmasked copy of those amino acid residues may be appended at the end of the concatenated protein sequences. InNAI-5007186271V1 26Attorney Ref.: 14786-058-228 (103963-228058) / P39076-WO-1some cases, the fill-in-the-middle (FIM) objective may include training the generative model to leverage the context (or generative context) around the masked amino acid residues, including one or more preceding and subsequent amino acid residues, when determining the identities of these amino acid residues. In instances where the generative context around the masked amino acid residues include other protein sequences in the same related group of protein sequences, the generative model may be trained to generate novel protein sequences (or portions thereof) that belong to the same related group of protein sequences. As noted, in some cases, the embeddings generated by the embedding computation model may be used to identify groups of related protein sequences exhibiting similarities in one or more specific regions with greater impact on certain properties of interest, such as complementarity determining regions (CDRs) in the case of the binding affinity and specificity of antibodies. Accordingly, the generative model trained on a related group of protein sequences may be capable of generating novel protein sequences exhibiting similarities in the one or more specific regions as the other protein sequences in the same group and are therefore more likely to exhibit the corresponding properties of interest.

[0102] In some example embodiments, the generative model pretrained on the fill-in-the-middle (FIM) objective may be fine-tuned (or post trained) with direct preference optimization (DPO). In the absence of direct preference optimization (DPO) finetuning, the pretrained generative model may generate, based on a generative context of protein sequences with similar embeddings, novel protein sequences that reflect a consensus across the protein sequences in the generative context. In some cases, the novel protein sequences generated in this manner may exhibit similar properties as the those in the generative context but not necessarily an improvement in any properties of interest. Doing so may ignore the hierarchy that may be present in the generative context in which one or more properties of interest are improved over successiveNAI-5007186271V1 27Attorney Ref.: 14786-058-228 (103963-228058) / P39076-WO-1generations of protein sequences in the generative context. For example, in some cases, the protein sequences forming the generative context may include parent and child antibodies resulting from a naive antibody repertoire undergoing affinity maturation. In some cases, direct preference optimization (DPO) finetuning may train the generative model to generate novel protein sequences that evolve more specifically from certain protein sequences in the generative context. As described in more detail below, with direct preference optimization (DPO) finetuning, the generative model may be capable of generating novel protein sequences with further improvement to one or more properties of interest rather than novel protein sequences with merely similar properties as those in the generative context.

[0103] FIG. 1 depicts a system diagram illustrating an example of a protein analysis system 100, in accordance with some example embodiments. Referring to FIG. 1, the protein analysis system 100 may include an analysis controller 110, a sequencing platform 120, and a client device 130. As shown in FIG. 1, the analysis controller 110, the sequencing platform 120, and the client device 130 may be communicatively coupled via a network 140. The sequencing platform 120 may apply a variety of sequencing techniques (e.g., next generation sequencing (NGS), third generation sequencing, and / or the like) to determine the sequence of nucleotide bases in one or more polynucleotides (e.g., nucleic acid molecules such as deoxyribonucleic acid (DNA), ribonucleic acid (RNA), and variants or derivatives thereof (e.g., single stranded DNA)). As described in more detail below, raw sequencing data from the sequencing platform 120 may include coding sequences (or protein-coding genes), which the analysis controller 110 may translate into protein sequences (or sequences of amino acid residues) for further processing and visualization. The client device 130 may be a processor-based device including, for example, a workstation, a desktop computer, a laptop computer, a smartphone, a tablet computer, a wearableNAI-5007186271V1 28Attorney Ref.: 14786-058-228 (103963-228058) / P39076-WO-1apparatus, and / or the like. The network 140 may be a wired network and / or a wireless network including, for example, a local area network (LAN), a virtual local area network (VLAN), a wide area network (WAN), a public land mobile network (PLMN), the Internet, and / or the like.

[0104] In some example embodiments, the analysis controller 110 may include a preprocessing engine 111, an embedding engine 113, a clustering engine 115, a visualization engine 117, and a generative model 119. The preprocessing engine 111 may translate, into protein sequences (or sequences of amino acid residues, raw sequencing data (e.g., next generation sequencing (NGS) data) from the sequencing platform 120. In order to do so, the preprocessing engine 111 may perform operations such as merging forward and reverse reads, translation, quality check, identifying overlaps between sequences (e g., all-pairs suffix prefix (APSP)), deduplication, alignment, and / or the like. In some cases, the embedding engine 113 generate an embedding for each protein sequence identified by the preprocessing engine 111 operating on the raw sequencing data from the sequencing platform 120. For example, as described in more detail below, the embedding engine 113 may apply an embedding computation model 114 that has been trained to generate, for a protein sequence, an embedding that captures the significance of diversity (or conservation) within different regions (e.g., third complementarity determining region (CDR3), framework region, and / or the like) of the protein sequence. In some cases, the embedding computation model 114 may include an encoder from a language model (or large language model) having a transformer architecture. In some cases, the embedding generated by the embedding engine 113 may be a hybrid embedding that integrates multiple embeddings including, for example, an embedding that captures the significance of diversity (or conservation) in different regions of the protein sequence, an additional embedding generalizing certain regions of the protein sequence, and / or the like. In some cases, the clustering engine 115 may leverage theNAI-5007186271V1 29Attorney Ref.: 14786-058-228 (103963-228058) / P39076-WO-1embedding of each protein sequence in order to identify related protein sequences within a population of protein sequences, such as the repertoire of antibodies generated in response to pathogen exposure. In some cases, identifying related protein sequences in this manner may emphasize similarities within certain regions of each protein sequence, such as the third complimentarity determining region (CDR3) of an antibody, with greater impact to the properties that are exhibited by each protein sequence. For instance, in some cases, the visualization engine 117 may project the embeddings of the protein sequences into two-dimensional space (or three-dimensional space) to yield clusters of related protein sequences for visualization of the relationship present therein. In some cases, the visualization engine 117 may generate

[0001] In some example embodiments, the generative model 119 may be pretrained on a fill-in-the-middle (FIM) objective in which the generative model 119 is trained to determine the identifies of a contiguous span of amino acid residues in a set of protein sequences, such as a group of related protein sequences identified based on the clustering of the embeddings generated by the embedding computation model 114. For example, in some cases, the individual protein sequence in the set of protein sequences may be concatenated. Moreover, in some cases, the identities of the individual amino acid residues in the continuous span of amino acid residues may be masked and, in some cases, an unmasked copy of those amino acid residues may be appended to the end of the concatenated protein sequences. In some cases, the fill-in-the-middle (FIM) objective may include training the generative model 110 to determine, based at least on the context (or generative context) around the masked amino acid residues, the identities of the masked amino acid residues. In some cases, the generative context may include one or more of the amino acid residues preceding and subsequent to the masked amino acid residues. In some cases, the masked amino acid residues may be an entire protein sequence (or one or more specific regions in a protein NAI-5007186271V1 30Attorney Ref.: 14786-058-228 (103963-228058) / P39076-WO-1sequence) in the set of protein sequences. Accordingly, where the set of protein sequences is a group of related protein sequences, the generative model 119 trained on the fill-in-the-middle (F1M) objective may be capable of generating novel protein sequences that belong to the same group of related protein sequences.

[0002] In some example embodiments, the generative model 119 pretrained on the fill-in-the-middle (FIM) objective may be fine-tuned (or post trained) with direct preference optimization (DPO). In the absence of direct preference optimization (DPO) finetuning (or post training), the pretrained generative model 119 may generate novel protein sequences that reflect a consensus across the protein sequences (e.g., group of related protein sequences) in the generative context. However, the novel protein sequences generated in this manner may merely exhibit similar properties as the those in the generative context but not necessarily an improvement in any properties of interest. This is because absent direct optimization (DPO) finetuning, the hierarchy that may be present in the generative context in which one or more properties of interest are improved over successive generations of protein sequences in the generative context is ignored. For example, in the context of antibodies, generating novel antibodies along the evolutionary path away from the germline may yield incremental improvements in one or more properties of interest. As such, in some cases, direct preference optimization (DPO) finetuning may train the generative model 119 to generate novel protein sequences that evolve more specifically from certain protein sequences in the generative context. For instance, the generative model 119 finetuned with direct preference optimization (DPO) may generate novel protein sequences that are children protein sequence of certain protein sequences in the generative context, such as antibodies with a threshold distance to the germline. Accordingly, with direct preference optimization (DPO) finetuning, the generative model 119 may be capable of generating novel protein sequences with further NAI-5007186271V1 31Attorney Ref.: 14786-058-228 (103963-228058) / P39076-WO-1improvement to one or more properties of interest rather than novel protein sequences with merely similar properties as those in the generative context.

[0003] In some example embodiments, one or more of the functionalities of the analysis controller 110 may be invoked by an agent 135 at the client device 130. In some cases, the agent 135 may be an autonomous artificial intelligence (Al) agent trained to invoke the one or more functionalities as a part of its orchestration of a drug molecule design workflow. For example, in some cases, the agent 135 may invoke the analysis controller 110 to generate one or more embeddings and apply the one or more embeddings to identify one or more groups of related protein sequences. In some cases, the agent 135 may invoke one or more other computational models (or workflows) and provide at least a portion of the resulting output to the analysis controller 110. For instance, in some cases, one or more additional class labels determined by the other computational models (or workflows) may be used, in addition to the one or more embeddings generated by the analysis controller 110 to identify the one or more groups of related protein sequences. One example of an additional class label includes canonical conformation, which refers to one of a limited number of possible conformations (or three-dimensional structures) of a class of protein molecules or a portion of that class of protein molecules. Another example of an additional class label include a deviation from known protein structures.

[0004] FIG. 2A depicts a flowchart illustrating an example of a process 200 for machine learning enabled protein analysis, in accordance with some example embodiments. In some example embodiments, the process 200 may be performed by the analysis controller 110 including, for example, the preprocessing engine 111, the embedding engine 113, the clustering engine 115, and the visualization engine 117 such that the embedding computation model 114 (e.g., an encoder from a transformer-based language model (or large language model)) may be trained and applied NAI-5007186271V1 32Attorney Ref.: 14786-058-228 (103963-228058) / P39076-WO-1to generate embeddings of protein sequences. As described in more detail below, in some cases, two or more similar protein sequences that exhibit a threshold similarity in a specific region (e.g., CDR3) of each protein sequence may be determined. Moreover, two or more dissimilar protein sequences that fail to exhibit the threshold similarity in the specific region of each protein sequence may also be determined. Accordingly, the embedding computation model 114 may be trained through contrastive learning to generate similar embeddings with a threshold edit distance for the two similar protein sequences and dissimilar embeddings with a different threshold edit distance for the two dissimilar protein sequences.

[0005] At 202, a selected protein sequence is determined. In some example embodiments, the selected protein sequence may include a plurality of amino acid residues. In some cases, the identity of each amino acid residue in the selected protein sequence may be determined by translating one or more corresponding coding sequences (or protein-coding genes) in raw sequencing data. In some cases, the selected protein sequence may be a part of an entire population of protein sequences, such as the repertoire of antibodies generated in response to exposure to a pathogen. In some cases, the selected protein sequence may be determined by at least determining a representation of the selected protein sequence. For example, in some cases, the representation of the selected protein sequence may encode the identity (or type) of each amino acid residue in the selected protein sequence. Furthermore, in some cases, the representation of the selected protein sequence may encode the relative position (or order) of each constituent amino acid residue.

[0006] At 204, two or more similar protein sequences that exhibit a threshold similarity in a specific region of each protein sequence are determined. In some example embodiments, two protein sequences may be identified as similar protein sequences if the edit distance between aNAI-5007186271V1 33Attorney Ref.: 14786-058-228 (103963-228058) / P39076-WO-1specific region in each protein sequence satisfies one or more threshold criteria. For example, in the context of antibodies, two protein sequences may be identified as similar protein sequences if the edit distance between the complementarity determining region (CDR) of each protein sequence satisfies one or more threshold criteria. As described in more detail below, two similar protein sequences may form a positive pair for training an embedding computation model, such as the encoder of a language model (or large language model) with a transformer architecture, to generate embeddings of protein sequences that prioritize similarities in certain regions (e.g., complementarity determining regions (CDRs)) of each protein sequence over similarities in other regions (e.g., framework region).

[0007] At 206, two or more dissimilar protein sequences that fail to exhibit the threshold similarity in the specific region of each protein sequence are determined. In some example embodiments, two protein sequences may be identified as dissimilar protein sequences if the edit distance between the specific region in each protein sequence satisfies a different threshold value. For example, while two protein sequences may be identified as similar protein sequences if the edit distance between the specific region in each protein sequence satisfies a first threshold criteria, two protein sequences may be identified as dissimilar protein sequence if the edit distance between the specific region in each protein sequence satisfies a second threshold criteria. As described in more detail below, two dissimilar protein sequences may form a negative pair for training an embedding computation model, such as the encoder of a language model (or large language model) with a transformer architecture, to generate embeddings of protein sequences that prioritize similarities in certain regions (e.g., complementarity determining regions (CDRs)) of each protein sequence over similarities in other regions (e.g., framework region).NAI-5007186271V1 34Attorney Ref.: 14786-058-228 (103963-228058) / P39076-WO-1

[0008] At 208, an embedding computation model that has been trained on the two or more similar protein sequences and the two or more dissimilar protein sequences is applied to generate an embedding of the selected protein sequence. In some example embodiments, the embedding computation model that is applied to generate the embedding of the selected protein sequence may have been trained through contrastive learning. For example, in some cases, contrastive learning may include training the embedding computation model to generate similar embeddings for similar protein sequences that exhibit the threshold similarity in a specific region (e g., CDR3) of each protein sequence and dissimilar embeddings for dissimilar protein sequences that do not exhibit the threshold similarity in the specific region of each protein sequence. In this context, two embeddings may be considered similar if the edit distance between the two embeddings satisfies one threshold criteria. Alternatively, two embeddings may be considered dissimilar if the edit distance between the two embeddings satisfies a different threshold criteria. Accordingly, training the embedding computation model through contrastive learning may include adjusting one or more parameters (e.g., weights, biases, and / or the like) of the embedding computation model such that the edit distance between embeddings generated by the embedding computation model satisfies a first threshold criteria for two similar protein sequences and a second threshold criteria for two dissimilar protein sequences.

[0009] In some example embodiments, the embedding computation model may be trained to generate embeddings that occupy a two-dimensional Euclidean space (e g., a flat surface). For example, in some cases, the output of the embedding computation model may pass through one or more linear layers that project embeddings into the two-dimensional Euclidean space. In some cases, the distances between embeddings in the two-dimensional Euclidean space may correspond to the extent of similarities in one or more specific regions of the underlyingNAI-5007186271V1 35Attorney Ref.: 14786-058-228 (103963-228058) / P39076-WO-1protein sequences. However, grouping similar protein sequences using embeddings in two-dimensional Euclidean space do not adequately capture the hierarchical relationships that may exist between protein sequences (e.g., parent and children antibodies in an antibody repertoire that has undergone affinity maturation). Accordingly, in some cases, instead of being trained to generate embeddings occupying a two-dimensional Euclidean space, the embedding computation model may be trained to generate embeddings that occupy a hyperbolic space (e.g., Lorentz manifold and / or the like) whose curvature better preserve the hierarchical relationship between protein sequences. For instance, in some cases, in addition to the one or more linear layers, the output of the embedding computation model may be mapped to the hyperbolic space (e.g., Lorentz manifold and / or the like) using an exponential map. In some cases, the exponential map may be a function mapping a tangent vector from a point's tangent space in two-dimensional Euclidean space to a point on the manifold of the hyperbolic space (e.g., the Lorentz manifold and / or the like). In some cases, the exponential map may perform the mapping by at least traveling along a unique geodesic (or a straight line) in the direction of the vector, for example, one unit at a time. It should be appreciated that mapping a point and a tangent vector to a new point by traveling along a geodesic (or a straight line) may serve to move from the "flat" tangent space of the two-dimensional Euclidean space to the curved hyperbolic space.

[0010] At 210, one or more related groups of protein sequences are identified based at least on the embedding of the selected protein sequence. In some example embodiments, one or more related groups of protein sequences within the population of protein sequences, such as the repertoire of antibodies generated in response to pathogen exposure, may be identified based at least on the embedding of each protein sequence. In some cases, the one or more related groups of protein sequences may be identifying by at least applying, to the embedding of each proteinNAI-5007186271V1 36Attorney Ref.: 14786-058-228 (103963-228058) / P39076-WO-1sequence, a clustering technique. Examples of clustering techniques include hierarchical clustering, centroid-based clustering, density-based clustering, distribution-based clustering, and / or the like. Applying the clustering technique to the embedding of a protein sequence may assign, to each protein sequence, a cluster label, thus generating one or more clusters representative of groups of related protein sequences. As the cluster labels are assigned based on the embedding of each protein sequence, which prioritizes similarities in the one or more specific regions (e.g., CDR3), the resulting clusters may group protein sequences that exhibit greater similarities in the one or more specific regions. For antibodies, each cluster may correspond to a group of antibodies exhibiting sufficient similarities in the complementarity determining region (CDR) of each antibody. However, based on the embeddings, which prioritize similarities in the complementarity determining region, antibodies with similar framework regions may not be grouped into a same cluster unless those antibodies also exhibit sufficient similarities in the complementarity determining region.

[0011] At 212, a visual representation is generated based at least on the one or more identified groups of related protein sequences. In some example embodiments, the one or more groups of related protein sequences may be projected into a lower dimensional space, such as a two-dimensional or three-dimensional space, which is suitable for visualization. For example, in some cases, the embeddings of each protein sequence may be further projected into a two-dimensional (or three-dimensional) space in order to generate a visual representation depicting one or more groups of related protein sequences. In some cases, the relationship between different protein sequences within the population of protein sequences may be indicated by the distance (e.g., Euclidean distance) therebetween. Accordingly, two protein sequences that are more closely related, for example, by the presence of greater similarities (e.g., quantified by edit distance) inNAI-5007186271V1 37Attorney Ref.: 14786-058-228 (103963-228058) / P39076-WO-1certain regions (e.g., CDR3) of each protein sequence may be depicted as more proximate to one another (e.g., neighbors) in two- or three-dimensional space. Contrastingly, two protein sequences that are less closely related, for example, by the presence of greater dissimilarities (e.g., quantified by edit distance) in certain regions (e g., CDR3) of each protein sequence may be depicted as more distant to one another (e.g., non-neighbors) in two- or three-dimensional space.

[0012] The distance placed between individual protein sequences may yield one or more clusters in the visual representation, each of which corresponding to a group of related protein sequence who exhibit sufficient similarities in certain regions (e.g., CDR3) to be disposed within close proximity of one another in two- or three-dimensional space. Meanwhile, groups of unrelated protein sequences may form separate clusters, with the distance between two clusters indicative of the relationship (e.g., similarities in certain regions such as CDR3) between the protein sequences in each cluster. Accordingly, in some cases, the visual representation of the one or more groups of related protein sequences may show the distribution of the one or more groups of related protein sequences in a two-dimensional (or three-dimensional) space. Furthermore, in some cases, each protein sequence may be shown as a visual indicator that is rendered in a shape, color, and / or size representative of one or more properties of the protein sequence. For instance, for antibodies, each antibody in the visual representation may be shown as either a first visual indicator representative of a binding antibody, a second visual indicator representative of a weakly-binding antibody, a third visual indicator representative of a non-binding antibody, and / or the like.

[0013] FIG. 2B depicts a flowchart illustrating an example of a process 250 for training an embedding computation model to generate an embedding of a protein sequence that emphasize similarities in certain regions of the protein sequence, in accordance with some example embodiments. In some example embodiments, the process 250 may be performed by the analysisNAI-5007186271V1 38Attorney Ref.: 14786-058-228 (103963-228058) / P39076-WO-1controller 110 such as, for example, by the embedding engine 113, in order to train the embedding computation model 114 to generate embeddings of protein sequences. As described in more detail below, in some cases, the embedding computation model 114 may be trained through contrastive learning to generate similar embeddings for two protein sequences exhibiting a threshold degree of similarity in one or more specific regions (e.g., CDR3) of each protein sequence and dissimilar embeddings for two protein sequences that fail to exhibit the threshold degree of similarities in the one or more specific regions.

[0014] At 252, a pair of protein sequences is identified as a positive pair based at least on the pair of protein sequences exhibiting a threshold similarity in one or more specific regions of each protein sequence. In some example embodiments, an embedding computation model, such as the encoder of a language model (or large language model) with a transformer architecture, may be trained to generate embeddings of protein sequences. In some cases, the embedding computation model (e.g., the encoder) may be trained based on a training dataset that includes one or more positive pairs. In this context, a positive pair may include a first protein sequence and a second protein sequence that exhibit a threshold degree of similarity in one or more specific regions of each protein sequence. For example, for antibodies, a positive pair may include a first antibody and a second antibody exhibiting a threshold degree of similarity in the third complementarity determining region (CDR3) of each antibody. In some cases, to be identified as a positive pair, the first protein sequence and the second protein sequence may be required to satisfy one or more additional criteria. For instance, in addition to the threshold degree of similarity in one or more specific regions of each protein sequence, the first protein sequence and the second protein sequence may be identified as a positive pair if the first protein sequence andNAI-5007186271V1 39Attorney Ref.: 14786-058-228 (103963-228058) / P39076-WO-1the second protein sequence exhibit a threshold degree of similarity across the entire length of each protein sequence.

[0015] In some example embodiments, the similarity (or dissimilarity) between two protein sequences may be quantified by a similarity metric such as an edit distance (e.g., Hamming distance) and / or the like. For example, in some cases, the first protein sequence and the second protein sequence may be identified as a positive pair if the similarity metric (e.g., edit distance and / or the like) quantifying the similarity between the entire length of each protein sequence and / or the similarity between one or more specific regions of each protein sequence satisfies one or more threshold criteria. In some cases, the one or more thresholds may be fixed values. However, as described in more detail below, the one or more thresholds may be adjusted dynamically based on the diversity of the population of protein sequences, such as the repertoire of antibodies generated in response to pathogen exposure, from which the two protein sequences originate.

[0016] In some example embodiments, the one or more thresholds for identifying positive pairs may be set dynamically based on the diversity of the population of protein sequences and one or more user inputs. In some cases, the one or more user inputs may specify one or more limits on the quantity of nearest neighbors for each protein sequence when the protein sequence are projected into a two-dimensional (or three-dimensional) space. In this context, the nearest neighbor of a protein sequence may be another protein sequence that is within a threshold distance (e.g., Euclidean distance) of the protein sequence in the two-dimensional (or three-dimensional) space. The projection into two-dimensional (or three-dimensional) space may be achieved by reducing the dimensionality of a representation of each protein sequence that includes the identity (or type) and, in some cases, the relative position (or order), of each constituent amino acid residue.NAI-5007186271V1 40Attorney Ref.: 14786-058-228 (103963-228058) / P39076-WO-1

[0017] In some cases, a subset of protein sequences may be randomly sampled from the population of protein sequences to establish the one or more thresholds to ensure that the average quantity of nearest neighbors per protein sequence match the user-specified limits on the quantity of nearest neighbors for each protein sequence. In some cases, once the one or more thresholds are established, for example, a positive pair may be identified by selecting the first protein sequence. In some cases, once the first protein sequence is selected, the second protein sequence may be selected from the nearest neighbors of the first protein sequence. In some cases, the second protein sequence may be selected as a part of the positive pair with the first protein sequence if the second protein sequence exhibits the threshold degree of similarity in the one or more specific regions (e.g., CDR3) and, in some cases, across its entire length as the first protein sequence. In some cases, the second protein sequence may be selected instead of a third protein sequence that is also nearest neighbor of the first protein sequence based on the similarity metric (e.g., edit distance) associated with the second protein sequence satisfying the one or more thresholds and the similarity metric (e.g., edit distance) associated with the third protein sequence failing to satisfy the one or more thresholds. In instances where the second protein sequence and the third protein sequence are associated with the same similarity metric (e.g., edit distance), an embedding computation model may be applied to identify the one protein sequence with a higher relative similarity. In some cases, neither the first protein sequence nor the second protein sequence may be identified to be a part of a positive pair if the quantity of nearest neighbors associated with either protein sequence fails to satisfy one or more thresholds. For instance, in some cases, a protein sequence that has fewer than a threshold quantity (e.g., two, three, and / or the like) nearest neighbors are excluded from the training dataset at least because such protein sequences do not contribute to meaningful placement in two-dimensional (or three-dimensional) space.NAI-5007186271V1 41Attorney Ref.: 14786-058-228 (103963-228058) / P39076-WO-1

[0018] At 254, the pair of protein sequences is identified as a negative pair based at least on the pair of protein sequences failing to exhibit the threshold similarity in the one or more specific regions of each protein sequence. In some example embodiments, the training dataset for training the embedding computation model may further include one or more negative pairs. In this context, a negative pair may include a first protein sequence and a second protein sequence that fails to exhibit a threshold degree of similarity in one or more specific regions of each protein sequence. In the case of antibodies, a negative pair may include a first antibody and a second antibody that fails to exhibit a threshold degree of similarity in the third complementarity determining region (CDR3) of each antibody. In some cases, to be identified as a negative pair, the first protein sequence and the second protein sequence may be required to satisfy one or more additional criteria. For example, in addition to lacking the threshold degree of similarity in one or more specific regions of each protein sequence, the first protein sequence and the second protein sequence may be identified as a negative pair if the first protein sequence and the second protein sequence fails to exhibit a threshold degree of similarity across the entire length of each protein sequence. As noted, in some cases, the similarity (or dissimilarity) between two protein sequences, such as the similarity across the entire length of each protein sequences or the similarity between specific regions therein, may be quantified by a similarity metric such as an edit distance (e.g., Hamming distance) and / or the like. The one or more thresholds associated with similarity metric may be fixed values or, alternatively, values that are set dynamically to ensure that the quantity of protein sequences in each group of related protein sequences resulting therefrom is consistent with one or more user inputs. Moreover, in some cases, the quantity of positive pairs and the quantity of negative pairs in the training dataset may be balanced to ensure effective contrastive learning.NAI-5007186271V1 42Attorney Ref.: 14786-058-228 (103963-228058) / P39076-WO-1In some cases, the quantity of positive pairs and the quantity of negative pairs in the training dataset may be balanced by setting, proportionately, the batch size and the quantity of seed sequences. For example, 8 seed sequences (e.g., 8 randomly selected seed sequences) may be needed where the batch size is set to 256 protein sequences in order to ensure that the quantity of positive pairs in the batch is balanced with the quantity of negative pairs in the batch. In some cases, the balance between the quantity of positive pairs and the quantity of negative pairs may include the batch being configured to contain a sufficient number of neighboring (or similar) protein sequences and non-neighboring (or dissimilar) protein sequences. For instance, for each of the 8 seed sequences, 32 of its neighboring protein sequences may be added to the batch. The resulting batch of 256 protein sequences may yield balanced quantities of positive pairs and negative pairs at least because the probability of encountering non-neighboring (dissimilar) sequences is generally much higher than the probability of encountering neighboring (or similar) sequences. Accordingly, a protein sequence that is a neighboring sequence of one seed sequence is unlikely to a neighboring sequence of another seed sequence. Moreover, two protein sequences in one positive pair may form negative pairs when combined with protein sequences from other positive pairs. As such, the foregoing configuration may ensure a minimum number of neighboring (or similar) protein sequences and non-neighboring (or dissimilar) protein sequences in each batch without explicitly adding non-neighboring (or dissimilar) protein sequences to the batch.

[0019] At 256, an embedding computation model is trained by at least adjusting one or more parameters of the embedding computation model to generate similar embeddings for each positive pair and dissimilar embeddings for each negative pair. In some example embodiments, an embedding computation model, such as the encoder of a language model (or a large language model) having a transformer architecture, may be trained to generate embeddings of proteinNAI-5007186271V1 43Attorney Ref.: 14786-058-228 (103963-228058) / P39076-WO-1sequences that captures the significance of diversity within different regions (e.g., third complementarity determining region (CDR3), framework region, and / or the like) of each protein sequence. In some cases, the embedding computation model (e.g., the encoder) may be trained through contrastive learning, meaning that the training of the embedding computation model may include adjusting one or more parameters (e.g., weights, biases, and / or the like) of the embedding computation model to increase the similarity of embeddings generated by the embedding computation model for positive pairs in the training dataset and the dissimilarity of embeddings generated by the embedding computation model for negative pairs in the training dataset. In some cases, the similarity (and dissimilarity) between two embeddings may be quantified by a similarity metric such as an edit distance (e.g., Hamming distance) and / or the like.

[0020] In some example embodiments, the training of the embedding computation model may include reducing (or minimizing) a contrastive loss quantifying the similarity between the embeddings of positive pairs and the dissimilarity between the embeddings of negative pairs. In some cases, the training of the embedding computation model may include training the embedding computation model (e.g., the encoder) by at least adjusting one or more parameters (e.g., weights, biases and / or the like) of the embedding computation model to reduce (or minimize) the edit distance (e.g., Hamming distance and / or the like) between the embeddings of positive pairs and increase (or maximize) the edit distance (e.g., Hamming distance and / or the like) between the embeddings of negative pairs. An example of the loss function is shown below, wherein lposdenotes the loss associated with positive pairs and lnegdenotes the loss associated with negative pairs. In some cases, the embedding computation model may be pretrained to recover the identity of one or more masked amino acid residues in protein sequences before undergoing contrastive learning to generate similar embeddings for protein sequences exhibiting a threshold degree of NAI-5007186271V1 44Attorney Ref.: 14786-058-228 (103963-228058) / P39076-WO-1similarity in one or more specific regions (e.g., CDR3) of each protein sequence and dissimilar embeddings for protein sequences that fail to exhibit the threshold degree of similarities in the one or more specific regions. Doing so may pretrain the embedding computation model to capture the patterns and dependencies that may be present in protein sequences.loss — (Ipos + InegIpos =—log(similarity score to neighbors)lneg=log(similarity score to non — neighbors)1similarity = - - - - - (1 + distance)

[0021] In some cases, the embedding computation model may be trained to reduce (or minimize) contrastive loss such that the distances between the resulting embeddings reflect the extent of similarities in one or more specific regions of the underlying protein sequences. However, the embeddings generated by the embedding computation model trained on contrastive loss alone may fail to adequately capture the hierarchical relationships that may exist between protein sequences, such as hierarchy of parent and children antibodies that arise when an antibody repertoire undergo affinity maturation. To the extent the embedding computation model may generate similar embeddings for parent and child protein sequences with sufficient similarities in specific regions, the mere proximity of embeddings in latent space (e.g., two-dimensional Euclidean space, hyperbolic space, and / or the like) do not necessarily reflect an evolutionary hierarchy of successive generations of children protein sequences descending from prior parent generations. Accordingly, in some example embodiments, in addition to or instead of the contrastive loss, the embedding computation model may be trained to reduce (or minimize) a hierarchical loss (e.g., tree loss, angle-based entailment loss, and / or the like). For example, as described in more detail below, the embedding computation model may be pretrained to reduceNAI-5007186271V1 45Attorney Ref.: 14786-058-228 (103963-228058) / P39076-WO-1(or minimize) contrastive loss before being finetuned (or post trained) to reduce (or minimize) hierarchical loss.

[0022] In some example embodiments, subsequent to undergoing contrastive learning in which the embedding computation model is trained to reduce (or minimize) contrastive loss, the embedding computation model may be finetuned on a hierarchical loss (e.g., tree loss, angle-based entailment loss, and / or the like) that geometrically quantifies the hierarchical relationship between the embeddings of protein sequences generated by the embedding computation model. In some cases, in order to finetune (or post train) the embedding computation model, a hierarchy of protein sequences may be defined. In some cases, the hierarchy may include successive generations of protein sequences, such as parent protein sequences and children sequences descending therefrom. In some cases, the hierarchy may be asymmetric, meaning that that the branches of the hierarchy may exhibit different depths (or quantity of generations). For example, in some cases, the hierarchy may be defined by at least identifying, for the embedding of a protein sequence, one or more of similar embeddings in latent space (e.g., fc-nearest neighbors (kNN) in two-dimensional Euclidean or hyperbolic space). In some cases, the hierarchy may be further defined by at least determining the hierarchical relationship between the similar embeddings. For instance, in some cases, the hierarchy may be the hierarchy of parent and child antibodies that exist within an antibody repertoire that has undergone affinity maturation. As such, in some cases, the hierarchy may be further defined based on an entailment score (e.g., a variable (V) gene and joining (J) gene entailment score in the case of antibodies) between protein sequences.

[0023] In some example embodiments, the embedding computation model may be finetuned (or post trained) based on the hierarchy defined, for example, in the aforementioned manner. In some cases, the embedding computation model may be finetuned (or post trained) to NAI-5007186271V1 46Attorney Ref.: 14786-058-228 (103963-228058) / P39076-WO-1reduce (or minimize) hierarchical loss. For example, in some cases, the embedding computation model may be finetuned (or post trained) to reduce (or minimize) an angle-based entailment loss between two protein sequences. In some cases, the finetuning (or post training) of the embedding computation model may include adjusting or more parameters (e.g., weights, biases, and / or the like) of the embedding computation model to increase (or maximize) the angle between the embeddings of a positive pair of protein sequences, which in this case may include two related protein sequences such as a parent protein sequence and a child protein sequence descending therefrom. In some cases, the finetuning (or post training) of the embedding computation model may further include adjusting one or more parameters (e.g., weights, biases, and / or the like) of the embedding computation model to reduce (or minimize) the angle between the embeddings of a negative pair of protein sequences, which in this case may include two unrelated protein sequences. Where the embeddings occupy a hyperbolic space (e.g., Lorentz manifold and / or the like), reducing (or minimizing) angle-based entailment loss may train the embedding computation model to generate embeddings of children protein sequences to fall within a geometric cone beneath the embeddings of the parent protein sequences while the embeddings of unrelated protein sequences fall outside of the geometric cone.

[0024] In some example embodiments, the embedding computation model (e.g., the encoder) may be trained over multiple phases. For example, in instances where the embedding computation model includes the encoder of a language model (or large language model) having a transformer architecture, the training of the embedding computation model may include training the encoder and a corresponding decoder. In some cases, the embedding computation model may undergo pretraining in order for the encoder to generate embeddings of a certain dimensionality (e.g., 256-dimensions) for each protein sequence ingested by the embedding computation model. NAI-5007186271V1 47Attorney Ref.: 14786-058-228 (103963-228058) / P39076-WO-1In some cases, the embedding computation model may be coupled with a linear transformation layer that ingest the embeddings generated by the embedding computation model and produces an output having the same number of dimensions. A projection layer, such as a linear projection layer (or head), may then map the output of the linear transformation layer from its higher dimensional space (e.g., 256-dimensions) to a lower dimensional space, such as a two-dimensional or three-dimensional space, which is more suitable for visualization and downstream analysis. In some cases, the training of the encoder may include training the linear transformation and projection layers. For instance, during an initial training phase, the embedding computation model including the encoder and the linear transformation layer may be trained to adopt the embeddings to a taskspecific context (e.g., antibodies) over a first quantity of epochs (e.g., 400 epochs). Thereafter, during a subsequent training phase, the projection layer (or head) is trained over a second quantity of epochs (e.g., 100 epochs) while the parameters (e.g., weights, biases) of the embedding computation model (e.g., the encoder) and the linear transformation layer are fixed. The objective of this phase is to refine the lower-dimensional projections generated by the linear projection layer to better reflect the relationships between protein sequences without affecting the embedding computation model or the linear transformation layer. During a final training phase, the embedding computation model, the linear transformation layer, and the linear project layer (or head) may be trained for a third quantity of epochs (e.g., 400 epochs).

[0025] At 258, the trained embedding computation model is applied to generate an embedding for a protein sequence. In some example embodiments, the embedding computation model, upon undergoing training through contrastive learning, may be applied to generate embeddings of protein sequences, such as protein sequences from a repertoire that is generated in response to exposure to a pathogen. In some cases, the embedding computation model may NAI-5007186271V1 48Attorney Ref.: 14786-058-228 (103963-228058) / P39076-WO-1generate, for each protein sequence, an embedding that emphasizes similarities in one or more specific regions of the protein sequence. Such embeddings may enable the identification of related protein sequences to prioritize similarities in regions with a greater impact on the properties of the protein sequences over similarities in those regions with less impact on the properties of the protein sequences. For example, in the case of antibodies, the embedding computation model may generate embeddings that emphasize similarities in the third complementarity determining region (CDR3) of each antibody. These embeddings may enable the identification of related antibodies to prioritize similarities in the third complementarity determining regions (CDR3) of each antibody over similarities in the framework regions of each antibody. Doing so may be consistent with the observation that diversity within the third complementarity determining region (CDR3) of antibodies, particularly the heavy chain complementarity determining region (CDR3), have a greater impact on binding affinity towards certain antigens than diversity within the framework region of antibodies.

[0026] FIG. 3A depicts a flowchart illustrating an example of a process 300 for machine learning enabled analysis of protein sequences, in accordance with some example embodiments. In some example embodiments, the process 300 may be performed by the analysis controller 110 including, for example, the preprocessing engine 111, the embedding engine 113, the clustering engine 115, the visualization engine 117, and / or the like. As described in more detail below, in some cases, while the preprocessing engine 111 may translate raw sequencing data (e.g., from the sequencing platform 120) into protein sequences (e.g., sequences of amino acid residues), the clustering engine 115 and the visualization engine 117 may leverage embeddings of these protein sequences generated by the embedding engine 113 in order to identify, analyze, and visualize the relationships therebetween.NAI-5007186271V1 49Attorney Ref.: 14786-058-228 (103963-228058) / P39076-WO-1

[0027] At 302, a protein sequence including a plurality of amino acid residues is determined. In some example embodiments, the identity of each amino acid residue in the protein sequence may be determined by translating one or more corresponding coding sequences (or protein-coding genes) in raw sequencing data. In some cases, the protein sequence may be a part of an entire population of protein sequences, such as the repertoire of antibodies generated in response to exposure to a pathogen. In some cases, the protein sequence may be determined by at least determining a representation of the protein sequence that encodes, for example, the identity (or type) of each amino acid residue in the protein sequence and, in some cases, the relative position (or order), of each constituent amino acid residue.

[0028] At 304, an embedding computation model is applied to generate an embedding of the protein sequence. In some example embodiments, an embedding computation model, such as the encoder from a language model (or large language model) having a transformer architecture, may be applied to generate an embedding of the protein sequence that prioritizes similarities in one or more specific regions of the protein sequence over similarities in other regions. For example, in some cases, the embedding computation model may have been trained to generate similar embeddings for protein sequences that exhibit a threshold degree of similarity in the one or more specific regions and dissimilar embeddings for protein sequences that fails to exhibit the threshold degree of similarity in the one or more specific regions. In the context of antibodies, the embedding that is generated by the embedding computation model may prioritize similarities in the complementarity determining regions (CDR) over similarities in the framework regions. In some cases, the embedding may be a hybrid embedding that combines the embedding prioritizing similarities in the one or more specific regions with another embedding generalizing the one or more specific regions, for example, by assigning each constituent amino acid residue to a NAI-5007186271V1 50Attorney Ref.: 14786-058-228 (103963-228058) / P39076-WO-1corresponding amino acid group (e.g., a non-polar group, a polar group, an aromatic group, a positive-charged (basic) group, a negatively-charged (acidic) group, or a special group).

[0029] At 306, one or more related groups of protein sequences are identified based at least on the embedding of the protein sequence. In some example embodiments, one or more related groups of protein sequences within the population of protein sequences may be identified based at least on the embedding of each protein sequence. For example, in some cases, a clustering technique, such as hierarchical clustering, centroid-based clustering, density-based clustering, distribution-based clustering, and / or the like, may be applied in order to assign, to each protein sequence, a cluster label. Doing so may generate one or more clusters, each of which representative of a group of related protein sequences. As the cluster labels are assigned based on the embedding of each protein sequence, which prioritizes similarities in the one or more specific regions (e g., CDR3), the resulting clusters may group protein sequences that exhibit greater similarities in the one or more specific regions. For antibodies, each cluster may correspond to a group of antibodies exhibiting sufficient similarities in the complementarity determining region (CDR) of each antibody. However, based on the embeddings, which prioritize similarities in the complementarity determining region, antibodies with similar framework regions may not be grouped into a same cluster unless those antibodies also exhibit sufficient similarities in the complementarity determining region.

[0030] At 308, a visual representation is generated based at least on the one or more identified groups of related protein sequences. In some example embodiments, the one or more groups of related protein sequences may be projected into a lower dimensional space, such as a two-dimensional or three-dimensional space, which is suitable for visualization. For example, in some cases, the embeddings of each protein sequence may be further projected into a two- NAI-5007186271V1 51Attorney Ref.: 14786-058-228 (103963-228058) / P39076-WO-1dimensional (or three-dimensional) space in order to generate a visual representation of the population of protein sequences in which the relationship between different protein sequences are indicated by the distance (e.g., Euclidean distance) therebetween. In some cases, the visual representation of the one or more groups of related protein sequences may show the distribution of the one or more groups of related protein sequences in a two-dimensional (or three-dimensional) space. Furthermore, in some cases, each protein sequence may be shown as a visual indicator that is rendered in a shape, color, and / or size representative of one or more properties of the protein sequence. For instance, for antibodies, each antibody in the visual representation may shown as either a first visual indicator representative of a binding antibody, a second visual indicator representative of a weakly-binding antibody, or a third visual indicator representative of a nonbinding antibody.

[0031] FIG. 3B depicts a flowchart illustrating another example of a process 350 for machine learning enabled analysis of protein sequences, in accordance with some example embodiments. In some example embodiments, the process 350 may be performed by the analysis controller 110 including, for example, the preprocessing engine 111, the embedding engine 113, the clustering engine 115, the visualization engine 117, and / or the like. As described in more detail below, in some cases, while the preprocessing engine 111 may translate raw sequencing data (e.g., from the sequencing platform 120) into protein sequences (e.g., sequences of amino acid residues), the clustering engine 115 and the visualization engine 117 may leverage embeddings of these protein sequences generated by the embedding engine 113 in order to identify, analyze, and visualize the relationships therebetween.

[0032] At 352, raw sequencing data may be translated into a plurality of protein sequences. In some example embodiments, raw sequencing data may include one or more coding NAI-5007186271V1 52Attorney Ref.: 14786-058-228 (103963-228058) / P39076-WO-1sequences (or protein-coding genes), which are the portions of a gene’s deoxyribonucleic acid (DNA) or a ribonucleic acid (RNA) sequence that code for a protein. In some cases, the translating of the raw sequencing data may include translating the nucleotide bases in each coding sequence (or protein-coding gene) into a corresponding protein sequence (or sequence of amino acid residues). In some cases, the translating of the raw sequencing data may include multiple operations including, for example, merging forward and reverse reads, translation, quality check, identifying overlaps between sequences (e.g., all-pairs suffix prefix (APSP)), deduplication, alignment, and / or the like. It should be appreciated that the protein sequence that results from the translating of a coding sequence (or protein coding gene) may be a representation of the protein sequence that includes the identity (or type and, in some cases, the relative position (or order), of each amino acid residue in the protein sequence. For example, in some cases, the representation of the protein sequence may include a one-hot encoding of the identity of each constituent amino acid residue and a positional encoding of the location (or position) of each amino acid residue in the protein sequence.

[0033] At 354, an embedding computation model may be applied to generate, for each protein sequence of the plurality of protein sequences, an embedding emphasizing similarities at one or more specific regions. In some example embodiments, an embedding of a protein sequence may be generated to capture the significance of diversity in different regions of the protein sequence. For example, in some cases, the embedding of the protein sequence may be generated by an embedding computation model (e.g., the encoder of a language model (or large language model) with a transformer architecture) that has been trained through contrastive learning to generate similar embeddings for protein sequences that exhibit a threshold degree of similarities in one or more specific regions and dissimilar embeddings for protein sequences that fail to exhibit NAI-5007186271V1 53Attorney Ref.: 14786-058-228 (103963-228058) / P39076-WO-1the threshold degree of similarities in the one or more specific regions. In some cases, in addition to the threshold degree of similarities in the one or more specific regions, the embedding computation model may be further trained to generate similar embeddings when two protein sequences exhibit a threshold degree of similarities across the entire length of each protein sequence. For instance, where the protein sequence is an antibody, the embedding of the protein sequence may be generated to prioritize similarities in the complementarity determining regions (e g., third complementarity determining (CDR3) of the heavy chain) over similarities in the framework regions.

[0034] At 356, an additional embedding generalizing the one or more specific regions may be generated for each protein sequence of the plurality of protein sequences. In some example embodiments, while the embedding of a protein sequence may be generated to emphasize similarities in one or more specific regions (e g., CDR3) of the protein sequence, an additional embedding of a protein sequence may be generated in order to generalize at least a portion of those specific regions. In some cases, a specific portion of the protein sequence may be generalized by at least assigning each constituent amino acid residue to a corresponding amino acid group such that the additional embedding of the protein sequence includes, for an amino acid residues present in the one or more specific regions, an amino acid group to which the amino acid residue belongs instead of the actual identity of the amino acid residue. For example, in instances where the protein sequence is an antibody, the additional embedding may generalize the third complementarity determining region (CDR3) by assigning each constituent amino acid residue to either a non-polar group, a polar group, an aromatic group, a positive-charged (basic) group, a negatively-charged (acidic) group, or a special group. For any alanine (Ala) present in the third complementarity determining region (CDR3) of the protein sequence, for instance, the additional embedding may NAI-5007186271V1 54Attorney Ref.: 14786-058-228 (103963-228058) / P39076-WO-1specify the non-polar group to which alanine (Ala) belongs instead of its actual identity. In some cases, the additional embedding of the protein sequence may be generated based on the fc-mer fragments, or substrings of length k, within the one or more specific regions (e.g., CDR3) of the protein sequence. In some cases, the additional embedding may be generated by subjecting the k-mer fragments to two rounds of dimensionality reduction. For instance, in some cases, the fc-mer fragments may undergo principal component analysis to yield a lower-dimensional representation of the fc-mer fragments before being projected, with Uniform Manifold Approximation and Projection (UMAP) or t-Distributed Stochastic Neighbor Embedding (t-SNE), into a two-dimensional (or three-dimensional) space. The resulting embedding is a two-dimensional (or three-dimensional) representation of the protein sequence that more accurately reflect the similarity present within one or more specific regions (e.g., CDR3).

[0035] At 358, a hybrid embedding may be generated for each protein sequence by integrating the embedding and the additional embedding. In some example embodiments, a hybrid embedding for a protein sequence may be a combination (e.g., union) of an embedding emphasizing the similarities in one or more specific regions of the protein sequence and an additional embedding generalizing the one or more specific regions of the protein sequence. In some cases, the combination (e.g., union) of the embedding and the additional embedding may prioritize the embedding over the additional embedding. For example, in some cases, the embedding may be assigned a higher weight (e.g., a 10 fold higher weight) than the additional embedding when the two embeddings are integrated to generate the hybrid embedding such that the embedding contributes more to the hybrid embedding than the additional embedding. In some cases, the incorporation of the additional embedding, which generalizes the one or more specific regions of the protein sequence, may increase the differentiation the protein sequence and NAI-5007186271V1 55Attorney Ref.: 14786-058-228 (103963-228058) / P39076-WO-1unrelated protein sequences. For instance, in some cases, the boundaries between different groups of related protein sequences, which are described in more detail below, may be refined by the inclusion of the additional embedding.

[0036] At 360, one or more groups of related protein sequences may be identified within the plurality of protein sequences based at least on the hybrid embedding of each protein sequence. In some example embodiments, one or more groups of related protein sequences may be identified within a population of protein sequences, such as the repertoire of antibodies generated in response to pathogen exposure. In some cases, the one or more groups of protein sequences may be identified based on the embedding of each protein sequence. For example, in some cases, the one or more groups of related protein sequences may be identified based on the hybrid embedding of each protein sequence, which integrates an embedding emphasizing similarities in one or more specific regions (e g., CDR3) and an additional embedding generalizing the one or more specific regions. Alternatively, instead of the hybrid embedding, it should be appreciated that the one or more groups of related protein sequences can also be identified based on either the embedding or the additional embedding alone without any integration of the two. In some cases, the one or more groups of related protein sequences may be identified by applying a clustering technique to the embeddings of the protein sequences in the population of protein sequences. In some cases, the clustering technique may be a hierarchical clustering technique, such as Leiden clustering, in order to increase (or maximize) the modularity between the one or more groups of related protein sequences resulting therefrom. However, it should be appreciated that other clustering techniques may also be applied including, for example, centroid-based clustering, density-based clustering, distribution-based clustering, and / or the like. Where the group assignment of a protein sequence is ambiguous, such as when the distance between a protein sequence and two or more clusters of NAI-5007186271V1 56Attorney Ref.: 14786-058-228 (103963-228058) / P39076-WO-1protein sequences fails to satisfy one or more thresholds, the protein sequence may be assigned to a cluster of protein sequences based at least on the similarity between the one or more specific regions (e.g., CDR3) of the protein sequence and those of the protein sequences within the clusters.

[0037] At 362, a visual representation may be generated based at least on the one or more identified groups of related protein sequences. In some example embodiments, the one or more groups of related protein sequences may be projected into a lower-dimensional space, such as a two-dimensional or a three-dimensional space, which is more suitable for visualization. In some cases, the visual representation of the one or more identified groups of related protein sequence may be generated to depict a distribution of the one or more identified groups of related protein sequences within a two-dimensional (or three-dimensional) space. In some cases, the visual representation of the one or more identified groups of related protein sequences may include visual indicators to enable a differentiation between different groups of related protein sequences. For example, in some cases, protein sequences in a first group of related protein sequences may be depicted using a first visual indicator rendered in a different color, shape, and / or size than a second visual indicator depicting protein sequences in a second group of related protein sequences. In some cases, in addition to visual indicators to differentiate between different groups of related protein sequences, the visual representation of the one or more identified groups of related protein sequences may further include visual indicators to enable a differentiation between protein sequences exhibiting different magnitudes of a property and / or different properties. For instance, in some cases, the visual representation of the one or more identified groups of related protein sequences may include a first visual indicator depicting protein sequences that are binders, a second visual indicator depicting protein sequences that are weak binders, and a third visual indicator depicting protein sequences that are non-binders. The first visual indicator, the second NAI-5007186271V1 57Attorney Ref.: 14786-058-228 (103963-228058) / P39076-WO-1visual indicator, and the third visual indicator may be rendered in different colors, shapes, and / or sizes in order to enable a differentiation of binding, weakly-binding, and non-binding protein sequences across the population of protein sequences. In some cases, the visual representation of the one or more groups of related protein sequences may depict the distribution of different properties and / or different magnitudes of a property across the one or more groups of related protein sequences. Examples of properties include expression, binding affinity towards another molecule (e.g., a viral antigen, a tumor antigen, and / or the like), specificity, stability, immunogenicity, human-ness, self-association, and / or the like. In the context of antibodies, for instance, the visual representation of the one or more groups of related protein sequences may depict the distribution of binding, weakly binding, and non-binding protein sequences across the one or more groups of related protein sequences. As described in more detail below, the visual representation generated based on various embodiments of the embeddings described herein may enable a clearer differentiation between groups of related protein sequences and the distribution of different properties and / or different magnitudes of a property therein.

[0038] In some cases, the visual representation of the one or more groups of related protein sequences may be generated as a part of a user interface, such as a graphic user interface (GUI), at a client device. In some cases, the user interface may be an interactive user interface to support configuration of the visual representation. For example, in some cases, one or more user inputs received via the user interface may specify one or more thresholds (e.g., minimum, maximum, and / or the like) on the quantity of protein sequences in each group of related protein sequences. In some cases, the one or more user inputs received via the user interface may specify a selection of properties that are included in the visual representation of the one or more groups of related protein sequences. For instance, in some cases, upon receiving one or more user inputsNAI-5007186271V1 58Attorney Ref.: 14786-058-228 (103963-228058) / P39076-WO-1selecting expression instead of binding affinity, the visual representation of the one or more groups of related protein sequences may be updated to depict the distribution of high-, low-, and nonexpression protein sequences across the one or more groups of related protein sequences. FIG. 4 depicts a flowchart illustrating an example of a process 400 for context-aware generation of novel protein sequences, in accordance with some example embodiments. Referring to FIGS. 1 and 4, in some case, the process 400 may be performed by the analysis controller 110 including, for example, the generative model 119. For example, in some cases, the process 400 may be performed to train (or pretrain) and finetune (or post train) the generative model 119 to perform context-aware generation of novel protein sequences. In some cases, the generative model 119 may be trained (or pretrained) on a fill-in-the-middle (FIM) objective before undergoing direct preference optimization (DPO). In some cases, the trained and finetuned generative model 119 may be applied to generate novel sequences based on a generative context that includes a group of related protein sequences identified based on the embeddings generated, for example, by the embedding computation model 114.

[0039] At 402, a generative model is trained on a fill-in-the-middle (FIM) objective. In some example embodiments, the generative model may be a language model with a deep learning architecture (e.g., recurrent neural network (RNN) and / or the like) adapted to modeling sequences, such as sequences of amino acid residues, with long range dependencies. In some cases, training the generative model on the fill-in-the-middle (FIM) objective may include applying the generative model to determine the identities of a continuous span of amino acid residues in a set of protein sequences. In some cases, the set of protein sequences may include a group of related protein sequences identified based on embeddings that emphasize similarities within certain regions of each protein sequence (e.g., CDR3 in the case of antibodies) with greater impact to one or moreNAI-5007186271V1 59Attorney Ref.: 14786-058-228 (103963-228058) / P39076-WO-1properties of interest (e.g., binding affinity and / or the like). In some cases, the protein sequences in the set of protein sequences may be concatenated, for example, with a special character (or token) designating the separation between successive protein sequences. In some cases, the identities of the individual amino acid residues in the continuous span of amino acid residues may be masked and, in some cases, an unmasked copy of those amino acid residues may be appended to the end of the concatenated protein sequences. In some cases, the generative model may be trained to determine the identities of the masked amino acid residues based on a context (or generative context) around the masked amino acid residues. For example, in some cases, the generative context around the masked amino acid residues may include one or more of the amino acid residues preceding or subsequent to the masked amino acid residues in the concatenated protein sequences. In some cases, the training of the generative model may include adjusting one or more parameters (e.g., weights, biases, and / or the like) of the generative model such that the identities of the masked amino acid residues determined by the generative model match the ground-truth identities of these amino acid residues. For instance, in some cases, the generative model may be trained to reduce (or minimize) a cross-entropy loss between the identities of the masked amino acid residues determined by the generative model and the corresponding groundtruth identities of these amino acid residues. In some cases, the cross-entropy loss may quantify the error (or discrepancy) present in the output of the generative model (e.g., the difference between the identities of the masked amino acid residues determined by the generative model and the corresponding ground-truth identities). In some cases, the cross-entropy loss may be evaluated as perplexity (e.g., the exponential of the cross-entropy loss) in order to quantify the uncertainty in the output of the generative model.NAI-5007186271V1 60Attorney Ref.: 14786-058-228 (103963-228058) / P39076-WO-1

[0040] At 404, the generative model is finetuned with direct preference optimization (DPO). In some example embodiments, the generative model may undergo finetuning (or post training) with direct preference optimization (DPO) in order for the generative model to generate novel protein sequences that evolve from a region of interest (e.g., a specific subset of protein sequences) in the generative context. In the absence of direct preference optimization (DPO) finetuning (or post training), the novel protein sequences generated by the generative model may merely reflect a consensus (e.g., average) across the protein sequences in the generative context. Where the generative context include related protein sequences that have evolved over successive generations to exhibit better properties (e.g., children antibodies that exhibit better binding affinity than parent antibodies as a results of affinity maturation), direct preference optimization (DPO) may train the generative model to leverage the hierarchy of parent and children protein sequences that results from this evolution. For example, in some cases, instead of the generative model generating novel protein sequences that merely reflect a consensus (e.g., average) across the entire generative context, the generative model may be finetuned (or post trained) to generate novel protein sequences that further evolve from a particular region of interest in the generative context. In some cases, the region of interest may be populated by a subset of protein sequences with better properties of interest (e.g., binding affinity) than other protein sequences in the generative context. For instance, in the case of antibodies, the region of interest may be populated by antibodies with greater distant to the germline, greater edit distance relative to the antibodies in the initial naive antibody repertoire, higher binding affinity towards a target antigen, and / or lower liabilities than other antibodies in the generative context. In instances where the generative context is formed by embeddings generated by an embedding computation model trained on contrastive loss and finetuned (or post trained) on hierarchical loss, the region of interest may be populated by childrenNAI-5007186271V1 61Attorney Ref.: 14786-058-228 (103963-228058) / P39076-WO-1protein sequences exhibiting better properties than prior generations of parent protein sequences from which they descend.

[0041] In some cases, each training sample used for post training (or finetuning) the generative model may be a tuple including a generative context, a positive sample, and a negative sample. In some cases, the positive sample may be a protein sequence that is related to a protein sequence in a region of interest of the generative context, such as being a child protein sequence of a parent protein sequence in the region of interest. Furthermore, in some cases, the negative sample may be a protein sequence that is unrelated to any protein sequences in the region of interest. In some cases, the generative model may be finetuned (or post trained) to increase (or maximize) the likelihood of positive samples in the distribution of the output of the generative model while decreasing (or minimizing) the likelihood of negative samples in the output distribution of the generative model. In some cases, the generative model may be finetuned (or post trained) such that the likelihood of positive samples in its output distribution exceeds the likelihood of negative samples. For example, in some cases, the generative model may be trained by further adjusting one or more parameters (e.g., weights, biases, and / or the like) such that the generative model is more likely to output a positive sample (e.g., a novel protein sequence that descends from a protein sequence in the region of interest in the generative context) than a negative sample (e.g., a novel protein sequence that is unrelated to a protein sequence in the region of interest in the generative context).

[0042] Equation (1) below is an example of a loss function for finetuning (or post training) the generative model for direct preference optimization (DPO).-Corots. Href) = - [10g O' ( / ? log^g^ - Plog^g^ “)] <0NAI-5007186271V1 62Attorney Ref.: 14786-058-228 (103963-228058) / P39076-WO-1wherein 7Tref denotes the generative model trained (or pretrained) on the fill-in-the-middle (FIM) objective, n0denotes the generative model finetuned (or post trained) with direct preferenceoptimization (DPO), wherein log js the likelihood of positive samples yw, and "refCywK.)ft logis thelikelihood of negative samplesIn Equation (1), logmay act asa regularization term to prevent the parameters of the finetuned (or post trained) generative model nefrom undergoing excessive updates from those of the original trained (or pretrained) generative model 7Tref.

[0043] In some example embodiments, instead of finetuning (or post training) the generative model to increase (or maximize) the likelihood of the generative model outputting positive samples and reduce (or minimize) the likelihood of the generative model outputting negative samples, the generative model may be trained for weighted direct preference optimization (DPO) such that the probability distribution of its output corresponds to the distribution of one or more properties of interest. For example, in some cases, the generative model may be trained such that the likelihood of the generative model outputting a particular antibody is proportional to the distance of the antibody to its germline. Accordingly, given three sequences and their corresponding distances to the germline (SI: 15, S2: 10, S3: 1), the generative model may be trained to output the first sequence SI with a greater likelihood than both the second sequence S2 and the third sequence S3 while the second sequence S2 may be output with a greater likelihood than the third sequence S3. Equation (2) is an alternate loss function for finetuning (or post training) the generative model for weighted direct preference optimization (DPO).L P* = -E^ ly -w^ gl1o0ggpPr9ef((yyWW|x|x))- l ioggZyt■ expp f^g lioggpPr9ef((yy(°W|x|x)i]j (1)NAI-5007186271V1 63Attorney Ref.: 14786-058-228 (103963-228058) / P39076-WO-1wherein w denotes the weights for weighing each individual sample based on their respective properties of interest (e.g., distance to germline, edit distance to naive repertoire, binding affinity, liabilities, and / or the like).

[0044] At 406, a generative context including a group of related protein sequences having a threshold similarity in one or more specific regions is determined based at least on a plurality of embeddings generated by an embedding computation model. In some example embodiments, the embedding computation model may be trained (or pretrained), for example, through contrastive loss, to generate embeddings that emphasize similarities in certain regions of the underlying protein sequences (e.g., CDR3 in the case of antibodies). As such, in some cases, the embeddings generated by the embedding computation model may capture the significance of diversity (or conservation) in different regions of protein sequences. Accordingly, the group of related protein sequences identified based on the corresponding embeddings may exhibit a threshold similarity in specific regions (e.g., CDR3 in the case of antibodies) rather than similarities in any indiscriminate region. In some cases, the embedding computation model may be further finetuned (or post trained) to reduce (or minimize) a hierarchical loss such that the embeddings generated by the embedding computation model captures the hierarchical relationships that may be present within the group of related protein sequences. For example, in some cases, the embeddings generated by the embedding computation model may capture the parent-child relationship that may exist between protein sequences, such as successive generations of antibodies that evolve through affinity maturation. As described in more detail below, the protein sequences forming the generative context may guide the generation of novel protein sequences by the generative model.NAI-5007186271V1 64Attorney Ref.: 14786-058-228 (103963-228058) / P39076-WO-1

[0045] At 408, the generative model is applied to generate, based at least on the generative context, one or more novel protein sequences. In some example embodiments, the generative model, which has been finetuned with direct preference optimization, may generate novel protein sequences that evolve from one or more regions of interest within the generative context. For example, in some cases, the generative model may generate novel protein sequences that evolve from children protein sequences that are a threshold distance from the germline or exhibit a threshold edit distance from the naive repertoire. In some cases, the protein sequences in the region of interest may exhibit certain properties of interest, such as a threshold binding affinity towards a target antigen. In some cases, the novel protein sequences generated by the generative model to further evolve the protein sequences in the region of interest may exhibit better properties than the protein sequences in the region of interest.

[0046] FIG. 5 depicts a schematic diagram illustrating an example of a workflow 500 for machine learning enabled analysis of protein sequences, in accordance with some example embodiments. Referring to FIG. 5, in some cases, an encoder 503 may be applied to generate, for each protein sequence in a population of protein sequences 502, an embedding 504. In some cases, the population of protein sequences 502 may be a repertoire of antibodies generated in response to pathogen exposure, such as during an immunization campaign. Alternatively, in some cases, the population of protein sequences 502 may originate from a synthetic library or are the products of in silica design efforts. In some cases, instead of the protein sequences 502, the encoder 503 may ingest and operate on nucleic acid sequences (e.g., deoxyribonucleic acid (DNA) sequences), ribonucleic acid (RNA), and / or the like), such as the coding sequences (or protein-coding genes) specifying the sequence of amino acid residue in the polypeptide chain of protein molecules. In some cases, the encoder 503 may be the encoder of a language model having a transformerNAI-5007186271V1 65Attorney Ref.: 14786-058-228 (103963-228058) / P39076-WO-1architecture. In some cases, the encoder 503 may have been trained, through contrastive learning, to generate similar embeddings for protein sequences exhibiting a threshold degree of similarity in one or more specific regions (e.g., CDR3) of each protein sequence and dissimilar embeddings for protein sequences that fail to exhibit the threshold degree of similarities in the one or more specific regions. In some cases, the encoder 503 may have been pretrained to recover the identity of one or more masked amino acid residues in protein sequences before undergoing contrastive learning to generate similar embeddings for protein sequences exhibiting a threshold degree of similarity in one or more specific regions (e.g., CDR3) of each protein sequence and dissimilar embeddings for protein sequences that fail to exhibit the threshold degree of similarities in the one or more specific regions. In some cases, the embedding 504 that is generated by the encoder 503 may emphasize similarities in the one or more specific regions of each protein sequence from the population of protein sequences 502.

[0047] Referring again to FIG. 5, in some cases, an additional embedding 506 may be generated for each protein sequence within the population of protein sequences 502. As described in more detail below, in some cases, the additional embedding 506 may generalize the one or more specific regions of each protein sequence from the population of protein sequences 502. In some cases, a hybrid embedding 508 may be generated for each protein sequence from the population of protein sequences 502 by at least combining the embedding 504 and the additional embedding 506. In some cases, one or more groups of related protein sequences may be identified within the population of protein sequences 502 by at least applying a clustering technique 509, such as a hierarchical clustering technique (e.g., Leiden clustering and / or the like), to the hybrid embedding 508 of each protein sequence in the population of protein sequences 502. In some cases, the application of the clustering technique 509 may result in the assignment of a cluster label 510 toNAI-5007186271V1 66Attorney Ref.: 14786-058-228 (103963-228058) / P39076-WO-1each protein sequence in the population of protein sequences 502. Furthermore, some cases, a visual representation 512 of the one or more groups of related protein sequences may be generated by at least projecting the one or more groups of related protein sequences to a two-dimensional (or three-dimensional) space.

[0048] FIG. 6 depicts a schematic diagram illustrating an example of a workflow for generating embeddings occupying a two-dimensional Euclidean space, in accordance with some example embodiments. Referring to FIG. 6, in some example embodiments, an embedding computation model, such as the embedding computation model 114 shown in FIG. 1, may be trained to generate embeddings of protein sequences that occupy a two-dimensional Euclidean space. In the example shown in FIG. 6, each protein sequence (e.g., of N protein sequence) may be tokenized, for example, to encode the identities (or types) and relative positions of the constituent amino acid residues. In some cases, the embedding computation model 114 may ingest the tokenized representation of a protein sequence and output an embedding of a certain dimensionality (e.g., 256-dimensions in the example shown in FIG. 6). In some cases, the embedding output by the embedding computation model 114 may be further operated upon by one or more linear layers, such as a linear layer 610 followed by a linear subsampling layer 615. In the example shown in FIG. 6, the linear layer 610 may ingest the embedding output by the embedding computation model 114 and produce an output having the same number of dimensions (e g., 256-dimensions in the example shown in FIG. 6). In some cases, a projection layer, such as the linear subsampling layer 615 shown in FIG. 6, may then map the output of the linear layer 610 from its original higher dimensional space (e.g., 256-dimensions) to a lower dimensional space. In the example shown in FIG. 6, the linear subsampling layer 615 may map the output of the linearNAI-5007186271V1 67Attorney Ref.: 14786-058-228 (103963-228058) / P39076-WO-1layer 610 to a two-dimensional Euclidean space 620 in which the distance between embeddings reflect the extent of similarities in one or more specific regions of the underlying protein sequences.

[0049] FIG. 7A depicts a schematic diagram illustrating an example of a workflow for generating embeddings occupying a hyperbolic space, in accordance with some example embodiments. Referring to FIG. 7A, in some cases, an embedding computation model, such as the embedding computation model 114 shown in FIG. 1, may be trained to generate embeddings of protein sequences that occupy a hyperbolic space (e.g., a Lorentz manifold) that better captures the hierarchical relationships that may exist between protein sequences. In the example shown in FIG. 7A, each protein sequence (e.g., of N protein sequence) may be tokenized, for example, to encode the identities (or types) and relative positions of the constituent amino acid residues. In some cases, the embedding computation model 114 may ingest the tokenized representation of a protein sequence and output an embedding of a certain dimensionality (e.g., 256-dimensions in the example shown in FIG. 6), which may be further operated upon by one or more linear layers, such as the linear layer 610 and the linear subsampling layer 615. In the example shown in FIG. 7A, the linear layer 610 may ingest the embedding output by the embedding computation model 114 and produce an output having the same number of dimensions (e.g., 256-dimensions in the example shown in FIG. 7A). In some cases, the linear subsampling layer 615 may subsequently map the output of the linear layer 610 from its original higher dimensional space (e.g., 256-dimensions) to a lower dimensional space, such as a three-dimensional space. In the example shown in FIG. 7A, the output of the linear subsampling layer 615 may be further mapped, for example, by an exponential map 710, to a hyperbolic space 715 (e.g., a Lorentz manifold and / or the like). For example, in some cases, the exponential map 710 may be a function mapping a tangent vector from a point's tangent space in two-dimensional Euclidean space to a point on theNAI-5007186271V1 68Attorney Ref.: 14786-058-228 (103963-228058) / P39076-WO-1manifold of the hyperbolic space 715. In some cases, the exponential map 710 may perform the mapping by at least traveling along a unique geodesic (or a straight line) in the direction of the vector, for example, one unit at a time. In some cases, mapping a point and a tangent vector to a new point by traveling along a geodesic (or a straight line) may serve to move from the "flat" tangent space of the two-dimensional Euclidean space to the curved hyperbolic space 715 (e.g., the Lorentz manifold and / or the like).

[0050] To further illustrate, Equation (3) below is an example of the exponential map 710.(r\ sinh(-L)■J=j O -I - r^-V (3)wherein O denotes the origin of the exponential.

[0051] Equation (4) below is an example of a distance function quantifying the distance between two embeddings x and y in a hyperbolic space (e g., a Lorentz manifold and / or the like) •, iwith a curvature of —.kd(x, y) = Vk arccosh ( )wherein x,y)His the inner product of the two embeddings x and y in the hyperbolic space (e.g., the Lorentz manifold and / or the like).

[0052] In some cases, the inner product (x,y)His the inner product of the two embeddings x and in the hyperbolic space (e.g., the Lorentz manifold and / or the like) may be computed in accordance to Equation (5).(x, y)H= ~xoyo+ Xi=i xiyt (5)NAI-5007186271V1 69Attorney Ref.: 14786-058-228 (103963-228058) / P39076-WO-1

[0053] Equation (6) below is an example of a contrastive loss function for a single “anchor” protein sequence and multiple positive protein sequences exhibiting a threshold degree of similarity in one or more specific regions as the anchor protein sequence. / expf-^^) \ ^Pilog; —(6)1 ll\X / ^(exp( —

[0054] Equation (7) below is an example of a function converting three-dimensional hyperbolic (or Lorentz) coordinates (t, x, y) to the corresponding two-dimensional (or Poincare)coordinates (u, v) with a curvature of —= <7’

[0055] FIG. 7B depicts a schematic diagram illustrating another example of a workflow in which an embedding computation model, such as the embedding computation model 114 shown in FIG. 1, is trained to generate embeddings that populate the hyperbolic space 715 (e.g., a Lorentz manifold. For example, as shown in FIG. 7B, the embedding computation model 114, which may be the encoder from a language model (or large language model) having a transformer architecture, may first undergo n-dimensional hyperbolic contrastive learning before being finetuned (or post trained) on a hierarchical loss (e.g., tree loss, angle-based entailment loss, and / or the like). In some cases, the n-dimensional hyperbolic contrastive learning may include training the embedding computation model 114 to generate, for a positive pair of protein sequences exhibiting a threshold similarity in one or more specific regions, similar embeddings that are proximately located to one another in the hyperbolic space 715. Furthermore, in some cases, the n-dimensional hyperbolic contrastive learning may include training the embedding computation model to generate, for a negative pair of protein sequences that fail to exhibit the threshold similarity in the one or moreNAI-5007186271V1 7AAttorney Ref.: 14786-058-228 (103963-228058) / P39076-WO-1specific regions, dissimilar embeddings that are distantly located to one another in the hyperbolic space 715.

[0056] In some example embodiments, the embedding computation model may be further finetuned (or post trained) subsequent to undergoing n -dimensional hyperbolic contrastive learning. For example, FIG. 7B shows the embedding computation model 114 being finetuned (or post trained) on an angle-based entailment loss, a type of hierarchical loss that geometrically quantifies the hierarchical relationship between the embeddings of protein sequences generated by the embedding computation model 114. In some cases, a hierarchy of protein sequences may be defined to include similar protein sequences (e.g., (e.g., k -nearest neighbors (kNN) in the hyperbolic space 715) that are also related, for example, with sufficient entailment (e.g., a threshold variable (V) gene and joining (J) gene entailment score in the case of antibodies). In some cases, a positive pair of protein sequences may include two related protein sequences (e.g., a parent protein sequence and a child protein sequence descending therefrom) while a negative pair of protein sequences may include two unrelated protein sequences. In some cases, the finetuning (or post training) of the embedding computation model 114 may include adjusting or more parameters (e.g., weights, biases, and / or the like) of the embedding computation model 114 to increase (or maximize) the angle between the embeddings of positive pairs of protein sequences and reduce (or minimize) the angle between the embeddings of negative pairs of protein sequences. In some cases, doing so may be geometrically tantamount to training the embedding computation model 114 to generate embeddings of children protein sequences to fall within a geometric cone beneath the embeddings of the parent protein sequences while the embeddings of unrelated protein sequences fall outside of the geometric cone. It should be appreciated that subjecting the embedding computation model 114 to contrastive learning may train the embedding computationNAI-5007186271V1 71Attorney Ref.: 14786-058-228 (103963-228058) / P39076-WO-1model 114 to learn a global clustering structure of similar embeddings while the subsequent finetuning (or post training) on the hierarchical loss may further train the embedding computation model 114 to learn an internal hierarchy (e.g., parent-child hierarchy) present within the clusters of similar embeddings.

[0057] FIG. 7C depicts two-dimensional visualizations of the embeddings of antibodies that illustrate the biological significance captured by the embeddings, in accordance with some example embodiments. As shown in FIG. 7C, a cluster of interest may be identified, for example, to include a group of related antibodies exhibiting a threshold similarity in one or more specific regions (e.g., CDR3). As noted, in some cases, the embedding computation model 114 may be trained to generate embeddings whose relative locations in latent space (e.g., a hyperbolic space such as a Lorentz manifold and / or the like) capture both similarities in the one or more specific regions and the hierarchical relationship between parent and children antibodies descending therefrom. As shown in FIG. 7C, the resulting embeddings may form a conical cluster exhibiting a distinct improvement in certain properties of interest (e g., increase in binding affinity towards a target antigen) along the same direction of increasing distance to germline, as quantified by the quantity of somatic hypermutations (SHMs) present in the heavy chain variable region (VH) of each antibody. That is, as shown in FIG. 7C, the embeddings of later generations of antibodies, which exhibit a larger quantity of somatic hypermutations (SHMs) than preceding generations, also exhibits improved properties of interest such as higher binding affinity towards the target antigen. As described in more detail below, the hierarchical relationship between parent and children protein sequences may be leveraged as generative context to guide the generation of novel protein sequences that further evolve the germline to improve upon the properties of interest (e.g., binding affinity towards the target antigen).NAI-5007186271V1 72Attorney Ref.: 14786-058-228 (103963-228058) / P39076-WO-1

[0058] FIG. 8 depicts a graph 800 illustrating an example of hierarchical loss, in accordance with some example embodiments. In some example embodiments, an embedding computation model, such as the embedding computation model 114 shown in FIG. 1, may be trained (or pretrained) with contrastive learning (e.g., n -dimensional hyperbolic contrastive learning) before being finetuned (or post trained) to reduce (or minimize) a hierarchical loss (e.g., tree loss, angle-based entailment loss, and / or the like) that geometrically quantifies the parent and child relationships that exist between protein sequences. For example, in some cases, the embedding computation model 114 may be finetuned (or post trained) on a batch D of protein sequences (e.g., antibodies). In some cases, the hierarchical loss associated with a protein sequence i from the batch D may be a bidirectional loss that includes a parent loss Lp^cassociated with the protein sequence i being a parent protein sequence. In the case of the parent loss Lp^cof the protein sequence i, the positive pairs of the protein sequence i may include the children protein sequence of the protein sequence i in the batch D. In some cases, the reduction (or minimization) of the parent loss Lp^cmay train the embedding computation model 114 to generate the embeddings of the children protein sequences of the protein sequence i to lie on one side of the embedding of the protein sequence i. For instance, in some cases, the embedding computation model 114 may be trained such that the embeddings of those children protein sequences form a geometric cone on one side of the protein sequence i opposite the embeddings of its parent protein sequences. In some cases, the bidirectional loss may further include a child loss Lc^passociated with the protein sequence i being a child protein sequence. For the child loss Lc^pof the protein sequence i, the positive pairs of the protein sequence i may include the parent protein sequences of the protein sequence i in the batch D. In some cases, the reduction (or minimization) of the child loss Lc^pmay train the embedding computation model 114 to generate the embeddings of NAI-5007186271V1 73Attorney Ref.: 14786-058-228 (103963-228058) / P39076-WO-1the parent protein sequences of the protein sequence i to lie in an opposite direction of the embeddings of its children sequences. Accordingly, if the embeddings of the children protein sequences are pulled (e.g., to form a geometric cone) beneath the embedding of the protein sequence i, the embeddings of the parent protein sequences are pushed ahead of the embedding of the protein sequence i. In some cases, the embedding of the protein sequence i itself may be further pulled beneath the embeddings of its parent protein sequences. Furthermore, in some cases, the embeddings of unrelated protein sequences (e.g., non-parent and non-children protein sequences) may be pushed away from the embedding of the protein sequence i.

[0059] Equation (8) below is an example of a loss function quantifying the angle loss Langle f°rthe batch D of protein sequences used as training data for finetuning (or post training) the embedding computation model 114.^angleC^ref) = L^C(. D, + Lc^(2, t2) + Lreg(D, 7Tref) wherein denotes the angle between the embedding of the embedding of the protein sequence i and its children protein sequences, a2denotes the angle between the embedding of the embedding of the protein sequence i and its parent protein sequences, and Lregdenotes a regularization term that penalizes embeddings that drift too far from their original positions in order to preserve the global cluster structure learned by the embedding computation model 114 undergoing contrastive learning.

[0060] Equation (9) below is an example formulation for computing parent loss Lp^c.LP^C( ), K) = - S(x. yi)e2)log (9)exp(K(X^,yi'l)+Zy- EM, exp^’Y1wherein K denotes a similarity metric for a pair of protein sequences.NAI-5007186271V1 74Attorney Ref.: 14786-058-228 (103963-228058) / P39076-WO-1

[0061] Equation (10) below is an example formulation for computing the similarity metric K for embeddings of the protein sequence i and its child protein sequence.K(X yi) = j8i(Xi,yi) = 7i - ext(x yi) (10)

[0062] Equation (11) below is an example formulation for computing the similarity metric K for the embeddings of the protein sequence i and its parent protein sequence.K(Yi<xi) = = ext(Yi,xf) (11)

[0063] In Equations (10) and (11), ext(y, x) denotes the external angle that is formed by a first line from an origin (e g., the point (0,0) in two-dimensional space) to a point corresponding to the embedding x of a parent protein sequence and the embedding y of the child protein sequence. Equation (12) below is an example formulation for computing the exterior angle ext(y,x).

[0064] ext(x,y) = COS’1f „ ytime+XtimeCCx^)\||xspace || V(c(x,y)H)2-l / (12)The effects of finetuning (or post training) an embedding computation model, such as the embedding computation model 114 shown in FIG. 1, are depicted in FIGS. 9A-9B. In FIG. 9A, the embeddings occupying the Poincare disk were generated by an embedding computation model that has undergone finetuning (or post training) on a hierarchical loss (e.g., a tree loss, an angle-based entailment loss, and / or the like), in addition to training (or pretraining) on a contrastive loss. Contrastingly, FIG. 9B depicts embeddings that were generated by an embedding computation model that has not undergone finetuning (or post training) on the hierarchical loss. As noted, training (or pretraining) an embedding computation model on a contrastive loss may train the embedding computation model to generate embeddings whose relative distance in the latent space (e.g., Poincare disk in FIGS. 9A-9B) reflect the similarities that are present in one or more specific regions of the underlying protein sequences. In some cases, NAI-5007186271V1 75Attorney Ref.: 14786-058-228 (103963-228058) / P39076-WO-1further finetuning (or post training) the embedding computation model on the hierarchical loss may train the embedding computation model to generate embeddings whose relative positions in the latent space (e.g., Poincare disk in FIGS. 9A-9B) reflect the hierarchical relationships (e.g., parent-child relationship) that may exist between the underlying protein sequences. The effect of the post finetuning (or post training) is evident in a comparison of FIGS. 9A-9B. Notably, the embeddings in FIG. 9A form not only clusters that group protein sequences with sufficient similarities in one or more specific regions (e.g., CDR3 in the case of antibodies) but those clusters exhibit a hierarchical structure corresponding to the parent-child relationships that exist between the protein sequences within each cluster. Despite the presence of clusters in FIG. 9B, the hierarchical strucutre that is present in FIG. 9A are absent from the clusters in FIG. 9B.

[0065] FIG. 10A depicts a schematic diagram illustrating the output of a generative model without finetuning for direct preference optimization (DPO), in accordance with some example embodiments. As noted, in some example embodiments, a generative model (e.g., the generative model 119 in FIG. 1) may be trained (or pretrained) on a fill-in-the-middle (FIM) objective to generate novel protein sequences based on a generative context that includes, in some cases, related protein sequences identified based on embeddings generated by an embedding computation model (e.g., the embedding computation model 114 in FIG. 1). In some cases, the embedding computation model may be trained (or pretrained) to generate embeddings that emphasize similarities in one or more specific regions of the underlying protein sequences, such as the third complementarity determining region (CDR3) of antibodies. Furthermore, in some cases, the embedding computation model may be finetuned (or post trained) to generate embeddings that capture the hierarchical relationships that may exist between protein sequences in which subsequent generations of child protein sequences exhibit incrementally better propertiesNAI-5007186271V1 76Attorney Ref.: 14786-058-228 (103963-228058) / P39076-WO-1than the parent protein sequences from which they descend. For example, in the case of a naive antibody repertoire undergoing affinity maturation upon pathogen exposure, child antibodies that are more distant from the germline or with greater edit distance relative to the naive repertoire may exhibit higher binding affinity and fewer liabilities than their parent antibodies. In some cases, this evolutionary trend may be leveraged to generate novel protein sequences that evolve from child protein sequences to further improve upon one or more properties of interest, such as binding affinity, liabilities, and / or the like. However, in instances where the generative model is not finetuned (or post trained) for direct preference optimization (DPO), the generative model may be incapable of leveraging these hierarchical relationships when generating novel protein sequences. This deficiency is illustrated in FIG. 10A in which the generative model pretrained on the fill-in-the-middle (FIM) objective generates novel protein sequences that reflect a consensus (e.g., average) across the protein sequences forming the generative context rather than further evolving a specific subset of protein sequences, such as child protein sequences in a particular region of interest. As such, the novel protein sequences may merely exhibit similar properties as the protein sequences in the generative context but not necessarily better properties.

[0066] FIG. 10B depicts a schematic diagram illustrating the output of a generative model with finetuning for direct preference optimization (DPO), in accordance with some example embodiments. In some cases, further finetuning (or post training) the generative model with direct preference optimization (DPO) may enable the generative model to leverage the hierarchical relationships present in the generative context when generating novel protein sequences. For example, as shown in FIG. 10B, the generative model finetuned (or post trained) with direct preference optimization (DPO) may generate novel protein sequences that evolve from a region of interest occupied by a subset of protein sequences in the generative context. In some cases, theNAI-5007186271V1 77Attorney Ref.: 14786-058-228 (103963-228058) / P39076-WO-1subset of protein sequences may be children protein sequences with high distances from the germline and better properties than the other protein sequences, including their parent protein sequences, that lie outside of the region of interest. As shown in FIG. 10B, the generative model may be trained with tuples that include, in addition to the generative context, a positive sample of a protein sequence that evolved from the protein sequences in the region of interest and a negative sample of a protein sequence that did not evolve from the protein sequences in the region of interest. In some cases, the generative model finetuned (or post trained) with direct preference optimization (DPO) may generate novel protein sequences that further improve the properties of the protein sequences in the region of interest.

[0067] FIG. 11 depicts a schematic diagram illustrating an example of a generalization scheme 1100 for generating an embedding of a protein sequence that generalizes one or more specific regions of the protein sequence, in accordance with some example embodiments. The example of the generalization scheme 1100 shown in FIG. 11 assigns each one of the twenty canonical amino acid residues to one of six amino acid groups. For example, as shown in FIG. 11, each one of the twenty canonical amino acid residues is assigned to either a non-polar group, a polar group, an aromatic group, a positive-charged (basic) group, a negatively-charged (acidic) group, and a special group. To generate an embedding of the protein sequence that includes a generalization of one or more specific regions of the protein sequence, each amino acid residue in the one or more specific regions of the protein sequence may be replaced with a corresponding amino acid residue group. For instance, to generate a generalized third complementarity determining region (CDR3) 1115, each amino acid residue in the third complementarity determining region (CDR3) 110 may be replaced with a corresponding amino acid residue group. In the example shown in FIG. 5, the alanine (A), valine (V), and isoleucine (I) in the thirdNAI-5007186271V1 78Attorney Ref.: 14786-058-228 (103963-228058) / P39076-WO-1complementarity determining region 1110 are replaced with the corresponding non-polar (H) group in the generalized third complementarity determining region 1115 while the tyrosine (Y) and tryptophan (W) are replaced with the corresponding aromatic (A) group.

[0068] FIG. 12A depicts a screenshot illustrating an example of a user interface 1200, in accordance with some example embodiments. The example of the user interface 1200 shown in FIG. 12A depicts a visual representation of one or more groups of related protein sequences. For example, in FIG. 12A, each dot in the user interface 1200 may represent a different protein sequence within a population of protein sequences, such as the repertoire of antibodies generated in response to pathogen exposure. In some cases, different groups of related protein sequences are depicted as different clusters of dots. In some cases, the dots may be depicted in different colors to identify protein sequences exhibiting different properties and / or different magnitudes of a property. In the example shown in FIG. 12A, a dot in a first color represents a protein sequence that is a binder, a dot in a second color represents a protein sequence that is a weak binder, and a dot in a third color represents a protein sequence that is a non-binder. The distribution of different colored dots across one or more clusters may indicate the distribution of binders, weak binders, and non-binders across different groups of related protein sequences. In the example shown in FIG. 12A, the population of protein sequences includes a first group of related protein sequences that are binders, a second group of related protein sequences that are weak binders, and a third group of related protein sequences that are non-binders.

[0069] FIG. 12B depicts a screenshot illustrating another example of a user interface 1210, in accordance with some example embodiments. The example of the user interface 1210 shown in FIG. 12B also depicts a visual representation of one or more groups of related protein sequences. In this example, the dots may be depicted in different colors corresponding to theNAI-5007186271V1 79Attorney Ref.: 14786-058-228 (103963-228058) / P39076-WO-1quantity of neighboring protein sequences (e.g., log(number of neighbors)). For example, a dot in a first color represents a protein sequence with a log(number of neighbors) of 0.6, a dot in a second color represents a protein sequence with a log(number of neighbors) of 1.2, a dot in a third color represents a protein sequence with a log(number of neighbors) of 1.5, a dot in a fourth color represents a protein sequence with a log(number of neighbors) of 1.8, and a dot in a fifth color represents a protein sequence with a log(number of neighbors) of 0.9. The distribution of different colored dots across one or more clusters may indicate the distribution of protein sequences with different numbers of neighboring protein sequences across the one or more groups of related protein sequences. A protein sequence in a denser cluster may have a higher number of neighboring protein sequences than a protein sequence in a less dense cluster.

[0070] FIG. 12C depicts a screenshot illustrating an example of a user interface 1220, in accordance with some example embodiments. The example of the user interface 1220 shown in FIG. 12C depicts a visual representation of one or more groups of related protein sequences in which the color of each dot is representative of the charge (e g., net charge) of the corresponding protein sequence. For example, a dot in a first color represents a protein sequence with a charge of -7.5, a dot in a second color represents a protein sequence with a charge of -2.5, a dot in a third color represents a protein sequence with a charge of 2.5, a dot in a fourth color represents a protein sequence with a charge of 5.0, a dot in a fifth color represents a protein sequence with a charge of -5.0, and a dot in a sixth color represents a protein sequence with a charge of 0.0. The distribution of different colored dots across one or more clusters may indicate the distribution of protein sequences with different charge (e.g., net charge) across the one or more groups of related protein sequencesNAI-5007186271V1 80Attorney Ref.: 14786-058-228 (103963-228058) / P39076-WO-1

[0071] FIG. 13 A depicts visual representations of groups of related protein sequences identified by imposing a different threshold for edit distance between two protein sequences, in accordance with some example embodiments. In some cases, the identification of groups of related protein sequences may be calibrated by at least setting one or more thresholds for edit distance between the embeddings of two or more protein sequences. For example, in some cases, clusters of related protein sequences within a population of protein sequences, such as a repertoire of antibodies generated in response to pathogen exposure, may be identified based on the extent of similarity to one another as quantified by an edit distance. Doing so may determine how similar two or more embeddings must be before the corresponding protein sequences are identified as related protein sequences (or neighbors within the same cluster). In the examples shown in FIG.13 A, a lower threshold for edit distance may impose a higher requirement for how similar two or embeddings must be before the corresponding protein sequences are identified as related protein sequences. For example, when the threshold for edit distance is set to 1, the embeddings of two related protein sequences cannot differ in more than one position.

[0072] Accordingly, as shown in FIG. 13 A, setting the threshold for edit distance to different fixed values may alter the morphology of the resulting clusters of related protein sequences. For instance, few protein sequences qualify as related protein sequences when the threshold for edit distance is fixed to 1. This phenomenon is observed in the visual representation 1310, in which an entire population of protein sequences are shown as a single, amorphous clusters with little differentiation. Contrastingly, when the threshold distance is set to incrementally higher values, the visual representation of the groups of related protein sequences depicts increasingly smaller and more modular clusters. For example, when the threshold for edit distance is set to 5, the corresponding visual representation 1318 depicts several clusters of protein sequences withNAI-5007186271V1 81Attorney Ref.: 14786-058-228 (103963-228058) / P39076-WO-1less defined boundaries than when the threshold is set to 6 in the visual representation 1320. However, the boundaries between the clusters in the visual representation 1318 is more defined than when the threshold for edit distance is set to 2 in the visual representation 1312, 3 in the visual representation 1314, and 4 in the visual representation 1316. When the threshold for edit distance is set to 7, meaning that the embeddings of two related protein sequences can differ in up to 7 positions, the resulting visual representation 1322 shows numerous small and diffuse clusters.

[0073] A similar shift in cluster morphology may also occur as a result of adjusting the average number of nearest neighbors for each protein sequence. This phenomenon is illustrated in FIG. 13B, which shows the clusters of protein sequences that are generated with the average quantity of nearest neighbors for each protein sequence set at 1 (in the visual representation 1330), 3 (in the visual representation 1332), 5 (in the visual representation 1334), 10 (in the visual representation 1336), and 600 (in the visual representation 1338). As shown in FIG. 13B, increasing the average quantity of nearest neighbors, along with a concomitant increase in the threshold for edit distance, increases the modularity of individual clusters of related protein sequences. For example, at a lower quantity of nearest neighbors and a lower threshold for edit distance, the protein sequences form a single, substantially monolithic cluster. At a higher quantity of nearest neighbors and a higher threshold for edit distance, the boundaries between different clusters become more refined. When the quantity of nearest neighbors is at 500 and the threshold for edit distance is at 7, the protein sequences form numerous small and diffuse clusters.

[0074] FIG. 13C depicts the effects of assigning different weights to an embedding emphasizing the similarities in one or more specific regions of a protein sequence when generating a hybrid embedding that combines the embedding and an additional embedding generalizing the one or more specific regions. As noted, in some cases, the embedding for a protein sequence mayNAI-5007186271V1 82Attorney Ref.: 14786-058-228 (103963-228058) / P39076-WO-1be a hybrid embedding that is a combination (e.g., union) of an embedding emphasizing the similarities in one or more specific regions (e.g., CDR3) of the protein sequence and an additional embedding generalizing the one or more specific regions (e.g., by amino acid groups). In some cases, the embedding may be assigned a weight to determine the extent to which the embedding contributes to the hybrid embedding relative to the additional embedding. For example, assigning a higher weight to the embedding than the additional embedding when integrating the two embeddings may result in a hybrid embedding that emphasizes the similarities in the one or more specific regions of the protein sequence over the generalization of the one or more specific regions. FIG. 13C shows that assigning a different weight to the embedding may also affect the clusters formed based on the corresponding hybrid embeddings. For instance, assigning a weight of 0 to the embedding, which effectively generates the clusters based on the additional embedding alone, generates different clusters of related protein sequences (in the visual representation 1340) than if the embedding is assigned a weight of 1 (in the visual representation 1342), a weight of 2 (in the visual representation 1344), and a weight of 5 (in the visual representation 1346).

[0075] The performance of protein analysis leveraging embeddings that emphasize similarities in one or more specific regions (e.g., CDR3) of each protein sequence was compared to that of conventional methodologies such as uniform manifold approximation and projection (UMAP) and t-distributed Stochastic Neighbor Embedding (t-SNE). FIG. 14A depicts a first visual representation 1400 of the clusters of related protein sequences identified based on embeddings emphasizing similarities in one or more specific regions (e.g., CDR3) of each protein sequence. In particular, the clusters of related protein sequences shown in the first visual representation 1400 are identified based on embeddings generated by an embedding computation model (e.g., the encoder of a language model having a transformer architecture) that has beenNAI-5007186271V1 83Attorney Ref.: 14786-058-228 (103963-228058) / P39076-WO-1trained to generate similar embeddings for protein sequences exhibiting a threshold degree of similarities (e.g., threshold edit distance) in the one or more specific regions and dissimilar embeddings for protein sequences failing to exhibit the threshold degree of similarities (e.g., threshold edit distance) in the one or more specific regions. Contrastingly, the second visualization 1402 and the third visualization 1404 show clusters of related sequences identified by uniform manifold approximation and projection (UMAP) and t-distributed Stochastic Neighbor Embedding (t-SNE), respectively. As shown in FIG. 14A, the use of embeddings emphasizing similarities in one or more specific regions (e.g., CDR3) of each protein sequence yields more clearly defined clusters than the conventional methods. With the addition of visual indicators identifying one or more properties of interest for each protein sequence, in this case binding affinity, the clusters of related sequences in the first visual representation 1400 more clearly delineate the distribution of the one or more properties of interest, in this case the distribution of binding protein sequences, weakly binding protein sequences, and non-binding protein sequences, across the clusters of related protein sequences. In FIG. 14B, the fourth visual representation 1410, the fifth visual representation 1412, and the sixth visual representation 1414 again show clusters of related protein sequences identified using embeddings emphasizing similarities in one or more specific regions (CDR3) and those identified using the conventional methodologies uniform manifold approximation and projection (UMAP) and t-distributed Stochastic Neighbor Embedding (t-SNE). In FIG. 14B, each protein sequence is annotated with a visual indicator of the degree of similarity (e.g., edit distance) to a reference sequence denoted as “X” in each visual representation. The quality of the embedding for a protein sequence may be evidenced by a correlation between the edit distance to the reference sequence and the distance (e.g., Euclidean distance) between their respective embeddings. As shown in FIG. 14B, this correlation is bestNAI-5007186271V1 84Attorney Ref.: 14786-058-228 (103963-228058) / P39076-WO-1observed in the visual representation 1410, which includes clusters of related protein sequences identified based on embeddings that emphasize similarities in one or more specific regions (e.g., CDR3) of each protein sequence.

[0076] FIG. 15 depicts a comparison of embeddings occupying a hyperbolic space and embeddings occupying a two-dimensional Euclidean space, in accordance with some example embodiments. As shown in FIG. 15(a), the amino acid residue sequences of the third complementarity determining region (CDR3) of antibodies targeting human epidermal growth factor receptor 2 (HER2) were clustered using uniform manifold approximation and projection (UMAP), t-distributed Stochastic Neighbor Embedding (t-SNE), and various embodiments of the embedding computation model described herein. More specifically, the heavy and light chain CDR3 sequences were clustered using uniform manifold approximation and projection (UMAP) on the Antibody Language Model (AbLang) embeddings, uniform manifold approximation and projection (UMAP) on conventional language model (LM) embeddings, t-distributed Stochastic Neighbor Embedding (t-SNE) on conventional language model (LM) embeddings, two-dimensional Euclidean space embeddings generated by the embedding computation model (deepNGS), and hyperbolic space embeddings generated by the embedding computation model (deepNGS (hyperbolic)). FIG. 15(b) shows the differentiation of non-binders, weak-binders, and binders of HER2 amongst the clusters generated using the aforementioned embeddings. FIG.15(d) depicts a violin chart illustrating the distribution of edit distances in each cluster while the entropy of non-binders, weak-binders, and binders labels in each cluster is shown in the violin chart depicted in FIG. 15(e).

[0077] FIG. 16 depicts graphs illustrating a comparison of the performance of a generative model trained on a fill-in-the-middle (FIM) objective and further finetuned onNAI-5007186271V1 85Attorney Ref.: 14786-058-228 (103963-228058) / P39076-WO-1hierarchical loss, in accordance with some example embodiments. FIG. 16 depicts a comparison of generative performance based on a total frequency of seeds per edit distance, binding rate per edit distance, and prediction probability distribution per edit distance. The generative performance of the generative model trained on different datasets as well as the generative performance of a state-of-the-art-model AbLang2 are evaluated in FIG. 16.

[0078] FIG. 17 depicts a block diagram illustrating an example of computing system 1700, in accordance with some example embodiments. Referring to FIGS. 1-17, the computing system 1700 may be used to implement the analysis controller 110, the sequencing platform 120, the client device 130, and / or any components therein.

[0079] As shown in FIG. 17, the computing system 1700 can include a processor 1710, a memory 1720, a storage device 1730, and an input / output device 1740. The processor 1710, the memory 1720, the storage device 1730, and the input / output device 1740 can be interconnected via a system bus 1750. The processor 1710 is capable of processing instructions for execution within the computing system 1700. Such executed instructions can implement one or more components of, for example, the analysis controller 110, the sequencing platform 120, the client device 130, and / or the like. In some example embodiments, the processor 1710 can be a singlethreaded processor. Alternatively, the processor 1710 can be a multi -threaded processor. The processor 1710 is capable of processing instructions stored in the memory 1720 and / or on the storage device 1730 to display graphical information for a user interface provided via the input / output device 1740.

[0080] The memory 1720 is a computer readable medium such as volatile or non-volatile that stores information within the computing system 1700. The memory 1720 can store data structures representing configuration object databases, for example. The storage device 1730 isNAI-5007186271V1 86Attorney Ref.: 14786-058-228 (103963-228058) / P39076-WO-1capable of providing persistent storage for the computing system 1700. The storage device 1730 can be a floppy disk device, a hard disk device, an optical disk device, or a tape device, or other suitable persistent storage means. The input / output device 1740 provides input / output operations for the computing system 1700. In some example embodiments, the input / output device 1740 includes a keyboard and / or pointing device. In various implementations, the input / output device 1740 includes a display unit for displaying graphical user interfaces.

[0081] According to some example embodiments, the input / output device 1740 can provide input / output operations for a network device. For example, the input / output device 1740 can include Ethernet ports or other networking ports to communicate with one or more wired and / or wireless networks (e.g., a local area network (LAN), a wide area network (WAN), the Internet).

[0082] In some example embodiments, the computing system 1700 can be used to execute various interactive computer software applications that can be used for organization, analysis and / or storage of data in various formats. Alternatively, the computing system 1700 can be used to execute any type of software applications. These applications can be used to perform various functionalities, e.g., planning functionalities (e.g., generating, managing, editing of spreadsheet documents, word processing documents, and / or any other objects, etc.), computing functionalities, communications functionalities, etc. The applications can include various add-in functionalities or can be standalone computing products and / or functionalities. Upon activation within the applications, the functionalities can be used to generate the user interface provided via the input / output device 1740. The user interface can be generated and presented to a user by the computing system 1700 (e.g., on a computer screen monitor, etc.).NAI-5007186271V1 87Attorney Ref.: 14786-058-228 (103963-228058) / P39076-WO-1

[0083] One or more aspects or features of the subject matter described herein can be realized in digital electronic circuitry, integrated circuitry, specially designed ASICs, field programmable gate arrays (FPGAs) computer hardware, firmware, software, and / or combinations thereof. These various aspects or features can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device. The programmable system or computing system may include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other.

[0084] These computer programs, which can also be referred to as programs, software, software applications, applications, components, or code, include machine instructions for a programmable processor, and can be implemented in a high-level procedural and / or object-oriented programming language, and / or in assembly / machine language. As used herein, the term “machine-readable medium” refers to any computer program product, apparatus and / or device, such as for example magnetic discs, optical disks, memory, and Programmable Logic Devices (PLDs), used to provide machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term “machine-readable signal” refers to any signal used to provide machine instructions and / or data to a programmable processor. The machine-readable medium can store such machine instructions non-transitory, such as for example as would a non-transient solid-state memory or aNAI-5007186271V1 88Attorney Ref.: 14786-058-228 (103963-228058) / P39076-WO-1magnetic hard drive or any equivalent storage medium. The machine-readable medium can alternatively or additionally store such machine instructions in a transient manner, such as for example, as would a processor cache or other random access memory associated with one or more physical processor cores.

[0085] To provide for interaction with a user, one or more aspects or features of the subject matter described herein can be implemented on a computer having a display device, such as for example a cathode ray tube (CRT) or a liquid crystal display (LCD) or a light emitting diode (LED) monitor for displaying information to the user and a keyboard and a pointing device, such as for example a mouse or a trackball, by which the user may provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well. For example, feedback provided to the user can be any form of sensory feedback, such as for example visual feedback, auditory feedback, or tactile feedback; and input from the user may be received in any form, including acoustic, speech, or tactile input. Other possible input devices include touch screens or other touch-sensitive devices such as single or multi-point resistive or capacitive track pads, voice recognition hardware and software, optical scanners, optical pointers, digital image capture devices and associated interpretation software, and the like.

[0086] In the descriptions above and in the claims, phrases such as “at least one of’ or “one or more of’ may occur followed by a conjunctive list of elements or features. The term “and / or” may also occur in a list of two or more elements or features. Unless otherwise implicitly or explicitly contradicted by the context in which it is used, such a phrase is intended to mean any of the listed elements or features individually or any of the recited elements or features in combination with any of the other recited elements or features. For example, the phrases “at least one of A and B;” “one or more of A and B;” and “A and / or B” are each intended to mean “A alone,NAI-5007186271V1 89Attorney Ref.: 14786-058-228 (103963-228058) / P39076-WO-1B alone, or A and B together.” A similar interpretation is also intended for lists including three or more items. For example, the phrases “at least one of A, B, and C;” “one or more of A, B, and C;” and “A, B, and / or C” are each intended to mean “A alone, B alone, C alone, A and B together, A and C together, B and C together, or A and B and C together.” Use of the term “based on,” above and in the claims is intended to mean, “based at least in part on,” such that an unrecited feature or element is also permissible.

[0087] The subject matter described herein can be embodied in systems, apparatus, methods, and / or articles depending on the desired configuration. The implementations set forth in the foregoing description do not represent all implementations consistent with the subject matter described herein. Instead, they are merely some examples consistent with aspects related to the described subject matter. Although a few variations have been described in detail above, other modifications or additions are possible. In particular, further features and / or variations can be provided in addition to those set forth herein. For example, the implementations described above can be directed to various combinations and subcombinations of the disclosed features and / or combinations and subcombinations of several further features disclosed above. In addition, the logic flows depicted in the accompanying figures and / or described herein do not necessarily require the particular order shown, or sequential order, to achieve desirable results. Other implementations may be within the scope of the following claims.NAI-5007186271V1 90

Claims

Attorney Ref.: 14786-058-228 (103963-228058) / P39076-WO-1CLAIMSWhat is claimed is:

1. A computer-implemented method, comprising:determining a selected protein sequence comprising a plurality of amino acid residues; determining two or more protein sequences that exhibit a threshold similarity in a specific region of each protein sequence;applying an embedding computation model to generate an embedding of the selected protein sequence, where the embedding computation model has been trained to generate embeddings for the two or more determined protein sequences;identifying, based at least on the embedding of the selected protein sequence, one or more related groups of protein sequences; andgenerating, based at least on the one or more identified groups of related protein sequences, a visual representation.

2. The method of claim 1, wherein the embedding computation model has been trained such that an edit distance between the two or more similar embeddings satisfy one or more threshold criteria.

3. The method of claim 2, further comprising:determining two or more dissimilar protein sequences, the dissimilar protein sequences comprising sequences that fail to exhibit the threshold similarity in the specific region of each protein sequence; andtraining the embedding computation model to generate, for the two or more dissimilar protein sequences, two or more embeddings whose edit distance satisfy a different threshold criteria.NAI-5007186271V1 91Attorney Ref.: 14786-058-228 (103963-228058) / P39076-WO-14. The method of any of claims 1 to 3, further comprising:determining two or more protein sequences that exhibit a threshold similarity across an entire length of each protein sequence;training the embedding computation model to generate, for the two or more protein sequences with the threshold similarity across the entire length of each protein sequence, similar embeddings whose edit distance satisfy one or more threshold criteria.

5. The method of claim 4, further comprising:determining two or more protein sequences that fail to exhibit the threshold similarity across an entire length of each protein sequence; andtraining the embedding computation model to generate, for the two or more protein sequences without the threshold similarity across the entire length of each protein sequence, two or more dissimilar embeddings whose edit distance satisfy a different threshold criteria.

6. The method of any of claims 1 to 5, further comprising:generating, for the selected protein sequence, an additional embedding that generalizes, to a corresponding amino acid group, an identity of each amino acid residue in the specific region.

7. The method of claim 6, further comprising:generating a hybrid embedding comprising a combination of the embedding and the additional embedding; andidentifying, based at least on the hybrid embedding, the one or more related groups of protein sequences.

8. The method of claim 7, wherein the combination comprises a union of the embedding and the additional embedding in which the embedding is weighted by one weight and the additional embedding is weighted by a different weight.NAI-5007186271V1 92Attorney Ref.: 14786-058-228 (103963-228058) / P39076-WO-19. The method of any of claims 6 to 8, wherein the additional embedding is generated by at leastassigning, to an amino acid residue group, each amino acid residue in the specific region of the selected protein sequence, andgenerating the additional embedding to include, for each amino acid residue in the specific region of the selected protein sequence, the amino acid residue group assigned to the amino acid residue.

10. The method of claim 9, wherein the amino acid residue group is one of a nonpolar group, a polar group, an aromatic group, a positively charged (or basic) group, a negatively-charged (or acidic) group, or a special group.

11. The method of any of claims 1 to 10, wherein the one or more related groups of protein sequences are identified by at least determining, based at least on the embedding of the selected protein sequence, a cluster label assigning the selected protein sequence to a group of related protein sequences.

12. The method of any of claims 1 to 11, wherein the identifying of one or more related groups of protein sequences includes applying a clustering technique to a plurality of embeddings of a plurality of protein sequences including the embedding of the protein sequences.

13. The method of claim 12, wherein the clustering technique comprises one of a hierarchal clustering technique, a centroid-based clustering technique, a density-based clustering technique, or a distribution-based clustering technique.

14. The method of any of claims 12 to 13, wherein the plurality of protein sequences comprise a repertoire of antibodies generated in response to exposure to a pathogen.NAI-5007186271V1 93Attorney Ref.: 14786-058-228 (103963-228058) / P39076-WO-115. The method of any of claims 1 to 14, wherein the determining the selected protein sequence includes translating, into the plurality of amino acid residues, one or more corresponding coding sequences comprising raw sequencing data.

16. The method of any of claims 1 to 15, wherein the visual representation is generated to depict, for each related group of protein sequences, a corresponding cluster.

17. The method of any of claims 1 to 16, wherein the visual representation is generated by at least projecting the one or more groups of protein sequences into a two-dimensional or a three-dimensional space.

18. The method of any of claims 1 to 17, wherein the visual representation is generated to include a visual indicator representative of the selected protein sequence, and wherein the visual indicator is rendered in a color, shape, and / or size representative of one or more properties of the selected protein sequence.

19. The method of any of claims 1 to 18, wherein the threshold similarity is quantified by an edit distance.

20. The method of any of claims 1 to 19, wherein the specific region includes a complementarity determining region (CDR).

21. The method of claim 20, wherein the embedding computation model has been trained to prioritize similarities in the complementarity determining region (CDR) of the selected protein sequence over similarities in a framework region of the selected protein sequence when generating the embedding of the selected protein sequence.

22. The method of any of claims 1 to 21, wherein the specific region includes a third complimentarity determining region (CDR3).

23. The method of any of claims 1 to 22, wherein the embedding computation modelNAI-5007186271V1 94Attorney Ref.: 14786-058-228 (103963-228058) / P39076-WO-1comprises an encoder of a language model having a transformer architecture.

24. The method of any of claims 1 to 23, wherein the embedding computation model has been trained to generate embeddings occupying a two-dimensional Euclidean space.

25. The method of any of claims 1 to 23, wherein the embedding computation model has been trained to generate embeddings occupying a hyperbolic space.

26. The method of any of claims 1 to 25, wherein the embedding computation model has been trained to generate the embedding of the selected protein sequence to lie below an embedding of a parent protein sequence of the selected protein sequence in a latent space occupied by a plurality of embeddings generated by the embedding computation model.

27. The method of claim 26, wherein the embedding computation model has been trained to generate the embedding of the selected protein sequence to lie above an embedding of a child protein sequence of the selected protein sequence in latent space.

28. The method of any of claims 1 to 27, wherein the embedding computation model has been trained with contrastive learning to generate the similar embeddings for the two or more similar protein sequences exhibiting the threshold similarity in the specific region and dissimilar embeddings for two or more dissimilar protein sequences without the threshold similarity in the specific region.

29. The method of claim 28, wherein the embedding computation model has been further finetuned to reduce an angle between embeddings of two or more related protein sequences exhibiting a parent-child hierarchical relationship and increase an angle between embeddings of two or more unrelated protein sequences absent the parent-child hierarchical relationship.

30. The method of any of claims 1 to 29, further comprising:NAI-5007186271V1 95Attorney Ref.: 14786-058-228 (103963-228058) / P39076-WO-1determining, based at least on an identified group of related protein sequences, a generative context; andapplying a generative model to generate, based at least on the generative context, one or more novel protein sequences.

31. The method of claim 30, wherein the generative model has been trained on a fill-in-the-middle (FIM) objective to generate the one or more novel protein sequences while leveraging a plurality of protein sequences forming the generative context.

32. The method of claim 31, wherein the generative model has been finetuned with direct preference optimization (DPO) to generate the one or more novel protein sequences by at least further evolving a subset of protein sequences occupying a region of interest in the generative context.

33. The method of claim 32, wherein the subset of protein sequences include children protein sequences exhibiting a greater distance to a germline and / or a greater edit distance to a naive repertoire than one or more parent protein sequences outside of the region of interest in the generative context.

34. A system, comprising:at least one data processor; andat least one memory storing instructions, which when executed by the at least one data processor, result in operations comprising the method of any of claims 1 to 33.

35. A non-transitory computer readable medium storing instructions, which when executed by the at least one data processor, result in operations comprising the method of any of claims 1 to 33.

36. A computer-implemented method, comprising:NAI-5007186271V1 96Attorney Ref.: 14786-058-228 (103963-228058) / P39076-WO-1identifying a pair of protein sequences as a positive pair based at least on the pair of protein sequences exhibiting a threshold similarity between a specific region of each protein sequence of the pair of protein sequences;identifying the pair of protein sequences as a negative pair based at least on the pair of protein sequences failing to exhibit the threshold similarity in the specific region of each protein sequence of the pair of protein sequences;training an embedding computation model by at least adjusting one or more parameters of the embedding computation model to generate embeddings for each positive pair and embeddings for each negative pair; andapplying the trained embedding computation model to generate an embedding for a protein sequence.

37. The method of claim 36, wherein the pair of protein sequence is identified as the positive pair further based at least on the pair of protein sequences exhibiting the threshold similarity across an entire length of each protein sequence.

38. The method of claim 37, wherein the pair of protein sequences are identified as the negative pair further based at least on the pair of protein sequences failing to exhibit the threshold similarity across the entire length of each protein sequence.

39. The method of any of claims 36 to 38, wherein the one or more parameters of the embedding computation model are adjusted to reduce an edit distance between embeddings of each positive pair.

40. The method of claim 39, wherein the one or more parameters of the embedding computation model are further adjusted to increase the edit distance between embeddings of each negative pair.NAI-5007186271V1 97Attorney Ref.: 14786-058-228 (103963-228058) / P39076-WO-141. The method of any of claims 36 to 40, wherein the threshold similarity is quantified by an edit distance.

42. The method of any of claims 36 to 41, wherein the embedding computation model comprises an encoder of a language model having a transformer architecture.

43. The method of any of claims 36 to 42, wherein the specific region includes a complementarity determining region (CDR).

44. The method of claim 43, wherein the embedding computation model is trained to prioritize similarities in the complementarity determining regions (CDRs) of each pair of protein sequences over similarities in a framework regions each pair of protein sequences.

45. The method of any of claims 36 to 44, wherein the specific region includes a third complementarity determining region (CDR3).

46. The method of any of claims 36 to 45, further comprising:projecting a plurality of protein sequences into a two-dimensional or a three-dimensional space;identifying, for a protein sequence of the plurality of protein sequences, one or more protein sequences within a threshold distance of the protein sequence as one or more neighboring protein sequences; andidentifying the protein sequence and a neighboring protein sequence as the positive pair based at least on the protein sequence and the neighboring protein sequence exhibiting the threshold degree of similarity in the specific region of each protein sequence.

47. The method of claim 46, further comprising:excluding the protein sequence from being a part of any positive pairs and negative pairs based at least on the protein sequence failing to have a threshold quantity of neighboring proteinNAI-5007186271V1 98Attorney Ref.: 14786-058-228 (103963-228058) / P39076-WO-1sequences.

48. The method of any of claims 36 to 47, further comprising:identifying the pair of protein sequences as the positive pair further based at least on the pair of protein sequences including a protein sequence and a child protein sequence of the protein sequence; andfinetuning the embedding computation model by further adjusting the one or more parameters of the embedding computation model to generate an embedding of the child protein sequence to lie beneath an embedding of the protein sequence in a latent space populated by a plurality of embeddings generated by the embedding computation model.

49. The method of any of claims 36 to 48, further comprising:identifying the pair of protein sequences as the positive pair further based at least on the pair of protein sequences including a protein sequence and a parent protein sequence from which the protein sequence descends; andfinetuning the embedding computation model by further adjusting the one or more parameters of the embedding computation model to generate an embedding of the protein sequence to lie beneath an embedding of the parent protein sequence in a latent space populated by a plurality of embeddings generated by the embedding computation model.

50. The method of any of claims 36 to 49, further comprising:identifying the pair of protein sequences as the negative pair further based at least on the pair of protein sequences including two unrelated protein sequences; andfinetuning the embedding computation model by further adjusting the one or more parameters of the embedding computation model to generate an embedding of one unrelated protein sequence to lie away from an embedding of the other unrelated protein sequence in aNAI-5007186271V1 99Attorney Ref.: 14786-058-228 (103963-228058) / P39076-WO-1latent space populated by a plurality of embeddings generated by the embedding computation model.

51. The method of any of claims 36 to 50, wherein the one or more parameters of the embedding computation model are adjusted such that the embedding computation model generates similar embeddings for each positive pair and dissimilar embeddings for each negative pair.

52. The method of any of claims 36 to 51, wherein the embedding computation model is trained to generate embeddings occupying a two-dimensional Euclidean space.

53. The method of any of claims 36 to 51, wherein the embedding computation model is trained to generate embeddings occupying a hyperbolic space.

54. A system, comprising:at least one data processor; andat least one memory storing instructions, which when executed by the at least one data processor, result in operations comprising the method of any of claims 36 to 53.

55. A non-transitory computer readable medium storing instructions, which when executed by at least one data processor, result in operations comprising the method of any of claims 36 to 53.NAI-5007186271V1 100